DETAILED ACTION
This action is responsive to the Applicant’s response filed 7/27/26.
As indicated in Applicant’s response, claims 1-3, 10-11, 18 have been amended. Claims 1-20 are pending prosecution by the following office action.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. Claim(s) 1 is/are directed to Abstract Idea The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because of the 2-step analysis as follows
Step I: Claim 1 is directed to an apparatus category
Step 2A
Prong 1:
The core limitation - analyzing source code for an "absence of cycles with at least one loop-carried true dependency", "marking" it, and "offloading" it -falls under the Mental Processes and Mathematical Concepts/Certain Organizing Human Activities concepts.
That is, determining loop dependencies, marking code, and deciding where to run code are logical operations that can theoretically be performed in the human mind or on paper with a set of rules. For instance, the steps recited as marking (a portion of the source code) can be viewed as activity that can be performed by a human mind or using pen and paper. MPEP 2106.04(a)(2) – and deemed analogous to identifying and categorizing data similar to rendering opinions, re-arrangement, performing judgement or observations by a human via a generic computer or via use of pen/paper.
Evaluating whether a loop contains a loop-carried true dependency is a logical analysis and data evaluation concept that can be performed in the human mind or by a human using pencil and paper. MPEP 2106.04(a)(2)
"Marking" and "offloading" data based on evaluated conditions represent generic acts of organizing, classifying, and routing data. MPEP 2106.04(a)(1)
Claim 1 is directed to a Judicial Exception of a Abstract Idea type.
Prong 2:
The elements recited as "compiler executing" and "offloading" (of code portion) can be viewed respectively as pre-activity to collect or prepare data for use by the mental process of "marking" and as a post-activity that makes use of data/result from the mental process via use of a generic computer.
The activity of preparing or pre-converting data and dispatching mentally processed data amount to a mere extra-activity of insignificant impact to a field of computer technology, or well-understood routines - see MPEP 2106 (d)(1) - which are not viewed as actually transforming a computer technology into a improvement in this field or inventive state thereof.
Merely executing a compiler on generic hardware (a standard processor core and PIM units) and using "marking" to offload code does not automatically integrate the concept into a practical application.
Per MPEP 2106.04(d)(1) The abstract mental process of code dependency analysis and marking is merely performed by a generic "compiler executing on the processor core."
Per MPEP 2106.04(d)(2), it is not apparent that the claim particularly recite how the compiler or processor core is physically modified, nor does it detail a specific improvement to the technical operation of the memory or PIM units themselves. Instead, the claim merely uses the existing compiler and processor core as a tool to execute the abstract logic of identifying loop dependencies and tagging code.
The claim does not integrate the abstract idea into a practical application
Step 2B
The recited structural components - "a memory", "processing-in-memory units", and "a processor core" – as additional elements are described generically, as these are well-understood, routine, and conventional computer operations in the compiler and parallel computing arts. MPEP 2106.05(d) and simply applying an abstract Idea using a conventional computer or a generic mention of a compiler is insufficient -MPEP 2106.05(a) – to render the Abstract Idea significantly more than itself.
Offloading execution to secondary processors (such as PIM units) based on dependency conditions is a conventional scheduling technique and does not impart an "inventive concept." The recited "offloading portion based on the marking" (as additional element) is considered a post-activity or extra-solution of insignificance- MPEP 2106.05(g) - since once the abstract Idea of marking has been made, the act of sending mental process data into a destination (PIM units) is viewed as routine or conventional consequence to the mental process - MPEP 2106.05(f); the “offloading” the additional elements are identified cannot amount to significantly more than the Judicial Exception found in step IIA.
The above-mentioned additional elements fail to provide a "inventive concept" to the Judicial Exception.
Considering the claim elements individually and as an ordered combination, the claim fails to recite significantly more than the abstract idea itself. Therefore, Claim 1 is non-statutory under 35 U.S.C. 101.
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a
judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without
significantly more. Claim 11 is/are directed to Abstract Idea. The claim(s) does/do not include
additional elements that are sufficient to amount to significantly more than the judicial exception
because of the 2-step analysis as following.
Step I:
The claim is directed to a process/method as statutory category
Step 2A
Prong 1:
The step of 'generating a read dependence graph' (representing one or more chains of dependent elements" and establishing "memory capacity being greater than or equal to number of elements represented by a longest chain" are acts that can be practically performed in the human mind or with the aid of pen and paper. MPEP 2106.04(a) – That is, a human programmer can look at a source code and manually trace the dependencies between data structures ("chains") and determine the length of those structure; e.g. a human can compare length to a known memory threshold (or "bank capacity"). As the above 'generating' and establishing memory capacity are viewed as concepts of organizing data and performing comparison; therefore, they fall within the Abstract Idea grouping of Mental Processes
Prong 2:
The step of "compiling" (portion of source code) is mere insignificant extra- solution activity, whereas compiling is a well-understood routine function of any generic system that serves as a stage for the Abstract dependency analysis. The processing-in-memory element expressed generically without details beyond its generic functional name together with the step of "offloading" (portion of code for execution) construed as an extra-solution or post-activity of insignificant impact cannot be viewed as capable of transforming the Abstract Idea into a patent-eligible Application, notably when the claim does not recite a specific change to the PIM hardware or specific improvement in the way performing PIM processes data.
The claim is more directed to logical dependence comparison than depicting a clear focus on a technical constraint, improvement associated with hardware (Process-in-Memory, processor); hence the elements of the claim cannot integrate the Abstract Idea into a Practical Application. The claim as such does not provide a "technical improvement' to a computer itself under MPEP 2106.04(d)(1).
Step 2B
The additional elements recited as "PIM units" and the "memory banks" are described in a most generic sense. The tracing of dependencies data (as additional element) during compilation is viewed as fundamental, well-understood practice in computer arts, whereas use of "memory threshold" is a routine in resource management. Looking at the elements recited as "compiling" and "offloading" of portion of code that is being traced via pen/paper as chains of dependencies, the claim as a whole simply invokes action sequence such as to take source code, find dependencies (Abstracted part) and if the data fits (conventional routine), send it to a processor, which all together, constitutes a conventional practice and comparing data or finding dependencies thereof and send that data to a destination; and as such, the claim construed in the recited order of the above elements do not amount to "significantly more" than the Abstract Idea itself. MPEP 2106.05(a)(b)(c)(e )
Claim 11 is deemed non-eligible under the 35 USC § 101 statute
Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without
significantly more. Claim 18 is/are directed to Abstract Idea. The claim(s) does/do not include
additional elements that are sufficient to amount to significantly more than the judicial exception
because of the 2-step analysis as following.
Step I:
Claim 18 is directed to a system/apparatus category.
Step 2A
Prong one:
The action of computing a metric (“duplication metric”) in capturing an amount of data duplication can be seen as a mathematical exercise that quantifies a relationship between data sets and under MPEP 2106.04(a)(1) this amount to a mathematical concept
In other words, the computing of a metric and comparing the metric with respect to a “threshold” so to decide whether to offload are viewed as acts that can be practically performed in the human mind or via use of pen/paper, in that, a human can inspect a portion of code, count duplicate data entries an apply a logical "if-then" rule to determine destination of that portion.
Prong two:
The element recited as “compiling” is a well-understood routine that pre-processes data to prepare them for an abstract calculation; whereas the step of "offloading” to PIM units based on the “duplication metric" is a mere post-solution activity, the latter being a generic function which cannot transform the Abstract Idea into a technical improvement of significance, notably when the claim does not describe how internal operation of the PIM units and memory banks is altered or improved. Instead, the claim is - under Electric Power Group LLC vs Alstom SA - about collection information, analyzing it and displaying or routing the results which is clearly not a improvement to computer functionality; hence the claimed compiling and offloading cannot integrate the Abstract Idea of Step IIA into a practical application. MPEP 2106.04(d)(1)
Step 2B
The components (as additional elements) recited as process-in-memory (PIM) units and “memory banks” are recited in a very high level of generality; whereas identifying “data duplication” is a long standing, conventional practice in computer sciences. Viewed as a whole, the compiling as a standard code preparation, the abstract calculation of a metric (mathematical concepts) and the conventional routing decision (i.e. offloading) without specific details cannot in combination add anything extraordinary beyond the Abstract Idea itself, as the claim only instructs the user to apply the abstract Idea of duplication analysis within the well-known environment of PIM architecture without explicit showing how the PIM units handle memory or execute instructions that are not already inherent in their conventional design. MPEP 2106.05(a)(b)(c)(e )
Claim 18 is therefore a non-eligible subject matter under the 35 USC § 101 statute
Analysis under step IIB of the dependent claims
Claim 2: the generating of dependence graph as a conventional routine to organize data
cannot add significantly more to the Abstract Idea of claim 1.
Claim 3: describes loop specificity associated with the dependency graph hence cannot be
viewed as internal computer-based improvement over the Abstract Idea of claim 1.
Claims 4 and 12: describes generating first data structures, second data structures and linked
data structures, which can be viewed as human arrangement of data (source code) by way of a mental process or via use of pen and paper.
Claim 5: describes marking based memory capacity with establishing a size comparison being well-understood routine in mental activities of organizing data, thus cannot add significantly more to the Abstract Idea itself as raised by Step 2A
Claims 6 and 13: describe basis of marking via a metric, hence a marking as conventional or well-known routine in mental data organizing/processing cannot add significantly more to the "marking" of code portion in the base claim.
Claims 7 and 20: describe duplication metric based on a comparison of elements represented
on a dependence graph, hence this limitation does not technically implement significantly more to the Abstract Idea of "marking" from the base claim.
Claim 8: describes basis of marking in terms of comparing graph elements and memory
elements; and this, as set forth above, cannot be seen a transformation to a computer technical field but rather as one more sub-functionality of a Abstract Idea.
Claim 9: recites marking use of a maximum number as basis; hence cannot be construed as
adding significantly more to the Abstract Idea of "marking"
Claim 10: describe marking of code portion and executing the portion, which can be
respectively viewed as a variety of a human process and post-activity using the result from that
process.
Claim 14: describes how a duplication metric is being based on from the dependence graph;
but using a metric as a means can be viewed as part of processes performed under an Abstract Idea, thus does not add significantly more to the Judicial Exception of base claim 11.
Claim 15: describes generating dependence graph and verifying absence of cycles, which in
all amount to activities than can be performed by a human mind or via pen and paper.
Claim 16: describes what a cycle amounts to and this cannot be seen as integrating the
Judicial Exception of claim 15 into a Practical Application.
Claim 17: describes verifying by comparing memory banks size with number of elements
acquired from the dependence graph; but verifying from size comparison does not add significantly
more to the abstract idea of claim 11.
Claim 19 describes the same generating steps of claim 4; hence cannot render the abstract
idea of the base claim significantly more than itself.
Claims 1-20 in all are deemed non-eligible under the 35 USC § 101 statute.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 15-16 is/are rejected under § 35 U.S.C. 103 as being unpatentable over Kalamatianos et al, WO 2023043711, 3-23-2023, 33 pgs (herein Kalamatianos) in view of Bertacco et al, USPubN: 2022/0019545, incorporating Nai: “GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks”, HPCA 2017- by reference, see para 0035 (herein Bertacco)
As per claims 1-3, Kalamatianos discloses a device, comprising: a memory that includes one or more processing-in-memory units (units 150 - para 0030; PIM execution unit - Fig. 2); a processor core (host processor 132 - para 0030; Fig. 3A, 3B) and
a compiler (para 0036, 0064-0065 – Note1: a work scheduler from a host computer - Fig. 5 - working with reservation of register through static analysis of a compiler - para 0065-0066 - and mapping command buffer with allocations destined for dispatch by the scheduler reads on compiler executing on the host environment in support for reserving resources for dispatch to a PIM execution) executing on the processor core, the compiler causing the processor core to perform operations including:
compiling source code (compiling - para 0064, 0067) of a software program to executable code;
during the compiling (see Note1), marking (tracking command buffer indices as invalid - para 0066; set of offloaded PIM instructions is marked by two special commands start command and end command - para 0036) a portion of the source code as suitable (instructions that will be written to the command buffer - para 0065) for execution using the one or more processing-in-memory units (indices in the command buffer that will be required to offload - par 0065; initiate an Offload of PIM instructions to a PIM device 410 - Fig. 6); and
offloading the executable code compiled from the portion of the source code (workload on the processor cores alleviated by offloading a PIM device - para 0025; for offloading PIM instructions for execution by the PIM units 150 - para 0034; completed offloading the PIM instructions - para 0035) for execution by the one or more processing-in-memory units (PIM execution units - para 0030) based on the marking (para 0036, 0066).
Kalamatianos does not explicitly disclose
(i) marking portion of the source code as suitable for execution using processing-in-memory units based on an absence of cycles with at least one loop-carried true dependency in the portion of the source code, as operations including generating a data dependence graph based on the portion of the source code, wherein the marking is based on the absence of cycles in the data dependence graph that include the at least one loop-carried true dependency.
(ii) wherein a cycle includes the at least one loop-carried true dependency based on a read access that is performed during a subsequent iteration of the cycle being dependent on a write access that is performed during a previous iteration of the cycle.
Similar to offloading off-chip execution to processing-in-memory execution per Kalamatianos, Bertacco discloses a GraphPIM-based approach of offloading (para 0030, 0035) all atomic operations to off-chip memory of a vertex management in combination with on-chip memory (para 0031, 0033) using a ranking algorithm (Fig. 1-2) of a graph analytics that assesses vertex data of the graph (para 0029; Fig. 7) to promulgate their data accesses with emphasis on favoring cache locality provided with identification of atomic operations for their off-chip memory (para 0030) execution as part of bypassing the cache coherence (para 0029) pressure from host CPU core; where each off-chip memory includes an atomic compute unit, a memory and a controller so that the atomic operations are handled by that compute unit using the respective memory module local to the off-chip memory, contributing to a high throughput typical to Processing-in-Memory execution (para 0003-0004) while minimizing read/write traffic with the host core; i.e. the graph-based vertex ranking to promulgate atomic operations at level of the off-chip memory as part of the cache coherence bypassing (para 0036) resulting in alleviating high-performance cost (when executed by the general host core) established between two memory requests, a read and a write (para 0031); in the sense that the off-chip execution would help minimize cache pollution and access latency that would otherwise exist in traffic between off-chip memory and the main core (para 0037)
Hence, analyzing code for maximizing atomic operations that can be carried out solely at the
level of off-chip memory entails scanning of data graph operations in heeding presence of possible
backward branch/loop by which a write to cache has to be effectuated before a newly stored data can be read again in the next iteration; i.e. without a cyclic dependency.
As for (ii)
The atomicity of operations that can be PIM offloaded as shown in Bertacco, is evidenced in Nail (incorporated by reference) motivation for the PIM, in which a PIM unit for such operation is slated to perform read-modify-write (RMW) atomically in a computation that locks away from other memory requests, such that the immediate value and a corresponding operand used with this atomic operation complete this atomicity in which the original memory data is returned along with the response and is not affected by more pending results (see Nail: II Background and Motivation, bottom L col. to top R col. pg. 2). Thus, atomic operation returned without depending on other memory requests in terms of indirectly nested iteration return values entails absence of a cycle in realizing a simple/direct RMW sequence. Accordingly, the scenario of true-dependency like Read-after-Write (RAW) can cause dependency across iterations, in which iteration i + 1 is strictly blocked until iteration i completes and potential bottlenecks and likely recourse thereby to the host CPU for intervention (see Nail: Bottlenecks in Graph Computing, bottom L col. to top R col., pg. 3) signify this cyclic dependency that will inhibit use of off-chip processing under a PIM mode endeavored in Bertacco, and this true dependency – e.g. RAW- cannot be typical a candidate sequence for a PIM offloading.
Thus, analyzing a data graph where the marking is based on acyclic operations (cache coherence with simple RMW realization) in the data dependence graph in terms of atomic operations with complete absence of at least one loop-carried true dependency (like a Read-after-Write) is recognized, in the sense that one such cycle is caused by a read access performed during a subsequent iteration of the cycle is being dependent on a write access that is performed during a previous iteration of the cycle.
Therefore, based on code marking with intent to employ local cache (para 001-0002) proximal to a local processing unit using memory banks allocated to respective PIM execution units (para 0030-0031) in Kalamatianos to alleviate additional traffic with the host core for updating its memory, it would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to implement analysis of code with data graph analytics - as in Bertacco via a GraphPIM associated algorithm that processes magnitude of data storing at the graph vertices - such that data dependence from traversing the graph nodes would enable marking as to where atomic operations (considered for offload) can be realized and internally completed in an atomic manner; i.e. without any cyclic dependency in form of a loop-carried true dependency – e.g. a RAW - from the graph interfering with completion of that atomic- e.g. a RMW - , and so, in the sense that a cyclic dependency (e.g. RAW) to be extricated entails a scenario where a read access performed during a subsequent iteration of the cycle is being dependent on a write access that is performed during a previous iteration of the cycle - as evidenced in Bertacco cache coherence and the atomicity of RMW sequence in Nail; because
marking code via a data graph analysis would enable identifying and consolidating the longest stretch of atomic operations that can be distributed as a long non-interrupted flow using local
processing unit (in-memory processor) and local memory bank provided as off-chip memory as intended with use of the off-loading in Kalamatianos for a desired off-memory processing throughput, and observance of any interruption thereof via graphPIM-based detection of presence of loop-carried true dependency (subsequent read depending on data value from write achieved from previous cycles) would enable marking and extrication of this form of dependency so to attain cache coherence, reduce performance cost on the host main CPU, and minimizing traffic complexities between the main CPU and its cache; the latter potentially causing a PIM execution to be halted, readjusted and otherwise deferred back to handling by the host core in accordance with the conventional Von Neuman approach, and so, in the sense that a flow of atomic operations at the level of local memory and PIM units - as intended in Bertacco - when disrupted by a read-after-write back loop associated with resolution of memory data access not only can hurt the desired off-memory throughput intended with a PIM execution offloading, but might also cause degradation by the main CPU, due to added traffic or latency that in turn would give way to unexpected cache pollution, jeopardize performance outcome of the offloaded execution approach, notwithstanding the fact that performance cost for this cache writeback can better off be resolved with a direct handover to the host core (on-chip memory) for this operation using a conventional Von Neuman approach.
As per claims 15-16, refer to rejection of claims 2-3 per rationale A in claim 1.
Claims 4-5, 8-12, 17 is/are rejected under § 35 U.S.C. 103 as being unpatentable over Kalamatianos et al, WO 2023043711, 3-23-2023, 33 pgs (herein Kalamatianos) in view of Bertacco et al, USPubN: 2022/0019545, incorporating Nai: “GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks”, HPCA 2017- by reference, see para 0035 (herein Bertacco) further in view of Lin et al, WO 2017076296 (translation), 05-11-2017, 17 pgs (herein Lin), Chang et al, USPubN: 2023/0119291 (herein Chang) and Alsop et al, CN 7063155(translation) 11-14-2023, 11 pgs (herein Alsop)
As per claims 4-5, Kalamatianos does not explicitly disclose device of claim 1, wherein the portion of the source code accesses a first data structure and a second data structure (see reading/writing - para 0030 ),
Kalamatianos does not explicitly disclose the operations further including:
(i) generating a first read dependence graph representing one or more first chains of dependent elements of the first data structure based on the portion of the source code;
generating a second read dependence graph representing one or more second chains of dependent elements of the second data structure based on the portion of the source code; and
generating a linked read dependence graph representing one or more linked chains of dependent elements by linking the one or more first chains with the one or more second chains based on the portion of the source code.
(ii) the marking is based on a memory capacity of a first number of banks communicatively coupled to respective ones of the one or more processing-in-memory units being greater than or equal to an amount of the memory to store a second number of elements represented by a longest chain of the one or more linked chains.
As for (i)
Implementation of a reduce map using traversal of a graph data is shown in Lin's graph processing and map simplification, whereby a Reduce phase processes the input data and intermediate calculation result, obtain simplified message thereof through a shuffle phase during which the intermediate results is/are taken out of the storage media (pg. 4). Hence, processing data result from processing data node of a data graph and reduce it by eliminating intermediate result from the node-edge flow is recognized.
Bertacco discloses identification of atomic operations by way of a GraphPIM analysis for forming an off-loadable sequence into a off-chip memory execution(memory 14 – Fig. 1; para 0035), with emphasis on bypassing or averting a cache coherence (para 0029-0030) required under Von Neuman architecture that causes additional delay/traffic not favorable to performance throughput (para 0003-0004) by the in-memory offloading approach, the identification of off-loadable instructions based on algorithm traversing a compiler time data graph and ranking (Fig. 1-2; para 0007) stored content of vertices to determine the nodes in terms of weight for offloading, where each offloaded atomic operation can involve a start address for read and a final address for write (para 0003)
Hence, based on the teaching by Lin, consideration of what is being finally collected into a
node without consideration of intermediate operation leading to the final stored value in the node entails determining by the compiler to the effect of linking sequences of more than two atomic operations sequence into a compacted representation of successive reads so to form merged chain of
atomic operations considered best candidate for offloading (without complying to a cache write-back policy) in which result from intermediate steps (i.e. edge computation) such as non-final values can be slighted (taken off) by effect of compaction by a compiler algorithm traversing a data flow graph and weighting the final content in its vertices. That is, generating a linked read dependence graph in terms of one or more linked chains of dependent elements by linking the one or more first chains ( first sequence of atomic read) with the one or more second chains (second sequence of atomic read) based on the portion of the source code represented on the DAG is recognized.
As for (ii),
Chang discloses a high-bandwidth memory (HBM) system operating under a FIM controller mode via Function-In-HBM control logic to coordinate which atomic operations (para 0007-0008) can be included for transferring from a (GPU) application to the HBM/FIM environment (Fig. 2)- e.g. using a scratchpad sector local at each RAM of the HBM environment to store result of a load/store instruction (para 0029-0030) of a ALU operation; where, according to which, a GPU control associated with the transfer effectuates analysis over compiler originated commands from the GPU source to determine source and destination memory addresses of the HBM RAM as to whether a GPU candidate instruction (para 0034) can be suited as FIM (Function-in-HBM) instruction or else, as non-FIM instruction.
Hence, effect of a controller associated with offloading atomic operations for execution via fast-memory environment to determine whether an operation can be FIM appropriate or denied as non-FIM ready via matching range of memory addresses from the HBM with scope of a candidate atomic operation is recognized.
Further, Alsop discloses offloading operations from a host processor onto a fast, high bandwidth memory or PIM execution environment (pg. 3) where instructions destined for PIM units must be mapped to the actual hardware or fall into the same memory partition of the PIM environment (pg .6) in terms of address-to-physical memory, or else addressing error can occur (pg. 3), the consecutive instructions from the processor cores should be each time mapped to the same offload target device (PIM module) as part of the divergence detection associated with declaration of unload operations destined for the offloading (pg. 7-8)
Therefore, based on marking in Kalamatianos in accordance with identifying a start and an end of a PIM command (para 0036) to be spanned with the maximum capacity range of registers, it would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to implement determination of code portion as candidate for their PIM offloading in Kalamatianos so that operations associated with the compiler marking includes
(1) generating a linked read dependence graph representing one or more linked chains of
dependent elements by linking the one or more first chains with the one or more second chains based
on the portion of the source code - as in Bertacco - where each generated first chain and second
chain represents a first read dependence graph composed of one or more first chains of dependent
elements of the first data structure based on the portion of the source code; and second read dependence graph composed of one or more second chains of dependent elements of the
second data structure based on the portion of the source code, each chain representing a atomic
operation chaining as set forth in Bertacco, each with start address for read and a final address for
write where disruption thereof by a write-back for cache coherence operation is being bypassed;
such that
(2) the ensuring effect of marking is based on a memory capacity of a first number of banks communicatively coupled to respective ones of the one or more processing-in-memory units being greater than or equal to an amount of the memory- as shown in Chang and Alsop - to store a second
number of elements represented by a longest chain of the one or more linked chains being realized
from the linking of atomic operations in Bertacco compiler approach that averts cache write-back
traffic and Lin map reduction that remove intermediate results from a data graph; because
use of a data graph to afford a compiler to mark simple atomic operations typical to a PIM execution offloading mode in which these simple operations can be achieved as locally stored data without necessitating of host CPU to consolidate a final result and/or coordinate the result with other cores established upon a cache coherence policy as set forth above, in that the compiler would be able to determine which stretch of atomic operation (read and write) can be linked into a longest contiguous sequence in accordance to which, only the final content resulting from individual internal
stage operation (like a load, read, store) is read into a next stage of the sequence without any latency caused by additional traffic awaiting intervention outside of the local context of the PIM memory would have for effect to avoid cache pollution, averting complexity/latency in relying upon a CPU layer to synchronize accesses by other runtime entities contending for a value that cannot be timely
established as final and proper read, thus boosting higher throughput for a type of data read using
spatial locality of the PIM memory banks, notably when scope of address range of the instruction to
handle under a PIM environment (e.g. a longer stretch of atomic operations) is being pre-mapped at
compiler time to the actual hardware/physical capacity of memory provision as set forth above in
Alsop and Chang in terms of matching of a desired offload code or contiguous instructions range with at least a larger or equal memory size by each bank in the PIM side, rendering throughput of the
offloaded execution largely enhanced with minimized delay caused by traffic and interaction with the host core.
As per claim 8, Kalamatianos does not explicitly disclose device of claim 4, wherein the marking is based on a first number of rows in a second number of banks communicatively coupled to respective ones of the one or more processing-in-memory units being greater than or equal to a maximum number of interacting elements of the linked read dependence graph.
But marking of offload code in consideration of what code size requires for proper local realization of result to be obtained within a local bank location under a PIM execution mode entails consideration of spatial and locality code marking by the compiler in regard to correlating a number of rows in a corresponding number of banks in relation to their being communicatively coupled to respective ones of the one or more processing-in-memory units, the latter enlisted for executing elements of the linked read dependence graph set forth per obviousness rationale in claim 4; hence the marking of rows of respective memory banks associated with a corresponding PIM unit as capacity deemed proper to carry out a PIM execution of a range of contiguous, acyclic operations in the context of observing spatial mapping to the PIM hardware and locality between processing unit and its memory would have been recognized for the same reasons set forth in the obviousness rationale of claims 4-6.
Thus, basis for a marking to the effect of identifying a first number of rows in a second number of banks that are communicatively coupled to respective ones of the one or more processing- in-memory units, the number of rows being greater than or equal to a maximum number of interacting elements of the linked read dependence graph would have been obvious for the same reasons set forth with rationale address claim 5 as set forth using obviousness of the linked read dependence graph of claim 4.
As per claim 9, Kalamatianos does not explicitly disclose device of claim 8, wherein the maximum number of interacting elements includes an element of the linked read dependence graph and one or more elements directly connected to the element in the linked read dependence graph, the element having a highest number of elements directly connected thereto in the linked read dependence graph.
However, joining elements of a dependency context from traversing a PIM graph per effect of generating chain of read instructions either as first graph of dependent elements and second graph of dependent elements so to have them merged into a highest number of elements connected under a combined read dependence graph to be marked for submission into a offloaded PIM execution has been addressed as obvious using the teachings by Bertacco and Map reduction of data graph in Lin, in accordance to obviousness of claim 4 from above.
Thus, generating a linked/combined read dependence graph per a compiler marking associated with traversing a data graph so that identification of maximum number of interacting elements based thereon takes into account of an element of the linked read dependence graph in conjunction with one or more elements directly connected to the element in the linked read dependence graph, for the element to have highest number of elements directly connected thereto in the linked read dependence graph would have been obvious for the same reasons set forth with the rejection of claim 4 in view of the intended spatial mapping raised as obvious in claim 5.
As per claim 10, Kalamatianos discloses device of claim 1, the operations further including:
during the compiling (see Note2), marking an additional portion of the source code as not
suitable (see below) for execution using the one or more processing-in-memory units (indices
required to offload the operations - para 0065); and
executing the executable code compiled from the additional portion of the source code based on the portion of the source code being marked as not suitable (operations that would otherwise be executed by a computer system’s primary processor core – para 0002 – e.g. processor 132- para 0024- Note2: all original instructions handled by the host controller- processor 132, Fig. 1 – but not designated for offloading or sent as indexed instructions to a command buffer – Fig. 6 - of the memory device 180 – Fig. 1-2, 3A – directed to the PIM execution unit reads on portion of code indicated at compilation- compiler based on static analysis – para 0064 - as not suitable for PIM execution and which is normally executed within the core host CPU; i.e. processor 132) for execution using the one or more processing-in-memory units.
As per claim 11, Kalamatianos discloses a method, comprising:
compiling a portion of source code of a software program to executable code (refer to claim 1);
during the compiling, generating a read dependence graph representing one or more chains of dependent elements of one or more data structures accessed by the portion of the source code; and
offloading the executable code compiled from the portion of the source code for execution by one or more processing-in-memory units based on a memory capacity of a number of banks communicatively coupled to respective ones of the one or more processing-in-memory units being greater than or equal to an amount of memory to store a number of elements represented by a longest chain of the one or more chains.
(refer to claims 4 and 5)
As per claim 12, Kalamatianos discloses wherein the portion of the source code accesses a first data structure and a second data structure, wherein generating the read dependence graph includes:
generating a first read dependence graph representing one or more first chains of dependent elements of the first data structure based on the portion of the source code;
generating a second read dependence graph representing one or more second
chains of dependent elements of the second data structure based on the portion of
the source code; and
generating the read dependence graph representing the one or more chains of
dependent elements by linking the one or more first chains with the one or more
second chains based on the portion of the source code.
(Refer to rationale of claim 4)
As per claim 17, Kalamatianos discloses method of claim 11, further comprising verifying
that an additional number of rows in the number of banks is greater than or equal to a maximum
number of elements that are operated on together in a single computation of the portion of the source
code based on the read dependence graph, wherein offloading the portion of the source code is
further based on the verifying.
(Refer to rationale of claim 5 and claim 8)
Allowable Subject Matter
Claims 6-7, 13-14, 18 are objected to as being dependent upon a rejected base claim, but would be allowable (pending resolution of any pending rejection to the base claims) if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
(claims 6-7) device of claim 4, wherein the marking is based on a duplication metric falling below a threshold, the duplication metric capturing an amount of data duplication in the memory to execute the portion of the source code using the one or more processing-in-memory units;
the operations further including computing the duplication metric based on a comparison of a first number of elements, including duplicated elements, represented by the linked read dependence graph to a second number of unique elements represented by the linked read dependence graph
(claims 13-14) method of claim 11, further comprising computing a duplication metric capturing an amount of data duplication in the memory to execute the portion of the source code using the one or more processing-in-memory units, wherein offloading the portion of the source code is further based on the duplication metric falling below a threshold; wherein
the duplication metric is based on a comparison of a first number of elements, including duplicated elements, represented by the read dependence graph to a second number of unique elements represented by the read dependence graph.
(claim 18) a system, comprising a memory that includes one or more processing-in-memory units; and a processor core to perform operations including:
compiling a portion of source code of a software program to executable code;
during the compiling, computing a duplication metric capturing an amount of data duplication in the memory to execute the portion of the source code using the one or more processing-in-memory units; and
offloading the executable code compiled from the portion of the source code for execution by the one or more processing-in-memory units based on the duplication metric falling below a threshold.
Response to Arguments
Applicant's arguments filed 7/27/26 have been fully considered but they are not persuasive. Following are the Examiner’s observations in regard thereto.
(A) The Applicant has submitted that compiling source code into executable and offloading the code into processing-in-memory units as claimed cannot be practically done by a human due to complexity, scale and deterministic processing of both operations (Applicant's Remarks pg. 10). The recited compiler operation does not provide details about how this improves upon the determination for offloading decision, but rather presents this offloading as a mere nomination of a result without depicting the “how”; therefore amounts to a mere context with no significant algorithmic or hardware impact to the offloading decision. The mention of processing-in-memory is mere reminder that some processing units are available for use by the off-loading determination process, when in fact this determination based on marking falls into well-understood activities that can be performed by a human.
(B) The Applicant has submitted that the claim provides an improvement to the functioning of computers notably in the functioning of PIM-based system, notably in view of the Specifications (para 0005, 0039, 0011, 0044, 0045, 0049) that automates identification of portions of code to offload during compilation, thereby improving accuracy and efficiency of the PIM system (Applicant's Remarks pg. 11-12). The PIM system is well-understood as capability that can perform off-chip on operations that can bypass traffic and latency complexities with the host CPU. The determination that leads to offloading of those operations in the absence of details showing “how” this specifically improves the low level of PIM in performance can be viewed as a mere determination which can be a mental process. Use of metric or comparative amount in form of “greater than” (claim 11) or “duplication metric” (claim 18) without mention of hardware cannot add significantly more to the abstract nature of these concepts. The mention that the “offloading” and generating of “executable” as recited “improves” upon this PIM hardware methodology appears to a) merely recite/rehash on a well-known concept/use of PIM and b) rely on the Specifications to advertise on a desired result as opposed to showing where in the claim is there detailed description of a technical improvement above the broadly mention of a well-understood capability.
( C) The Applicant has submitted that claim 1 as amended to include part of claim 2 differ from the GraphPIM approach by Bertacco which emphasizes on cache coherence, not on loop-carried true dependencies as claimed as read-after-write (RAW) situations. The underlying reason as to why offloading is being diverting from the standard Von Neuman methodology is driven by latency cost, and complexities that result from frequent traffic with the core CPU; and one of the most costly scenario include the RAW instances because of which an atomic operation that typically concludes using its own immediate value and operand incurs dependency setback caused by change to the value just registered by a write. The newly added limitation about detecting of all cyclic dependency in form of RAW has been addressed in the rationale A of claim 1.
( D) The Applicant has submitted that in the rejection of claim 11, Chang does not disclose read dependence graph and chaining of a any longest chain thereof; nor is there offloading based on “a number of banks communicatively coupled to …processing-in-memory units … represented by a longest chain” (Applicant's Remarks pg. 17). The mapping by Chang is to show effect of a controller associated with offloading atomic operations for execution via fast-memory environment to determine whether an operation can be FIM appropriate or else, denied as non-FIM ready via matching range of memory addresses from the HBM with scope of a candidate atomic operation; and matching here include fitting the number of offloaded instructions as the longest chain with the amount of memory banks having equal to or greater capacity to match that chain.
(E ) The Applicant has submitted that Alsop as cited, discloses providing convergence between addresses of a PIM operation with those of the PIM modules; and this is far different from the language of claim 11 regarding offloading based on “a number of banks communicatively coupled to …processing-in-memory units … represented by a longest chain” (Applicant's Remarks pg. 17-18). Broad interpretation of claim 11 has this limitation construed as a process that maps as much as possible the chain of PIM-intended operations with a minimal or optimal number of memory addresses in the PIM memory. This desired correspondence is shown in Alsop per effect of address-to-physical memory correspondence, or else will have to address a mismatch error, the constraint thereof being that consecutive instructions from the processor cores should be each time mapped to the same offload target device (a off-chip memory module) as part of the divergence/convergence detection associated in Alsop declaration of unload operations destined for the offloading.
Alsop and Chang are deemed sufficient to support the obviousness type rationale as presented.
(D ) The Applicant has submitted that Bertacco does not perform dependency in accordance
Read-after-Write; but rather describes high-degree vertices for on-chip handling and low-degree vertices for off-chip execution. (Applicant's Remarks pg. 19). The language of claim 3 falls under the same mechanism of identifying those graph operations affected by a cyclic dependency so to mark them off the offloading mechanism. The rejection of claim 3 has been currently integrated with the rejection of the now-amended claim 1 and any patentability merits thereof is deemed MOOT.
In all, the claims as submitted stand rejected as set forth above.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Tuan A Vu whose telephone number is (571) 272-3735. The examiner can normally be reached on 8AM-4:30PM/Mon-Fri.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Chat Do can be reached on (571)272-3721.
The fax phone number for the organization where this application or proceeding is assigned is (571) 273-3735 ( for non-official correspondence - please consult Examiner before using) or 571-273-8300 ( for official correspondence) or redirected to customer service at 571-272-3609.
Any inquiry of a general nature or relating to the status of this application should be directed to the TC 2100 Group receptionist: 571-272-2100.
/Tuan A Vu/
Primary Examiner, Art Unit 2193
Septembre 08, 2026