DETAILED ACTION
Claims 1, 10, 13 are amended. Claims 5, 14 were previously canceled.
Claims 1-4, 6-13, 15-20 are pending.
Priority: August 08, 2023
Assignee: Samsung
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 1-4, 6-13, 15-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Note: In the Remarks, the Applicant does not mention the relevant specification paragraph(s) that recite the amendment(s).
1.Amended claims 1, 13 are rejected for reciting limitations that are unclear, ambiguous and indefinite.
Amended Claim 1 recites, ‘wherein the probe….to inquire whether an address….stored in the coherence directory is in the cache of any of the….computing elements’, ….’wherein the coherence….controller is configured ….to clean the coherence directory to remove the address….in response to an acknowledgement indicating that the address is not in the cache of any of the plurality of computing elements’.
The limitation is contradictory. By definition, a probe to ‘any’ of the caches inherently requires the directory controller to collect and evaluate all acknowledgments from all queried computing elements/cores etc., before proceeding. Relying on only one acknowledgment leaves the state of the remaining cores unresolved.
The limitation recites that the directory controller sends a probe to inquire whether an address is 'in any of the caches' of the cores, and deletes the address in response to 'an' (a single) acknowledgement that the address is not in the cache of 'any' of the cores. This limitation is indefinite and logically contradictory. Because the directory controller must ascertain whether the address is in any of the caches, relying on only a single acknowledgment leaves the directory controller unable to determine the status of the remaining computing elements, rendering the scope of the deleting step ambiguous.
Accordingly the limitation and claim 1 are rejected as being indefinite because they fail to set out the boundaries and operation of the claimed system node. Claim 13 has the same issue and is also rejected.
2.Amended Claims 1, 13 are rejected for reciting a limitation that is unclear, inconsistent and indefinite.
Amended Claim 1 recites, ‘a coherence directory controller…to send a probe….during periods of computation of the plurality of computing elements’.
The spec does not recite this limitation. The spec does not define ‘periods of computation’ or ‘periods of computation of….elements’.
In Para-0010, the spec recites, ‘….send the probe when the computing elements are busy’.
The spec explicitly recites sending the probe when the cores are ‘busy’, whereas the claim's recitation of ‘periods of computation’ is unquantifiable and broad. ‘Periods of computation’ does not inherently mean busy (e.g., a core could be computing in a halted or idle state, or performing operations not relevant to the coherence directory).
Because the term ‘periods of computation’ can be interpreted in multiple ways, the limitation is indefinite. Hence claim 1 is rejected. Claim 13 has the same issue. For examination the spec is used.
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Note: In the Remarks, the Applicant does not mention the relevant specification paragraph(s) that recite the amendment(s).
Claim(s) 1-4, 6-13, 15-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
1.Amended claims 1, 13 are rejected for lack of written description support, as it finds no basis in or support from the spec as filed.
Amended Claim 1 recites, ‘a coherence directory controller configured to send a probe to the….computing elements during a free cycle of the network-on-chip and during periods of computation of the….computing elements’.
The spec does not teach this limitation.
Para-0006, Para-0043, Fig. 2, step 210 of the spec recite, ‘a coherence ….controller configured to send a probe to the computing elements during a free cycle of the network-on-chip’. And Para-0010 recites, ‘the coherence….controller may be configured to send the probe when the computing elements are busy’.
Because the spec only teaches sending probes when the cores are busy (NOC idle/free cycle), adding a limitation that covers sending probes ‘during periods of computation of the computing elements’ suggests the cores are actively computing while the NOC is busy. Accordingly the limitation expands the scope beyond what was actually taught and introduces new matter.
With reference to the disclosed cache coherent architecture which includes a directory controller, a main memory, a plurality of cores and/or accelerators connected to a NOC and an interconnect, the spec does not define what constitutes a ‘NoC free cycle’. Neither does the spec define
The spec fails to teach in detail how the coherent controller determines a NoC free cycle and determines periods of computation of the cores, simultaneously. The spec also fails to teach in detail how the coherent controller determines a NOC free cycle or how the coherent controller determines periods of computation of the cores.
The spec fails to describe any hardware structure, software logic, or a specific algorithm that correlates to the claimed limitation. This lack of disclosure demonstrates that the inventor was not in possession of the claimed limitation, especially sending the snoop instruction ‘during a NOC free cycle’ and ‘during periods of computation of cores or accelerators’, at the time of filing.
The spec only uses functional phrases (e.g., ‘send a probe’, ‘during a free cycle of NoC’, ‘when the cores are busy’), but fails to convey with adequate details that the inventor had actually conceived of how the coherence directory controller performs these functions, at the time of filing.
Since the limitation overreaches the scope of the inventor's actual contribution to the disclosure, the limitation is unsupported by the spec and is rejected for lack of possession. Claim 13 has the same issue and is also rejected. For examination, the spec is used.
Note: Based on the amendments and arguments, the rejection has been clarified and maintained.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6, 11, 13, 15, 19 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Salisbury et al (20160350220) in view of Hagersten et al (20200364144), Ringe et al (20180217932) and Morton et al (20070055826).
As per Claim 1, Salisbury discloses a cache-coherent computer system node (Salisbury, [0064,0066 - Fig. 1 shows a data processing apparatus with data handling nodes 10, 12, 14, 16, 18, 20 and interconnect circuitry 30 connected to the nodes. The apparatus/C-C system node can be implemented as a single integrated chip]) comprising:
a network-on-chip (Salisbury, [0052 - A cache coherent data communication device, such as a network-on-chip/NoC or a cache coherent interconnect comprising a cache coherency controller]);
a plurality of computing elements connected in a fully cache-coherent manner in communication with the network-on-chip (Salisbury, [0067 - Fig. 1 is a memory system comprising: a cache coherent data communication device/NoC and a group of one or more cache memories each connected to the cache coherent data communication device]; [0068 – In Fig. 1, nodes such as 12, 14 are processing elements/CPUs having their own respective caches]);
a coherence directory comprising a plurality of addresses for cache of each of the plurality of computing elements (Salisbury, [0006 - A directory for memory addresses cached by a group of cache memories connectable in a coherent cache structure, wherein each of the cache memories are caching those memory addresses]; [0003 - An m-way associative structure is provided for such a directory]),
a plurality of coherence states of the plurality of addresses (Salisbury, [0008 - The control circuitry being responsive to status data indicating whether each cache memory in the group is currently subject to cache coherency control so as to take into account, in the detection of the directory entry relating to the memory address to be accessed, only those cache memories in the group which are currently subject to cache coherency control]; [0122 – In Fig. 8, step 810, snoop filter 300 creates a mask which is a representation of the status/state of each cache memory/tag-address detected at step 800]),
and a plurality of tracking vectors (Salisbury, [0091 – In Fig. 5, the snoop directory entry 500 comprises two portions: a tag 510/address and a snoop vector 520]; [0089 - Fig. 4 shows a snoop filter directory 310 organized as an m-way associative group of n×m entries]; [0041 - The snoop vector returned from the snoop filter for the original access indicates caching agents still enabled in the coherency domain]) of the plurality of addresses (Salisbury, [0090 - The directory is m-way associative so that multiple memory addresses map to an associative set of m directory entries]; [0102 - The snoop vectors show each directory entry comprising information derived from the respective cached memory address and information indicating, for each cache memory in the group of cache memories, whether that cache memory is currently caching that memory address]);
a coherence directory controller (Salisbury, [0071 – Fig. 3 shows the operation of a cache coherency controller including a snoop filter]) configured to send a probe (Salisbury, [0080 - A snoop operation consists of sending a message to a cache memory indicated by the directory, to be caching a memory address being accessed by another cache memory and receiving a response indicating if the cache memory is actually caching that memory address]) to the plurality of computing elements during a free cycle of the network-on-chip (Salisbury, [0074 – Fig. 3 shows transaction router 320 in communication with the cache memories, thereby implying monitoring of directed coherence traffic; It is well-known that the transaction router monitors network buffers and tracks pipeline activity. Because the NoC is a shared medium, less coherence traffic tells the controller that the network links and routers are currently free of congestion. Thereby the controller determines that NoC is idle; Since neither the claim nor the spec disclose determining the NoC idle state, the citation is a valid interpretation]) and during periods of computation ([See 112(a)]) of the plurality of computing elements (Salisbury, [See Fig. 1]; [Figs. 6-12 are flowcharts initiated by the controller, thereby implying that the cores are not generating lots of coherence memory requests; So the controller determines that cores are locally ‘busy’ processing their own instructions. Since neither the claim nor the spec disclose determining that the cores are busy or periods of computation of cores, the citation is a valid interpretation]; [0082 - If a snoop operation is needed to enquire the current status of the data at one or more caches, then the snoop filter 300 carries out that enquiry as a unicast or multicast communication; The snoop operation does not involve data transfer. Since the interconnect is idle and cores are ready to receive the new message/snoop, it implies that the controller sends the snoop when the NoC is idle and the cores are busy]),
wherein the probe is configured to inquire whether an address of the plurality of addresses stored in the coherence directory (Salisbury, [0080 - A snoop operation/probe provides an example of sending a message to a cache memory]) is in the cache of any of the plurality of computing elements (Salisbury, [0002 – The cache coherency controller oversees accesses to memory addresses and uses a snoop filter/directory for checking whether a cached version of a memory address to be accessed is held by another cache in the cache coherent system]);
wherein the coherence directory controller is configured to clean the coherence directory to remove the address from the coherence directory (Salisbury, [0043 - Cleaning the directory of entries relating to cache memories no longer under cache coherency control]; [0104 – In Fig. 6, step 630 the snoop filter deletes the directory entry corresponding to that cache line at that cache, and sets the relevant bit of the snoop vector of that cache line to zero]; [0107 - Fig. 7 shows an operation to release an entry and store a newly created entry in snoop directory; Since the claim does not recite the reason for the deletion from the directory, the citations are valid interpretation]) in response to an acknowledgement ([See 112(b)]) indicating that the address is not in the cache of any of the plurality of computing elements (Salisbury, [0104 - In Fig. 6, step 620, the cache agent/core notifies the cache coherency controller 302 of the eviction, thereby implying the acknowledgement indicating that the address is not in the cache]).
Hagersten clarifies the architecture, NoC free cycle/idle state and busy state of computing elements as follows,
a network-on-chip (Hagersten, [Fig. 6: NoC 650]);
a plurality of computing elements connected in a fully cache-coherent manner in communication with the network-on-chip (Hagersten, [0057 – In Fig. 6, the nodes are connected together with each other through network on chip/NoC 650 circuit. NoC 650 also connects the nodes to the directory/DIR 660/controller, global LLC 670 and memory 660. DIR plays a central role in the coherence protocol that keep the contents of the caches and the CLBs coherent and consistent, thereby implying that the nodes/elements are connected in a fully cache-coherent way with the NoC]);
a coherence directory controller (Hagersten, [0057 – In Fig. 6, DIR 660/CDC plays a central role in the coherence protocol that keep the contents of the caches and the CLBs coherent and consistent]) configured to send a probe (Hagersten, [0023 - A coherence message/probe which is sent activates the blocking function to block other coherence messages if the other coherence messages are for the same address region as the coherence message]) to the plurality of computing elements (Hagersten, [0061 – Fig. 6 shows one CPU in each node, and each node may contain any number of CPUs, GPUs, accelerators, etc.]) during a free cycle of the network-on-chip (Hagersten, [0171 - Fig. 15: step 1506, block some coherence messages from being sent on the network, thereby implying NoC is idle because there is less network traffic; Since neither the claim nor the spec disclose determining the NoC idle state, the citation is a valid interpretation]) and during periods of computation ([See 112(a)]) of the plurality of computing elements (Hagersten, [0171 – In Fig. 15, step 1504, coherence of values stored in the caches is maintained by a distributed cache coherence protocol which sends coherence messages on the network]; [0073 - Evictions local to a node, such as eviction from L-1 620 to L-Y 640 are handled locally, tracked by its local CLB entries and are not visible outside the node, thereby implying that the computing elements are busy; Since neither the claim nor the spec disclose determining that the cores are busy or periods of computation of cores, the citation is a valid interpretation]),
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the blocking function of Hagersten into the cache coherence of Salisbury for the benefit of blocking some
coherence messages from being sent on the network. A coherence message which is sent activates the blocking function to block other coherence messages if the other coherence messages are for the same address region as the coherence message thereby reducing the number of NoC transactions (Hagersten, 0023).
Ringe discloses a SoC optimized for minimum latency or optimal bus utilization to show that the NoC is idle and cores are busy as follows,
a coherence directory controller (Ringe, [Fig. 1: Home node/HNF 112 includes snoop filter 138 and alignment logic 148]) configured to send a probe (Ringe, [0055 – In Fig. 3, If the home node determines that node RNF-128 has a copy of the requested data, the ADDR is passed as a snoop request/SNP_REQ [01] 304 towards snoop target RNF-128 without any changes]) to the plurality of computing elements during a free cycle of the network-on-chip (Ringe, [0043 – In Fig. 1, interconnect system 11 connects devices having different bus widths by the inclusion of data combiner module 118 and data splitter module 120]; [0033 - In an on-chip network, a ‘packet’ is the meaningful unit of the upper-layer protocol, such as the cache-coherence protocol]) and during periods of computation of the plurality of computing elements (Ringe, [0030 - A device such as 102,104,106,108, may be a core, a cluster of processing cores, and accelerators like GPP, DSP, FPGA or an ASIC device]; [0072 – In Fig. 7, if the system is optimized for minimum latency, as depicted by step 720, the snoop address is sent, at step 714, to the snoop target/cache without modification; The snoop being sent when the system is optimized for minimum latency implies that the NoC is idle and cores are busy. The system achieves optimal low/min latency because the snoop traveled from controller to cache with zero queuing delays]; [0044 - Data combiner module 118 combines multiple incoming flits containing narrow beats, say 128-bit, into a single flit containing a wider beat, say 256-bit, provided that all narrow beats are contiguous in time, do not cross the wider beat boundary, and arrive in a given time window]; [0062 – Snoop address alignment]);
It is well-known that data splitters and data combiners/multiplexers and demultiplexers are used in a NoC router to distribute packets into specific queues. The NoC router uses a split-and-merge stage to route packets, wherein an incoming data packet is evaluated, split into individual flits, sent across the chip, and safely combined at the destination.
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the data splitter and data combiner of Ringe into the cache coherence of Salisbury, Hagersten for the benefit of achieving an efficient trade-off between signal latency and data bus utilization in a data processing apparatus, such as a System-on-a-Chip (Ringe, 0022).
Morton clarifies,
wherein the probe is configured to inquire whether an address of the plurality of addresses (Morton, [0011 - The probe is evaluated with the probe filtering unit/PFU/coherence controller to determine whether a valid copy of the memory line/address is in any of the cache memories. The probe is transmitted from the PFU only to selected ones of the processing nodes identified by the evaluating, thereby implying that the probe is sent to any of the cores]) stored in the coherence directory is in the cache of any of the plurality of computing elements (Morton, [0082 – In Fig. 7, coherence directory 701 includes state information 713, dirty data owner information 715, and an occupancy vector 717 associated with the memory lines 711/addresses. The memory line states are modified, owned, shared, and invalid/unused]),
wherein the coherence directory controller (Morton, [cache coherence controller+home memory controller/HMC]; [0107 - The coherence directory is an associative memory which associates the memory line addresses with their remote cache locations]) is configured to clean the coherence directory to remove the address from the coherence directory (Morton, [0112 - If a directory entry to be purged indicates that the cached line is in the clean state, then the mechanism invalidates the memory line in each of the remote caches in which the line is cached]; [0129 - When the cache coherence directory associated with the cache coherence controller in a cluster determines that it needs to evict an entry which corresponds to one or more remotely cached clean memory lines, it generates a validate block request for the HMC. The HMC then generates invalidating probes to all the local nodes in the cluster. The local nodes invalidate their copies of the memory line and send confirming responses/ack messages to the HMC indicating that the invalidation took place]; [0131 - The cache coherence directory then transmits a ‘source done’ to the MC/HMC in response to which the memory line is freed up for subsequent transactions, thereby implying the cleanup of the memory line/address]; [0144 - In Fig. 17, the eviction manager 1702 is part of the cache coherence directory 1701 which is a functional block within the cache coherence controller 1700]) in response to an ([See 112(b)]) acknowledgement indicating that the address is not in the cache of any of the plurality of computing elements (Morton, [0124-0126 - The eviction of a cache coherence directory entry of a clean line in a cache requires that the cache invalidate its copy. The transaction involves: 1. The copy of the line in the cache is invalidated via the probe, and the eviction is notified/ack message when the invalidation is complete from cache]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the cache coherence controller/PFU of Morton into the cache coherence of Salisbury, Hagersten, Ringe for the benefit of using a coherence directory where global memory line state information is maintained and accessed by a memory controller or a cache coherence controller in a particular cluster. The coherence directory tracks and manages the distribution of probes as well as the receipt of responses (Morton, Para-0080).
As per Claim 6, the rejection of claim 1 is incorporated, and Salisbury discloses,
wherein the coherence directory controller (Salisbury, [0009,0010 - A cache coherency controller/CCC stores a directory 310 indicating memory addresses cached by a group of cache memories connectable in a coherent cache structure]) is configured to send the probe not in response to a command from the plurality of computing elements (Salisbury, [0082 - When a potential snoop operation is initiated, snoop filter 300 consults directory 310 to detect whether the information in question is held in one or more of the caches. If a snoop operation is indeed needed to enquire as to the current status of the data at one or more caches, then the snoop filter 300 can carry out that enquiry as a unicast or multicast communication, thereby implying that the CCC sends the probe not in response to a command from the plurality of computing elements]).
As per Claim 11, the rejection of claim 1 is incorporated, and Salisbury, discloses,
an interconnect (Salisbury, [Fig. 1: interconnect circuitry 30]),
wherein the coherence directory controller (Salisbury, [0006 - A cache coherency controller comprising a directory for memory addresses cached by a group of cache memories connectable in a coherent cache structure]) is configured to send the probe (Salisbury, [0080 - A snoop operation sends a message to a cache memory which is indicated, by the directory, to be caching a memory address being accessed by another cache memory]) over the interconnect (Salisbury, [0064 – In Fig. 1, the interconnect circuitry comprises a plurality of interfaces 40, 42, 44, 46, 48, 50 each associated with a respective one of the data handling nodes, and data routing circuitry 60 for controlling and monitoring data handling transactions as between the various data handling nodes]).
As per Claim 13, it is similar to claims 1-4 and therefore the same rejections are incorporated.
As per Claim 15, it is similar to claim 6 and therefore the same rejections are incorporated.
As per Claim 19, it is similar to claim 11 and therefore the same rejections are incorporated.
Claims 2-4, 7, 9, 12, 16, 18, 20 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Salisbury et al (20160350220) in view of Hagersten et al (20200364144), Ringe et al (20180217932), Morton et al (20070055826) and Daya et al (‘SCORPIO: A 36-core research chip demonstrating snoopy coherence on a scalable mesh NoC with in-network ordering’, IEEE, 2014, Pgs. 1-13) and
As per Claim 2, the rejection of claim 1 is incorporated, and Salisbury, Hagersten, Ringe, Morton disclose a cache-coherent system.
Daya further discloses,
wherein the plurality of computing elements is homogenous and comprises a plurality of cores (Daya, [Pg. 7, Col. 1, Sec. 4 - In Fig. 5, the 36-core fabricated multicore processor is arranged in a grid of 6×6 tiles]; [Pg. 2, Col. 1, Abstract,Para-3 - Scorpio architecture comprises 36 Freescale e200 Power Architecture cores]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the scalable mesh NoC of Daya into the cache coherence of Salisbury, Hagersten, Ringe, Morton for the benefit of using Scorpio, an ordered mesh NoC architecture with a separate fixed-latency, buffer-less network to achieve distributed global ordering. Message delivery is decoupled from the ordering, allowing messages to arrive in any order and at any time, and still be correctly ordered (Daya, Abstract).
As per Claim 3, the rejection of claim 1 is incorporated, and Salisbury, Hagersten, Ringe, Morton disclose a cache-coherent system.
Daya further discloses,
wherein the plurality of computing elements is homogenous and comprises a plurality of accelerators (Daya, [Pg. 7, Col. 2, Table 1 - FPGA controller 1× Packet-switched flexible data-rate controller]; [Pg. 5, Col. 1, Para-2 - Allow requests to fork through multiple router output ports in the same cycle, thus providing efficient hardware broadcast support]; [Pg. 7, Col. 2, Para-4.1, Para-2 - Snooping hardware is present at both L1 and L2 caches]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the scalable mesh NoC of Daya into the cache coherence of Salisbury, Hagersten, Ringe, Morton for the benefit of using Scorpio, an ordered mesh NoC architecture with a separate fixed-latency, buffer-less network to achieve distributed global ordering. Message delivery is decoupled from the ordering, allowing messages to arrive in any order and at any time, and still be correctly ordered (Daya, Abstract).
As per Claim 4, the rejection of claim 1 is incorporated, and Salisbury, Hagersten, Ringe, Morton disclose a cache-coherent system.
Daya further discloses,
wherein the plurality of computing elements is heterogeneous (Daya, [Pg. 3, Col. 2, Sec. 3, Para-2 - The broadcast coherence requests from different source nodes may arrive at the network interface controllers/NIC of each node in any order]) and comprises a combination of a plurality of cores (Daya, [Pg. 7, Col. 1, Sec. 4 - In Fig. 5, the 36-core fabricated multicore processor is arranged in a grid of 6×6 tiles]; [Pg. 2, Col. 1, Abstract,Para-3 - Scorpio architecture comprises 36 Freescale e200 Power Architecture cores]) and a plurality of accelerators (Daya, [Pg. 7, Col. 2, Table 1 - FPGA controller/accelerator 1× Packet-switched flexible data-rate controller]; [Pg. 5, Col. 1, Para-2 - Allow requests to fork through multiple router output ports in the same cycle, thus providing efficient hardware broadcast support]; [Pg. 7, Col. 2, Para-4.1, Para-2 - Snooping hardware/accelerator is present at both L1 and L2 caches]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the scalable mesh NoC of Daya into the cache coherence of Salisbury, Hagersten, Ringe, Morton for the benefit of using Scorpio, an ordered mesh NoC architecture with a separate fixed-latency, buffer-less network to achieve distributed global ordering. Message delivery is decoupled from the ordering, allowing messages to arrive in any order and at any time, and still be correctly ordered (Daya, Abstract).
As per Claim 7, the rejection of claim 1 is incorporated, and Salisbury, Hagersten, Ringe, Morton disclose a cache-coherent system.
Daya further discloses,
wherein the probe is incorporated into a standard coherence protocol (Daya, [Pg. 7, Table 1 - Coherence protocol MOSI]; [Pg. 7, Col. 2, Sec. 4.2 - The standard MOSI protocol is adapted to reduce the writeback frequency and to disallow the blocking of incoming snoop requests/probes]; [Pg. 3, Col. 2, Sec. Notification Network - For every coherence request/snoop/probe sent on the main network, a notification message encoding the source node’s ID/SID is broadcast on the notification network to notify all nodes that a coherence request from this source node is in-flight and needs to be ordered]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the scalable mesh NoC of Daya into the cache coherence of Salisbury, Hagersten, Ringe, Morton for the benefit of using the standard MOSI protocol to reduce the writeback frequency and to disallow the blocking of incoming snoop requests. To achieve this, an additional O_D state instead of a dirty bit per line is added to permit on-chip sharing of dirty data (Daya, Pg. 7, Col. 2, Last Para).
As per Claim 9, the rejection of claim 7 is incorporated, and Salisbury, Hagersten, Ringe, Morton disclose a cache-coherent system.
Daya further discloses,
wherein the coherence directory controller is configured to send the probe utilizing an existing probe/snoop channel (Daya, [Pg. 7, Col. 2, Table-1 - Channel width 137 bits (Ctrl packets – 1 flit, data packets – 3 flits)]; [Pg. 5, Col. 2, Para-1 - To prevent the deadlock scenario, one reserved virtual channel/rVC is added to each router and NIC, reserved for the coherence request/probe with SID equal to ESID of the NIC attached to that router]) of the standard coherence protocol (Daya, [Pg. 7, Col. 2, Sec. 4.2:Coherence Protocol - The standard MOSI protocol is adapted to reduce the writeback frequency and to disallow the blocking of incoming snoop requests]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the scalable mesh NoC of Daya into the cache coherence of Salisbury, Hagersten, Ringe, Morton for the benefit of using the standard MOSI protocol to reduce the writeback frequency and to disallow the blocking of incoming snoop requests. To achieve this, an additional O_D state instead of a dirty bit per line is added to permit on-chip sharing of dirty data (Daya, Pg. 7, Col. 2, Last Para).
As per Claim 12, the rejection of claim 1 is incorporated, and Salisbury, Hagersten, Ringe, Morton disclose a cache-coherent system.
Daya further discloses,
wherein the coherence directory controller is configured to send the probe intermittently (Daya, [Pg. 3, Col. 2, Sec. 3, Para-2:Notification Network - Synchronized time windows are maintained, greater than the latency bound, at each node in the system. The notification messages are synchronized and sent only at the beginning of each time window, thus guaranteeing that all nodes received the same set of notification messages at the end of that time window; Here ‘synchronized time windows’ implies sending the probe/message intermittently. Since the claim does not define ‘send the probe intermittently’, the citation is a valid interpretation]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the scalable mesh NoC of Daya into the cache coherence of Salisbury, Hagersten, Ringe, Morton for the benefit of using Scorpio, an ordered mesh NoC architecture with a separate fixed-latency, buffer-less network to achieve distributed global ordering. Message delivery is decoupled from the ordering, allowing messages to arrive in any order and at any time, and still be correctly ordered (Daya, Abstract).
As per Claim 16, it is similar to claim 7 and therefore the same rejections are incorporated.
As per Claim 18, it is similar to claim 9 and therefore the same rejections are incorporated.
As per Claim 20, it is similar to claim 12 and therefore the same rejections are incorporated.
Claims 8, 10, 17 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Salisbury et al (20160350220), Hagersten et al (20200364144), Ringe et al (20180217932), Morton et al (20070055826), Daya et al (‘SCORPIO: A 36-core research chip demonstrating snoopy coherence on a scalable mesh NoC with in-network ordering’, IEEE, 2014, Pgs. 1-13), and Tune (20140281180).
As per Claim 8, the rejection of claim 1 is incorporated, and Salisbury, Hagersten, Ringe, Morton, Daya disclose the MOSI protocol, an extension of the basic MSI cache coherency protocol.
Tune further discloses the MESI protocol which is also an extension of the basic MSI protocol,
wherein the standard coherence protocol is a MESI protocol (Tune, [0041 - The action of snoop requests in managing data coherence within a system such as in Fig. 1 for coherency control, e.g. MESI, MOESI, ESI, MEI etc., using snoop requests are employed]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the MESI coherency protocol of Tune into the cache coherence of Salisbury, Hagersten, Ringe, Morton, Daya for the benefit of using MESI for the snoop requests to determine which is the most up-to-date copy of the cache line available and return this to the original requesting cache memory. This snoop request may also invalidate some of the existing copies of the cache line as appropriate (Tune, 0041).
As per Claim 10, the rejection of claim 1 is incorporated, and Salisbury, Hagersten, Ringe, Morton disclose a main memory.
Tune clarifies,
main memory connected to the cache of each of the plurality of computing elements (Tune, [0040 – In Fig. 1, snoop control circuitry 10 is connected to the level 2 cache memories 8 and serves to receive memory access requests issued to main memory 12 when a cache miss occurs within one of the level 2 cache memories; Please note: Fig. 1 of the spec shows the main memory connected to the cache(s) via the interconnect]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the main memory of Tune into the cache coherence of Salisbury, Hagersten, Ringe, Morton for the benefit of include a main memory from which the plurality of main cache memories cache data (Tune, 0022).
Daya further clarifies,
main memory connected to the cache of each of the plurality of computing elements (Daya, [Pg. 8, Col. 1, Sec. 4.3 - coherency between L1s, L2s and main memory]; [Pg. 8, Col. 2, Sec. Directory baselines - The ownership bit indicates if the main memory has the ownership; that is, none of the L2 caches own the requested line and the data should be read from main memory]).
Therefore it would have been obvious to a person of ordinary skill at the time of filing to incorporate the main memory of Daya into the cache coherence of Salisbury, Hagersten, Ringe, Morton, Tune for the benefit of reading the data from the main memory if none of the L2 caches own the requested line (Daya, Pg. 8, Col. 2, Para-Last).
As per Claim 17, it is similar to claim 8 and therefore the same rejections are incorporated.
Response to Arguments
The Applicant's arguments filed on June 11, 2026 have been fully considered, but they are not persuasive. The broad amendments stretch the scope beyond what is originally disclosed, resulting in 112(a)’s and 112(b)’s.
Applicant argues:‘Because the specification does not describe it as essential for the coherence directory controller to determine the free cycle of the network-on-chip or to determine the computation cycle of the compute elements (cores and/or accelerators), these features are not essential and are properly omitted from the claims’. (Rem, Pg. 7)
Response: Applicant arguments cannot take the place of evidence or adequate spec support.
The applicant is confusing ‘unclaimed essential matter’ with ‘unsupported claimed matter’. The latter, as shown below, is a direct 112(a) violation.
The claim requires the probe to be sent by the coherence controller during the free cycle of the NoC/idle NoC and during periods of computation of the computing elements. Determining the two conditions via the coherence controller is therefore the necessary pre-requisite for sending the probe.
Without this disclosure, the claim is fatally broad because it covers any and all ways to determine the free cycle and periods of computation, while providing no teaching on how to achieve the result.
The phrase ‘send a probe during….’, is a functional limitation. The spec fails to provide a hardware configuration, or software algorithm to define what constitutes a NoC free cycle and periods of computation of cores. A claim cannot rely on purely functional language while trying to establish non-obviousness when the spec completely lacks written description support.
That said, the spec fails to demonstrate possession of how the coherence controller determines these two conditions. Accordingly the spec does not demonstrate possession of when to send the probe by the coherence controller. Please see the 112(a).
As an aside, the disclosure also fails the enablement prong of 112(a).
Applicant further argues:‘Independent claim 1 has been amended herein to recite…., "a coherence directory controller…to send a probe to the ….during a free cycle of the network-on-chip and during periods of computation of the plurality of computing elements’. (Rem, Pg. 10)
Response: The ‘free cycle of the NoC’ is not even defined in the spec, let alone ‘determining a free cycle of the NoC’ in the context of the disclosed cache coherent architecture. The same is true of ‘periods of computation of the…computing elements’, no definition, no determination disclosed in the spec.
The limitation is based on functional timing but there is no disclosure of the physical or logical means (e.g., clock cycles, timestamps etc.) used to determine the free cycle of the NOC and/or periods of computation of the computing elements.
The inventor relies on overly broad, one-line limitations in the spec to lock in the entire technology of NOC controlled cache coherence management, regardless of whether they possess it.
That being said, the limitation is unsupported by the spec. Please see the 112(a).
Applicant further argues:‘…..configured to clean the coherence directory to remove the address …..in response to an acknowledgement indicating that the address is not in the cache (Rem, Pg. 10)
Response: The limitation is contradictory. Please see the 112(b).
Applicant further argues:‘Claim 13….a probe to cache of each of the plurality of cores and/or….accelerators during a free cycle of the network-on-chip and during periods of computation ….of cores and/or….... and cleaning the address from the coherence directory in response to …..’ (Rem, Pg. 10)
Response: Amended Claim 13 has the same issues as Amended Claim 1. Please see the 112(a) and 112(b)’s.
Applicant further argues:‘FIG. 15….Hagersten also discloses that some coherence messages are blocked from being sent on the network (step 1506) and that sending a coherence message activates the blocking function to block other coherence messages …..(step 1508)’ (Rem, Pg. 11).
Response: Regarding Fig. 15, steps 1506, 1508, the blocking function acts as a traffic controller/router because it ensures that the first request finishes processing and resolves to a stable state before subsequent requests can interact with that memory. Doing so avoids race conditions. By stalling messages directed to the same address region, the coherence protocol enforces strict sequential ordering of operations on that specific block thereby preserving transaction order. In essence, both steps reduce network traffic.
That said, according to the spec, NoC free cycle is equivalent to NoC being idle. Therefore, blocking some coherence messages from being sent on the network, implies there is less network traffic thereby determining that the NoC is idle.
The relationship between the NoC, cores, and probes relies on this traffic observation. Because the NoC is a shared medium, a reduction of coherence traffic (fewer packets) tells the coherence controller that the network links and routers are currently free of congestion, there are lower queueing delays etc. This makes the interconnect idle and highly responsive, meaning a new message/probe can reach its destination/cache immediately.
The coherence controller also recognizes that the cores are locally ‘busy’ processing their own instructions because they are not simultaneously generating lots of coherence memory requests. See Hagersten, Para-0073.
Knowing the NoC is idle and the target core(s) are ready to process new messages without delay, the coherence controller sends the claimed probe. Hence the combination of the prior art disclose the claimed requirement. Please see O/A.
As mentioned above, the spec does not disclose determining any of the conditions to send the probe. The prior art discloses the requirement.
Applicant further argues:‘Thus, neither Salisbury nor Hagersten, whether taken alone or in combination, discloses,…."a coherence….controller….to send a probe to the….computing elements," as recited in amended… claim 1, "sending, from a coherence….controller, a probe to…. plurality of …," as recited in amended claim 13’. (Rem, Pg. 11)
Response: Relying on unsubstantiated amendments to mischaracterize the prior art, invalidates the argument. Please see the 112(a), 112(b)’s and the O/A.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ARVIND TALUKDAR whose telephone number is (303)297-4475. The examiner can normally be reached M-F, 10 am-6pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hosain Alam can be reached at 571-272-3978. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Arvind Talukdar
Primary Examiner
Art Unit 2132
/ARVIND TALUKDAR/ Primary Examiner, Art Unit 2132