DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Applications, No.KR10-2025-0023790 filed on 02/24/2025, and KR10-2025-0041076 filed on 03/31/2025.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 is/are rejected under 35 U.S.C. 103 as being unpatentable over Pitchumani et al. (US2025/0357426), hereinafter Pitchumani in view of Bari et al. (US2024/0004690), hereinafter Bari.
Regarding claim 1, Pitchumani teaches a data processing system comprising:
a host (Pitchumani, [0027], a machine 105 (e.g., a host) includes a processor 110);
a first memory device (Pitchumani, [0027], memory device 140) configured to communicate with the host via an interface (Pitchumani, [0029], a server 132 which includes one or more compute/memory trays 134 having compute and/or memory resources that may be communicatively coupled to the machine 105 … The compute/memory tray 134 may include one or more system-in-packages 136 which can include one or more memory devices 140 and one or more compute devices 160); and
a second memory device configured to communicate with the first memory device via the interface (Pitchumani, [0049], a system-in-package 136 may include one or more interposers 505, one or more memory devices 140; [0051], die-to-die interface 310; [0053], a memory device 140 may be configured as a request initiator to enable data sharing between the memory devices 140; [0065], Fig.5; Fig.6; Fig.7), store data used in a memory-intensive computation, and perform the memory- intensive computation instead of the first memory device (Pitchumani, [0029], the memory device 140 is configured to provide compute and/or memory resources; [0030]; [0053], a memory device 140 may be configured as a request initiator to enable data sharing between the memory devices 140).
Pitchumani teaches a first memory device acts as a request initiator between memory devices which implies the second memory device performs memory intensive computation instead of the first memory device. To advance prosecution, Pitchumani in view of Bari explicitly teaches a second memory device … perform the memory- intensive computation instead of the first memory device (Bari, [0094]; [0096]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Pitchumani to select memory devices to perform specific operations based on their capabilities, strengths, weaknesses, etc. As such, a first memory device acts as a request initiator while a second memory device performs more memory-intensive computations. A person of ordinary skill in the art would have been motivated to combine the teachings of Pitchumani with Bari because it improves efficiency and performance of the storage system disclosed in Pitchumani by ensuing memory devices are assigned with tasks based on the capabilities and characteristics of the memory devices.
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Pitchumani and Bari as applied to claim 1 above, and further in view of O (US2021/0225430), hereinafter O.
Regarding claim 2, the combination of Pitchumani teaches all the features with respect to claim 1 as outlined above. The combination of Pitchumani does not explicitly teach the data processing system of claim 1, wherein the second memory device includes a computing circuit and a memory bank, wherein the memory bank stores the data used in the memory-intensive computation, and wherein the computing circuit performs the memory- intensive computation, as claimed.
However, the combination of Pitchumani in view of O teaches the data processing system of claim 1, wherein the second memory device includes a computing circuit and a memory bank, wherein the memory bank stores the data used in the memory-intensive computation, and wherein the computing circuit performs the memory- intensive computation (O, [0015], The memory device 1000 may correspond to a computational memory device including a random access memory (RAM) and a processing element (PE) integrated in the same die … the PE may be also referred to as a “processor” or a “processing circuit”.; [0021], The PIM die 1100 may include bank groups BG0 to BG3; [0024]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of O to include a computational memory device as the second memory device that performs memory-intensive computation as the computation memory device comprises one or more processing circuits and one or more memory banks. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with O because it improves efficiency and performance of the storage system disclosed in the combination of Pitchumani by ensuing memory devices are assigned with tasks based on capabilities and characteristics of the memory devices.
Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Pitchumani, Bari, and O as applied to claim 2 above, and further in view of Lee et al. (US2024/0112708), hereinafter Lee.
Regarding claim 3, the combination of Pitchumani teaches all the features with respect to claim 2 as outlined above. The combination of Pitchumani does not explicitly teach the data processing system of claim 2, wherein the memory-intensive computation is a part of an inference computation of generating an output corresponding to a received input based on knowledge learned by a language model, as claimed.
However, the combination of Pitchumani in view of Lee teaches the data processing system of claim 2, wherein the memory-intensive computation is a part of an inference computation of generating an output corresponding to a received input based on knowledge learned by a language model (Lee, [0044], a computational memory device 100 may include a weight memory block 110; [0050]; [0052]; [0067]; Bari, [0094], the first chips 718 are configured to execute rules-based control logic and other processing tasks, while the second chips 720 are specialized artificial intelligence chips particularly adapted for use in executing artificial intelligence (AI) algorithms (e.g., deep learning models, large language models).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Lee to use a computation memory device (i.e. the second memory device) for inference computation of generating an output corresponding to a received input in a large language model. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Lee because it improves efficiency and performance to use computational memory devices for executing artificial intelligence (AI) algorithms.
Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Pitchumani, Bari, O, and Lee as applied to claim 3 above, and further in view of Gandhi et al. (US2026/0050805), hereinafter Gandhi, Xiong et al. (US 2026/0161894), hereinafter Xiong, and Yang et al. (US2026/0203368), hereinafter Yang.
Regarding claim 4, the combination of Pitchumani teaches all the features with respect to claim 3 as outlined above. The combination of Pitchumani does not explicitly teaches the data processing system of claim 3, wherein the data processing system comprises the inference computation model, wherein the inference computation comprises a plurality of computation units arranged sequentially, wherein each of the plurality of computation units includes an embedding layer, a plurality of decoder layers, and a head layer, wherein each of the plurality of decoder layers includes a multi-head attention block and a feed-forward block, and wherein the memory-intensive computation includes a matrix multiplication performed in the multi-head attention block, as claimed.
However, the combination of Pitchumani in view of Gandhi teaches the data processing system of claim 3, wherein the data processing system comprises the inference computation model (Pitchumani, [0045]; [0075], the compute/memory tray 134 may be configured as a deployable platform for a large language model (LLM) with compute and/or memory resources capable of performing inference using the LLM), wherein the inference computation comprises a plurality of computation units arranged sequentially (Gandhi, [0049], The plurality of sequentially chained task planners is a plurality of large language models (LLMs); [0105], llama-inference).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Gandhi to include a sequentially arranged language models. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Gandhi because it improves efficiency and performance of the storage system disclosed in the combination of Pitchman by providing the ability to create specialized pipelines where each model builds upon the output of the last, enabling the processing of highly complex, multi-step reason tasks that a single model cannot handle.
The combination of Pitchumani does not explicitly teach wherein each of the plurality of computation units includes an embedding layer, a plurality of decoder layers, and a head layer, wherein each of the plurality of decoder layers includes a multi-head attention block and a feed-forward block, and wherein the memory-intensive computation includes a matrix multiplication performed in the multi-head attention block, as claimed.
However, the combination of Pitchumani in view of Xiong teaches wherein each of the plurality of computation units includes an embedding layer, a plurality of decoder layers, and a head layer (Xiong, [0061], Language model 510 includes token embedding layers 512, one or more decoding layers, such as decoder layer 514, and language model head layers 518.), wherein each of the plurality of decoder layers includes a multi-head attention block (Xiong, [0062], attention block; [0067]) and a feed-forward block (Xiong, [0062], feed-forward block).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Xiong to include an embedding layer, a plurality of decoder layers, and a head layer for a computation unit and each of the plurality of decoder layers include an attention block and a feed-forward block. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Xiong because it improves performance of the storage system disclosed in the combination of Pitchumani by identifying patterns and making inferences from training data (Xiong, [0021]).
The combination of Pitchumani does not explicitly teach multi-head attention block and wherein the memory-intensive computation includes a matrix multiplication performed in the multi-head attention block, as claimed.
However, the combination of Pitchumani in view of Yang teaches each of the plurality of decoder layers includes a multi-head attention block and a feed-forward block (Yang, [0014], LLMs are generally built with sequential transformer layers, where each transformer layer contains a multi-head attention block followed by a feed forward block);
wherein the memory-intensive computation includes a matrix multiplication performed in the multi-head attention block (Yang, [0014], In both multi-head attention block and feed forward block, the primary computations are general matrix multiplication (GEMM) operations or mixed-precision general matrix multiplication (mpGEMM) operations with weight quantization).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Yang to include multi-head attention blocks and feed forward blocks in each of a plurality of decoder layers and use the blocks to perform matrix multiplication computations. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Yang because it improves efficiency of the storage system disclosed in the combination of Pitchumani by allowing models to capture many different types of relationships in data at the same time while learning complex patterns.
Claim(s) 5 and 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Pitchumani, Bari, O, Lee, Gandhi, Xiong, and Yang as applied to claim 5 above, and further in view of Kriman et al. (US2025/0371333), hereinafter Kriman and Zhang et al. (US2026/0236744), hereinafter Zhang.
Regarding claim 5, the combination of Pitchumani teaches all the features with respect to claim 4 as outlined above. The combination of Pitchumani does not explicitly teach the data processing system of claim 4, wherein the memory bank includes a key cache and a value cache, and wherein the computing circuit is configured to: receive query data, first key data, and first value data from the first memory device; generate third key data based on second key data stored in the key cache and the first key data; and generate intermediate data by performing the matrix multiplication between the query data and the third key data, as claimed.
However, the combination of Pitchumani in view of Kriman teaches wherein the memory bank includes a key cache and a value cache (Kriman, [0045], key-value (KV) cache 320), and wherein the computing circuit is configured to: receive query data, first key data, and first value data from the first memory device; generate third key data based on second key data stored in the key cache and the first key data (Kriman, [0048], The computed key Kj and value Vj may be stored in KV cache 320, for use in subsequent iterations; [0133]; [0134]; Fig.11B; Fig.11C); and generate intermediate data by performing the matrix multiplication between the query data and the third key data (Kriman, [0049]; [0132], a self-attention score may be calculated for pairs of tokens by taking the dot product of the query vector with the corresponding key vectors).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Kriman to include a cache memory for key/value and generate multiple levels of key/value data as well as an intermediate data by performing matrix multiplication on query and key. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Kriman because it improves efficiency of the AI model by providing clear inputs which leads to outputs with higher accuracy.
The combination of Pitchumani does not explicitly teach receive query data, first key data, and first value data from the first memory device, as claimed.
However, the combination of Pitchumani in view of Zhang teaches receive query data, first key data, and first value data from the first memory device (transmitting device) (Zhang, [0064]; [0066]; [0067]; [0071]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Pitchumani to incorporate teachings of Zhang to receive query data, first key data, and first value data from a first memory device. A person of ordinary skill in the art would have been motivated to combine the teachings of Pitchumani with Zhang because it improves efficiency of AI model disclosed in Pitchumani by allowing different devices with different characteristics to perform different functions.
Regarding claim 6, the combination of Pitchumani teaches all the features with respect to claim 5 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 5, wherein the second key data is key data generated by a preceding computation unit in a sequence of the plurality of computation units, or key data generated by a preceding decoder layer in a sequence of the plurality of decoder layers (Kriman, [0133], the decoder(s) 1145 form a decoder stack, where each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses the attention vectors (keys and values) from the encoder to focus on relevant parts of the input sequence, and a feedforward network … the encoder-decoder attention layer; Xiong, [0062], decoder layers).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Kriman to include a cache memory for key/value and generate multiple levels of key/value data as well as an intermediate data by performing matrix multiplication on query and key. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Kriman because it improves efficiency of the AI model by providing clear inputs which leads to outputs with higher accuracy.
Claim(s) 7-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Pitchumani, Bari, O, Lee, Gandhi, Xiong, Yang, Kriman, and Zhang as applied to claim 5 above, and further in view of Agranovich et al. (US2025/0348692), hereinafter Agranovich and Yeri et al. (US12,530,510), hereinafter Yeri.
Regarding claim 7, the combination of Pitchumani teaches all the features with respect to claim 5 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 5, wherein the computing circuit is further configured to: generate third value data based on second value data stored in the value cache and the first value data (Kriman, [0048], The computed key Kj and value Vj may be stored in KV cache 320, for use in subsequent iterations; [0133]; [0134]; Fig.11B; Fig.11C);
generate a context vector by performing the matrix multiplication between the intermediate data and the third value data (Kriman, [0019]); and
output the context vector to the first memory device.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Kriman to generate a context vector based on query data, key data, value data, and calculated scores. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Kriman because it improves efficiency of the AI model by providing clear inputs which leads to outputs with higher accuracy.
The combination of Pitchumani does not explicitly teach generating, by the second memory device, the context vector by performing the matrix multiplication between the intermediate data and the third value data, and output the context vector to the first memory device, as claimed.
However, the combination of Pitchumani in view of Agranovich teaches generating, by the second memory device, the context vector by performing the matrix multiplication between the intermediate data and the third value data (Agranovich, [0142], The similarity function may comprise, e.g., a dot product, cosine similarity, or other similarity measure … the attention mechanism can generate an attention context, e.g., a context vector, by multiply the corresponding value of the encoded audio frame, e.g., one of the k-preceding audio frames, by a similarity function of the query; Kriman, [0019], score).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Agranovich to generate a context vector by performing a matrix multiplication between a similarity value and a third value data. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Agranovich because it improves efficiency of AI model by dynamically weighting and blending information based on what is most relevant at an exact moment.
The combination of Pitchumani does not explicitly teach output the context vector to the first memory device, as claimed
However, the combination of Pitchumani in view of Yeri teaches output the context vector to the first memory device (Yeri, col.10, lines 30-40, the apparatus 200 includes means, such as processor 202, memory 204, context vector generation circuitry 212, and/or the like, for generating a plurality of context vectors; col.12, lines 44-59, context vectors may be stored by the modeling system 102 in memory 204).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Yeri to transfer and store context vector in a fast-access memory. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Yeri because it improves efficiency of the storage system disclosed in the combination of Pitchumani by allowing quick access to the most recent context vector.
Regarding claim 8, the combination of Pitchumani teaches all the features with respect to claim 7 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 7, wherein the second value data is value data generated by a preceding computation unit in a sequence of the plurality of computation units, or value data generated by a preceding decoder layer in a sequence of the plurality of decoder layers (Kriman, [0133], the decoder(s) 1145 form a decoder stack, where each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses the attention vectors (keys and values) from the encoder to focus on relevant parts of the input sequence, and a feedforward network … the encoder-decoder attention layer; Xiong, [0062], decoder layers).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Kriman to include a cache memory for key/value and generate multiple levels of key/value data as well as an intermediate data by performing matrix multiplication on query and key. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Kriman because it improves efficiency of the AI model by providing clear inputs which leads to outputs with higher accuracy.
Regarding claim 9, the combination of Pitchumani teaches all the features with respect to claim 7 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 7, wherein the computing circuit is further configured to store the generated third key data in the key cache, and the generated third value data in the value cache (Kriman, [0048], The computed key Kj and value Vj may be stored in KV cache 320, for use in subsequent iterations; [0133]; [0134]; Fig.11B; Fig.11C).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Kriman to include a cache memory for key/value and generate multiple levels of key/value data as well as an intermediate data by performing matrix multiplication on query and key. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Kriman because it improves efficiency of the AI model by providing clear inputs which leads to outputs with higher accuracy.
Regarding claim 10, the combination of Pitchumani teaches all the features with respect to claim 9 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 9, wherein the first memory device includes a plurality of memory dies and a base die (Pitchumani, [0031], base die 150; [0038], the die-to-die interfaces 310 are configured to interface with one or more additional dies and/or various types of compute and/or memory resources), the base die including a controller which controls the plurality of memory dies (Pitchumani, [0033], the memory die 155 includes a memory 202; [0035], In an example in which the memory die 155 illustrated in FIG. 2 is a DRAM die, the first controller 330 may be a memory controller (e.g., a DRAM controller) configured to control the memory 202; Fig.3), and wherein the first memory device is configured to communicate with the host (Pitchumani, [0029]) and the second memory device via the base die (Pitchumani, [0031]; Fig.5; Fig.6).
Regarding claim 11, the combination of Pitchumani teaches all the features with respect to claim 10 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 10, wherein the second memory device is configured as a single package in which the computing circuit and the memory bank are integrated (O, [0015], The memory device 1000 may relate to the PIM … The memory device 1000 may correspond to a computational memory device including a random access memory (RAM) and a processing element (PE) integrated in the same die; [0021], The PIM die 1100 may include bank groups BG0 to BG3).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of O to include a computational memory device as the second memory device that performs memory-intensive computation as the computation memory device comprises one or more processing circuits and one or more memory banks. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with O because it improves efficiency and performance of the storage system disclosed in the combination of Pitchumani by ensuing memory devices are assigned with tasks based on capabilities and characteristics of the memory devices.
Claim(s) 12 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Pitchumani, Bari, O, Lee, Gandhi, Xiong, Yang, Kriman, Zhang, Agranovich, and Yeri as applied to claim 11 above, and further in view of Oh et al. (US2025/0259954), hereinafter Oh.
Regarding claim 12, the combination of Pitchumani teaches all the features with respect to claim 11 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 11, wherein the host and the first memory device are physically coupled via an interposer substrate (Pitchumani, [0029]; [0049]; Fig.4), wherein the first memory device and the second memory device are physically coupled via a packaging substrate, wherein the interposer substrate and the packaging substrate are electrically connected via a connection terminal, and wherein the interposer substrate is disposed on the packaging substrate.
The combination of Pitchumani does not explicitly teach wherein the first memory device and the second memory device are physically coupled via a packaging substrate, wherein the interposer substrate and the packaging substrate are electrically connected via a connection terminal, and wherein the interposer substrate is disposed on the packaging substrate, as claimed.
However, the combination of Pitchumani in view of Oh teaches wherein the first memory device and the second memory device are physically coupled via a packaging substrate (Oh, [0079], The first memory device 400 a and the second memory device 400 b may be mounted on the interposer 210; [0080]; Fig.6;), wherein the interposer substrate and the packaging substrate are electrically connected via a connection terminal (Oh, [0084], The interposer connection terminal 214 may be configured to electrically connect the interposer 210 to the package substrate 110.), and wherein the interposer substrate is disposed on the packaging substrate (Oh, [0080]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Oh to include an interposer substrate disposed on a packaging substrate. The first memory device and the second memory are physically coupled via the packaging substrate. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Oh because it improves reliability of physical memory devices by preventing cracks caused by loads applied to semiconductor substrate (Oh, [0003]-[0004]).
Regarding claim 13, the combination of Pitchumani teaches all the features with respect to claim 1 as outlined above. The combination of Pitchumani further teaches the data processing system of claim 12, wherein the second memory device is disposed at one of sides of the first memory device (Oh, Fig.6, see first memory device 400 a and second memory device 400 b)).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Pitchumani to incorporate teachings of Oh to include an interposer substrate disposed on a packaging substrate. The first memory device and the second memory are physically coupled via the packaging substrate. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Pitchumani with Oh because it improves reliability of physical memory devices by preventing cracks caused by loads applied to semiconductor substrate (Oh, [0003]-[0004]).
Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US2026/0030866), hereinafter Li in view of Zhang et al. (US2026/0236744), hereinafter Zhang, and further in view of Kriman et al. (US2025/0371333), hereinafter Kriman.
Regarding claim 14, Li teaches a method of operating a data processing system which includes a first memory device processing an inference computation and a second memory device communicating with the first memory device, the method comprising:
receiving, by the second memory device, query data, first key data, and first value data from the first memory device (Li, [0040], First query 310, first key 312, and first value 314 are then forwarded to DM-decoder 350 as initial inputs; fig. 3); and
generating, by the second memory device, a context vector by performing a memory-intensive computation based on data stored in the second memory device, the query data, the first key data, and the first value data (Li, [0042], [0046]-[0050]).
Li does not explicitly teach receiving query data, first key data, and first value data from a first memory device, and generating a context vector based on in the second memory device, the query data, the first key data, and the first value data, as claimed.
However, Li in view of Zhang teaches receiving, by the second memory device (receiving device), query data, first key data, and first value data from the first memory device (transmitting device ) (Zhang, [0064]; [0066]; [0067]; [0071]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li to incorporate teachings of Zhang to receive query data, first key data, and first value data from a first memory device. A person of ordinary skill in the art would have been motivated to combine the teachings of Li with Zhang because it improves efficiency of AI model disclosed in Li by allowing different devices with different characteristics to perform different functions.
The combination of Li does not explicitly each first value data from a first memory device, and generating a context vector based on data stored in the second memory device, the query data, the first key data, and the first value data, as claimed.
However, the combination of Li in view of Kriman teaches first value data from a first memory device, and generating a context vector based on in the second memory device, the query data, the first key data, and the first value data (Kriman, [0045]; [0132]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Li to incorporate teachings of Kriman to generate a context vector based on query data, key data, value data, and calculated scores. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Li with Kriman because it improves efficiency of the AI model by providing clear inputs which leads to outputs with higher accuracy.
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Li, Zhang, and Kriman as applied to claim 14 above, and further in view of Agranovich et al. (US2025/0348692), hereinafter Agranovich.
Regarding claim 15, the combination of Li teaches all the features with respect to claim 14 as outlined above. The combination of Li further teaches the method of claim 14, wherein generating, by the second memory device, the context vector by performing the memory-intensive computation based on the data stored in the second memory device, the query data, the first key data, and the first value data comprises:
generating, by the second memory device, third key data based on second key data stored in the second memory device, and the first key data (Li, [0046], a second key and a second value are generated based at least on the first correlation and are forwarded to second multi-head cross-attention layer 330 as inputs; [0048], The output of second multi-head cross-attention layer 330 may be transformed by fourth Add&Norm layer 332 to generate third key 338 and third value 340; Fig.3; Kriman, Fig.11B; Fig.11C);
generating, by the second memory device, intermediate data by performing a matrix multiplication between the third key data and the query data (Kriman, [0049]; [0132], a self-attention score may be calculated for pairs of tokens by taking the dot product of the query vector with the corresponding key vectors);
generating, by the second memory device, third value data based on second value data stored in the second memory device, and the first value data (Li, [0046], second value; [0048], third value; Fig.3; Kriman, Fig.11B; Fig.11C); and
generating, by the second memory device, the context vector by performing the matrix multiplication between the intermediate data and the third value data (Kriman, [0019], score; [0132]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Li to incorporate teachings of Kriman to generate a context vector based on query data, key data, value data, and calculated scores. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Li with Kriman because it improves efficiency of the AI model by providing clear inputs which leads to outputs with higher accuracy.
The combination of Li does not explicitly teach generating, by the second memory device, the context vector by performing the matrix multiplication between the intermediate data and the third value data, as claimed.
However, the combination of Li in view of Agranovich teaches generating, by the second memory device, the context vector by performing the matrix multiplication between the intermediate data and the third value data (Agranovich, [0142], The similarity function may comprise, e.g., a dot product, cosine similarity, or other similarity measure … the attention mechanism can generate an attention context, e.g., a context vector, by multiply the corresponding value of the encoded audio frame, e.g., one of the k-preceding audio frames, by a similarity function of the query; Kriman, [0019], score).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Li to incorporate teachings of Agranovich to generate a context vector by performing a matrix multiplication between a similarity value and a third value data. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Li with Agranovich because it improves efficiency of AI model by dynamically weighting and blending information based on what is most relevant at an exact moment.
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Li, Zhang, Kriman, and Agranovich as applied to claim 15 above, and further in view of Yeri et al. (US12,530,510), hereinafter Yeri.
Regarding claim 16, the combination of Li teaches all the features with respect to claim 15 as outlined above. The combination of Li does not explicitly teach the method of claim 15, further comprising outputting, by the second memory device, the context vector to the first memory device, as claimed.
However, the combination of Li in view of Yeri teaches the method of claim 15, further comprising outputting, by the second memory device, the context vector to the first memory device (Yeri, col.10, lines 30-40, the apparatus 200 includes means, such as processor 202, memory 204, context vector generation circuitry 212, and/or the like, for generating a plurality of context vectors; col.12, lines 44-59, context vectors may be stored by the modeling system 102 in memory 204).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Li to incorporate teachings of Yeri to transfer and store context vector in a fast-access memory. A person of ordinary skill in the art would have been motivated to combine the teachings of the combination of Li with Yeri because it improves efficiency of the storage system disclosed in the combination of Li by allowing quick access to the most recent context vector.
Allowable Subject Matter
Claims 17-20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Claim 17 recites “[t]he method of claim 16, further comprising: storing, by the first memory device, the context vector received from the second memory device; and performing, by the first memory device, a subsequent computation of the inference computation, wherein the subsequent computation includes processing the context vector to generate an output tensor”.
The above-noted limitation, in combination with the other limitation of the claims, are neither disclosed nor suggested by the prior art of record. Therefore, in the context of claim 14, 15, 16, and 17 as a whole, the prior art does not teach the claimed subject matter. Thus, the subject matter of claim 17 is allowable. Claims 18-20 depend on claim 17 and they are objected for the same reasons.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Engel (US2026/0241360) teaches generating context vector using key, value, query, and attention score.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NANCI N WONG whose telephone number is (571)272-4117. The examiner can normally be reached Monday-Friday 9am -6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Arpan Savla can be reached at 571-272-1077. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NANCI N WONG/Primary Examiner, Art Unit 2137