DETAILED ACTION
This Office Action is in response to communications filed on May 5th, 2026 for Application No. 17/740,497, in which claims 1-20 are presented for examination. The amendments filed May 5th, 2026 have been entered, where claims 1, 3-4, 6-7, 9, 11, and 14 are amended.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-4, 6-7, 10-11, 14-16, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Lu et al. (hereinafter Lu) (“Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture”) in view of Beltagy et al. (hereinafter Beltagy) (“Longformer: The Long-Document Transformer”) and Tay et al. (hereinafter Tay) (“Synthesizer: Rethinking Self-Attention in Transformer Models”).
Regarding Claim 1, Lu teaches a method comprising: generating a task-specific machine-learning model by training a generic machine learning model on task specific training data associated with a specific task (Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”, where the “pre-trained checkpoint” is a generic model, see also Pg. 988, Col. 2, Para. 2, “Model: We use BERT-Base-Uncased (420.1MB), GPT2-Small (522.7MB) and BART-Base (532.1MB). Our scripts download pre-trained checkpoints from the Hugging Face Model Hub https://huggingface.co/models automatically” where at least the model “BERT” is well-established as a generically trained machine-learning model because it is trained on data that is not task-specific, which is used to generate a task-specific model by “fine-tun[ing]” on “downstream tasks”; Pg. 989, Col. 1, Para. 5, “Data sets. We evaluate models on three datasets, namely GLUE, SQuAD, and CLOTH. They correspond to three different NLP tasks . . . Our script automatically downloads the GLUE and SQuAD datasets before training”, where the “data sets” used for “training” correspond to three specific “different NLP tasks”);
identifying [aspects in] each . . . [of the rows] and [aspects in] each . . . [of the columns] in an attention matrix of the task-specific machine-learning model (Pg. 981, Col. 1, Para. 6, “after obtaining a quantized approximation Sˆ of the attention matrix, we generate a binary attention mask M according to the sparsity pattern it exhibits”, where the “binary attention mask” identifies aspects of the “attention matrix” by masking unidentified values, such that aspects in each of the rows and aspects in each of the columns are identified, which is of the task-specific machine-learning model, see Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”; Pg. 980, Col. 2, Para. 1, “In the inference phase . . . Sanger generates sparse masks dynamically”; and Pg. 985, Col. 1-2, Para. 3-1, “We obtain the attention mask S by applying a binary threshold on a low-precision estimation ˆP of the attention matrix . . . The results on BERT [16] are shown in Table 3. The column headings are NLP tasks. For a given Transformer network, it can be applied to multiple tasks. The row labeled "baseline" corresponds to the original BERT-base (with dense attention) for these tasks”; for more information see Pg. 980, Figure 2 and Pg. 980, Algorithm 1)
having an importance score for the specific task (Pg. 978, Col. 1, Para. 2, “The first step of computing attention is to obtain a score matrix . . . often referred to as the attention matrix”; Pg. 979, Col. 1, Para. 1, “the score matrix . . . represents the importance of each input token when producing an output element”, which is for the specific “inference” “task”, see Pg. 980, Col. 2, Para. 1, “In the inference phase . . . Sanger generates sparse masks dynamically” and Pg. 985, Col. 1-2, Para. 3-1, “We obtain the attention mask S by applying a binary threshold on a low-precision estimation ˆP of the attention matrix . . . The results on BERT [16] are shown in Table 3. The column headings are NLP tasks. For a given Transformer network, it can be applied to multiple tasks. The row labeled "baseline" corresponds to the original BERT-base (with dense attention) for these tasks”)
above a threshold importance score to form a plurality of . . . [aspects], wherein the importance score is generated . . . (Pg. 985, Col. 1, Para. 3, “We obtain the attention mask S by applying a binary threshold on a low-precision estimation Pˆ of the attention matrix”; Pg. 981, Equation
PNG
media_image1.png
78
430
media_image1.png
Greyscale
, where “importance” scores are generated and compared to a “threshold” to form a plurality of aspects that have not been “mask[ed]” “by applying a binary threshold”, see Pg. 978, Col. 1, Para. 2, “The first step of computing attention is to obtain a score matrix . . . often referred to as the attention matrix”; Pg. 979, Col. 1, Para. 1, “the score matrix . . . represents the importance of each input token when producing an output element”)
. . . training of the task-specific machine-learning model with the task-specific training data (Pg. 988, Col. 2, Para. 2, “Model: We use BERT-Base-Uncased (420.1MB), GPT2-Small (522.7MB) and BART-Base (532.1MB). Our scripts download pre-trained checkpoints from the Hugging Face Model Hub https://huggingface.co/models automatically” and Pg. 984, Col. 2, Para. 2, “We use MKL DNN[23] and CuDNN[10] as operator libraries for Intel CPUs and NVIDIA GPUs. Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”, where at least the model “BERT” is well-established as a generically trained machine-learning model because it is trained on data that is not task-specific, which is used to generate a task-specific model by “fine-tun[ing]” on “downstream tasks”; Pg. 989, Col. 1, Para. 5, “Data sets. We evaluate models on three datasets, namely GLUE, SQuAD, and CLOTH. They correspond to three different NLP tasks . . . Our script automatically downloads the GLUE and SQuAD datasets before training”, where the “data sets” used for “training” correspond to three specific “different NLP tasks”; notably, training of the task-specific model is interpreted as the training that generated the task specific model)
and without pre-training the model on generic training data to adapt to the attention matrix (Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”);
including the plurality of . . . [aspects] in an adaptive attention pattern, wherein the adaptive attention pattern is . . . determined . . . (Pg. 980, Col. 1, Para. 2, “The resulting attention mask exhibits an unstructured sparsity pattern”; Pg. 985, Col. 1, Para. 3, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”, where “dynamic patterns” “conditioned on individual input[s]”, is within the broadest reasonable interpretation of adaptive; Pg. 979, Table 1, “Existing sparse attention patterns”, where the “sparsity pattern” is an attention pattern, which is determined through inclusion of the plurality of aspects as its pattern, “The resulting attention mask exhibits an unstructured sparsity pattern” and “the sparsity patterns we generate”)
the training of the task-specific machine-learning model (Pg. 984, Col. 2, Para. 2, “We use MKL DNN[23] and CuDNN[10] as operator libraries for Intel CPUs and NVIDIA GPUs. Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”, where a task-specific model is generated through training by “fine-tun[ing]” on “downstream tasks”; Pg. 989, Col. 1, Para. 5, “Data sets. We evaluate models on three datasets, namely GLUE, SQuAD, and CLOTH. They correspond to three different NLP tasks . . . Our script automatically downloads the GLUE and SQuAD datasets before training”);
storing the adaptive attention pattern (Pg. 985, Col. 1-2, Para. 3-1, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns can better adapt to data, thereby improving the sparsity accuracy curve. The results on BERT [16] are shown in Table 3. The column headings are NLP tasks. For a given Transformer network, it can be applied to multiple tasks. The row labeled "baseline" corresponds to the original BERT-base (with dense attention) for these tasks”, where the adaptive attention pattern, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”, must be stored, at least temporarily, to be used to achieve “results”; see also Pg. 980, Col. 2, Para. 1, “In the inference phase . . . Sanger generates sparse masks dynamically”, where, similarly, the “spare masks” must be stored, at least temporarily, to be used “[i]n the inference phase”); and
in response to an input, generating a task-specific inference for the input using the task-specific machine-learning model (Pg. 985, Table 3, where the model’s “accuracy”, which requires output in response to an input, for ten specific tasks is displayed; where the tasks are inferencing tasks, see Pg. 980, Col. 2, Para. 1, “In the inference phase”; and where, as discussed above, the “model” is a task-specific machine learning model, see Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”)
with the adaptive attention pattern (Pg. 989, Col. 1, Para. 8, “Experimental workflow . . . Train a model with Sanger sparse attention”).
Lu does not explicitly disclose . . . row . . . column . . . rows and/or columns . . . only during the . . . rows and/or columns . . . only . . . during . . . (where the aspects are not specifically described as rows or columns and the importance scores and adaptive patterns are not specifically described as only determined during training).
However, Beltagy teaches . . . [identifying a] row . . . [and] . . . column . . . [to form a plurality of] . . . rows and/or columns [in an attention matrix, for inclusion of the] . . . rows and/or columns [in an attention pattern used by a machine learning model] (Pg. 3, figure 2(d), where the “Longformer” “attention pattern” includes selected rows and columns, identified by green shading; for more information, see Pg. 3, Col. 2, Para. 2, “The original Transformer model has a self-attention component with O(n2) time and memory complexity . . . To address this challenge, we sparsify the full self-attention matrix according to an attention pattern specifying pairs of input locations”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the identifying, to form a plurality of aspects and based on a threshold importance score for a specific task, aspects in each of the rows and aspects in each of the columns in an attention matrix of the task-specific machine-learning model and including the plurality of aspects in an adaptive attention pattern of Lu with the identifying rows and columns of the attention matrix to form a plurality of rows and columns, which are selected for inclusion in an attention pattern of Beltagy, in order to utilize Lu’s threshold-based aspect selection method at the column and row specificity level, which will maintain the performance improvements of both methods (Lu, Pg. 985, Table 3, where “Sanger” has improved “accuracy” and “sparsity”; Beltagy, Pg. 10, Col. 2, “pretrained, Longformer consistently outperforms RoBERTa on long document tasks and sets new state-of-the-art results on WikiHop and TriviaQA.”), while preserving the model’s ability to handle long input sequences (compare Beltagy, Pg. 10, Col. 1, Para. 2, “Longformer . . . perform[s] . . . NLP tasks without chunking/shortening the long input . . . while also scaling linearly with the sequence length” with Lu, Pg. 985, Col. 2, Para. 2, “While [longformer and bigbird] are originally proposed for processing long sequences (e.g., text length 4096), we scale them for standard benchmarks with shorter contexts”), and allowing rows or columns to be selected, which will result in easy and simple inclusion of inductive information (Beltagy, Pg. 4, col. 1-2, Para. 5-1, “While specifying global attention is task specific, it is a easy way to add inductive bias to the model’s attention, and it is much simpler than existing task specific approaches that use complex architecture to combine information across smaller input chunks”).
Additionally, Tay teaches . . . [an attention pattern generation method] (Pg. 4, Fig. 1, “Our proposed SYNTHESIZER model architecture”, where both “(b) Synthesizer (Dense)” and “Synthesizer (Random)” generate attention patterns; see also Pg. 3, Para. 2, “Our work is a novel take on the self-attention mechanism in Transformer models. We delve deeper, starting with replacing the pairwise dot products with what we call synthesizing functions that learn attention matrices that may or may not depend on the input tokens. The most closely related work is [Raganato et al., 2020], in which the authors propose using fixed (i.e., not learned) attention patterns . . . Our work takes this intuition further and expands on this narrative” and Pg. 2, Para. 5, “we generate the alignment matrix independent of token-token dependencies and explore a potpourri of parameterized functions for synthesizing attention matrices”)
[, wherein, an importance score for a specific task is generated] only during the [training of the model and the attention pattern is determined] . . . only . . . during . . . [the training of the model] (Pg. 3-4, Para. 8-1, “We consider another variation of SYNTHESIZER where the attention weights are not conditioned on any input tokens. Instead, the attention weights are initialized to random values. These values can then either be trainable . . . Let R be a randomly initialized matrix. The Random Synthesizer is defined as: Y = Softmax(R)G(X) . . . The basic idea2 of the Random Synthesizer is to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, where the importance scores, “the attention weights”, are generated only during training, “These values can then . . . be trainable . . . to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, and, as a result, the attention patterns that are made up of the “attention weights” are generated only during training, see Pg. 4, Fig. 1, “Our proposed SYNTHESIZER model architecture”, where both “(b) Synthesizer (Dense)” and “Synthesizer (Random)” generate attention patterns).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the identifying of each row and each column in an attention matrix of a task-specific machine-learning model having an importance score for a specific task above a threshold, including the rows and columns in an adaptive attention pattern, and training the task-specific machine learning model on specific training data without pre-training on generic training data of Lu in view of Beltagy with the attention pattern generation method, wherein, an importance score for a specific task is generated only during the training of the model and the attention pattern is determined only during the training of the model of Tay in order to learn task-specific attention patterns that work well across global samples (Tay, Pg. 4, Para. 1, “The basic idea2 of the Random Synthesizer is to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”), which reduces computational complexity and parameter costs while preserving performance (Tay, Pg. 6, Para. 1, “In general, the performance of SYNTHESIZER variants are competitive with standard Transformers for this task. Furthermore, SYNTHESIZER variants have reduced computational complexity and parameter costs that are about 10% lower than Transformers”; see also Tay, Pg. 2, Para. 3, “Aside from generalizing the standard Transformer model, we show that it is possible to achieve competitive results with fully global attention weights that do not consider token-token interactions or any instance-level (local) information at all. More specifically, a random matrix SYNTHESIZER model achieves a 27.27 BLEU score on WMT 2014 English-German1”).
Regarding Claim 3, Lu in view of Beltagy and Tay teach the method of claim 1, wherein the adaptive attention pattern (Lu, Pg. 985, Col. 1, Para. 3, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”; where, in view of Beltagy, adaptive during training to determine rows and columns, see Beltagy, Pg. 3, figure 2(d))
assigns global attention to tokens in a row or a column of the plurality of rows and/or columns (Beltagy, Pg. 4, Col. 1, Para. 5, “we add global attention on few pre-selected input locations . . . Fig. 2d shows an example of a sliding window attention with global attention at a few tokens at custom locations”; Beltagy, Pg. 3, Fig. 2d, where example rows and columns with “global attention” tokens are shaded green).
The reasons of obviousness have been noted in the rejection of Claim 1 above and remain applicable here.
Regarding Claim 4, Lu in view of Beltagy and Tay teach the method of claim 1, wherein the adaptive attention pattern is a merger of a row or a column of the plurality of rows and/or columns with a diagonal attention pattern (Beltagy, Pg. 3, Fig. 2d, where the “attention pattern” includes rows and columns merged with “Sliding window attention”, which is a diagonal attention pattern; Beltagy, Pg. 2, Col. 1, Para. 3, “Longformer’s attention mechanism is a combination of a windowed local-context self-attention and an end task motivated global attention that encodes inductive bias about the task”).
The reasons of obviousness have been noted in the rejection of Claim 1 above and remain applicable here.
Regarding Claim 6, Lu in view of Beltagy and Tay teach the method of claim 1, wherein the task-specific machine-learning model is a transformer model (Lu, Pg. 984, Col. 2, Para. 2, “Our code is based on the BERT implementation by NVIDIA [39] and the evaluation code is from Hugging Face’s Transformers library [60]”, where “BERT” is Bidirectional Encoder Representations from Transformers; Lu, Pg. 979, Col. 1, Para. 1, “The attention mechanism is the key operation in the Transformer models . . . Figure 1 (a) depicts the computation stages of the self-attention (abbreviated as attention) mechanism”, where, as discussed above, the “model” is a task-specific machine learning model, see Lu, Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”).
Regarding Claim 7, Lu in view of Beltagy and Tay teach a non-transitory computer-readable medium storing computer-executable instructions (Lu, Pg. 988, Col. 2, Section “A.2 Artifact check-list (meta-information)”, “How much disk space are required (approximately)?: The codebase and downloaded datasets take up about 1.5GB in total”, where “the codebase” is computer-executable instructions that must be downloaded to a “disk”, which is a non-transitory computer-readable medium)
that, when executed by a processing device, cause the processing device to perform operations (Lu, Pg. 988, Col. 2, Section “A.2 Artifact check-list (meta-information)”, “Hardware: NVIDIA Tesla V100-PCIE-16GB GPU, AMD Ryzen Threadripper 3970X CPU”) comprising:
generating a sparse-attention model by adding a sparse attention pattern (Lu, Pg. 989, Col. 1, Para. 8, “Train a model with Sanger sparse attention”)
to a pre-trained machine-learning model having a self-attention operation (Lu, Pg. 984, Col. 1, Para. 4, “Implementation Details The first stage of the attention mechanism is to calculate queries”, where the “Sanger” model is implemented with a “self-attention (abbreviated as attention)” operation, see Pg. 979, Col. 1, Para. 1, and is “pre-trained”, see Lu, Pg. 984, Col. 2, Para. 2);
generating a tuned sparse-attention model by fine tuning the sparse- attention model to perform a task with task-specific training (Lu, Pg. 994, Col. 2, Para. 2, “we directly fine-tune a pre-trained checkpoint on downstream tasks”; Beltagy, Pg. 9, Col. 1, Para. 1, “Longformer can learn to use long range context in task specific fine-tuning with large training datasets such as WikiHop”),
wherein the sparse- attention model is an adaptive attention pattern (Lu, Pg. 980, Col. 1, Para. 2, “The resulting attention mask exhibits an unstructured sparsity pattern”; Lu, Pg. 985, Col. 1, Para. 3, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”, where “dynamic patterns” “conditioned on individual input[s]” is within the broadest reasonable interpretation of adaptive),
and wherein the adaptive attention pattern is learned only during fine-tuning of the sparse-attention model using task- specific training data (Lu, Pg. 985, Col. 1, Para. 3, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”, where, in view of Tay, the adaptive attention pattern is learned only during fine-tuning of the model, Tay, Pg. 3-4, Para. 8-1, “We consider another variation of SYNTHESIZER where the attention weights are not conditioned on any input tokens. Instead, the attention weights are initialized to random values. These values can then either be trainable . . . Let R be a randomly initialized matrix. The Random Synthesizer is defined as: Y = Softmax(R)G(X) . . . The basic idea2 of the Random Synthesizer is to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, where the importance scores, “the attention weights”, are generated only during training, “These values can then . . . be trainable . . . to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, and, as a result, the attention patterns that are made up of the “attention weights” are generated only during training, see Tay, Pg. 4, Fig. 1, “Our proposed SYNTHESIZER model architecture”, where both “(b) Synthesizer (Dense)” and “Synthesizer (Random)” generate attention patterns; Lu, Pg. 989, Col. 1, Para. 5, “Data sets. We evaluate models on three datasets, namely GLUE, SQuAD, and CLOTH. They correspond to three different NLP tasks . . . Our script automatically downloads the GLUE and SQuAD datasets before training”, where the “data sets” used for fine-tuning are task-specific because they correspond to three specific “different NLP tasks”);
storing the adaptive attention pattern (Lu, Pg. 985, Col. 1-2, Para. 3-1, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns can better adapt to data, thereby improving the sparsity accuracy curve. The results on BERT [16] are shown in Table 3. The column headings are NLP tasks. For a given Transformer network, it can be applied to multiple tasks. The row labeled "baseline" corresponds to the original BERT-base (with dense attention) for these tasks”, where the adaptive attention pattern, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”, must be stored, at least temporarily, to be used to achieve “results”; see also Lu, Pg. 980, Col. 2, Para. 1, “In the inference phase . . . Sanger generates sparse masks dynamically”, where, similarly, the “spare masks” must be stored, at least temporarily, to be used “[i]n the inference phase”);
storing the tuned sparse-attention model (Lu, Pg. 988, Col. 2, Section “A.2 Artifact check-list (meta-information)”, “you need to make sure you have enough space for the checkpoints. Each checkpoint takes up about 500 MB of space”, where “checkpoints” include the tuned sparse-attention model, see Lu, Pg. 989, Col. 1, Para. 6, “Models . . . you can also download fine-tuned checkpoints and evaluate them directly”); and
in response to an input, generating a task-specific inference for the input using the tuned sparse-attention model with the adaptive attention pattern (Lu, Pg. 985, Table 3, where the model’s “accuracy”, which requires output in response to an input, for ten specific tasks is displayed; and where the tasks are inferencing tasks, see Lu, Pg. 980, Col. 2, Para. 1, “In the inference phase”, which uses the tuned sparse-attention model, Lu, Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”, with the adaptive attention pattern, see Lu, Pg. 980, Col. 2, Para. 1, “In the inference phase . . . Sanger generates sparse masks dynamically” and Tay, Pg. 4, Para. 1, “The basic idea2 of the Random Synthesizer is to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”).
The reasons of obviousness have been noted in the rejection of Claim 1 above and remain applicable here.
Regarding Claim 10, Lu in view of Beltagy and Tay teach the non-transitory computer-readable medium of claim 7, wherein the adaptive attention pattern includes a row or a column (Beltagy, Pg. 3, figure 2(d), where the “Longformer” “attention pattern” includes selected rows and columns, identified by green shading)
in an attention matrix (Lu, Pg. 981, Col. 1, Para. 6, “after obtaining a quantized approximation Sˆ of the attention matrix, we generate a binary attention mask M according to the sparsity pattern it exhibits”, where the “binary attention mask” identifies aspects of the “attention matrix” by masking unidentified values)
with a task-specific importance score (Lu, Pg. 978, Col. 1, Para. 2, “The first step of computing attention is to obtain a score matrix . . . often referred to as the attention matrix”; Lu, Pg. 979, Col. 1, Para. 1, “the score matrix . . . represents the importance of each input token when producing an output element”, which is task specific)
that is above a threshold importance score (Lu, Pg. 985, Col. 1, Para. 3, “We obtain the attention mask S by applying a binary threshold on a low-precision estimation Pˆ of the attention matrix”; Lu, Pg. 981, Equation
PNG
media_image1.png
78
430
media_image1.png
Greyscale
);
The reasons of obviousness have been noted in the rejection of Claim 1 above and remain applicable here.
Regarding Claim 11, the additional elements of the dependent claim are substantially the same as the limitations of Claim 3, therefore it is rejected under the same rationale.
Regarding Claim 14, Lu in view of Beltagy and Tay teach a system (Lu, Pg. 988, Col. 2, Section “A.2 Artifact check-list (meta-information)”, “Run-time environment: . . . Hardware: . . .”; for more information see Lu, Pg. 982 – 983, Section “Hardware Dataflow”) comprising:
a memory component (Lu, Pg. 984, Table 2, “Memory 128KB query buffer, 128KB key buffer, 128KB value buffer, 128KB output buffer”, where the “buffers” must be stored in a memory component); and
a processing device coupled to the memory component (Lu, Pg. 988, Col. 2, Section “A.2 Artifact check-list (meta-information)”, “Hardware: NVIDIA Tesla V100-PCIE-16GB GPU, AMD Ryzen Threadripper 3970X CPU”, which is known to be coupled to memory), the processing device to perform operations comprising:
identifying, during a task-specific fine tuning operation of a generically trained machine- learning model (Beltagy, Pg. 9, Col. 1, Para. 1, “Longformer can learn to use long range context in task specific fine-tuning with large training datasets such as WikiHop”; Tay, Pg. 3-4, Para. 8-1, “We consider another variation of SYNTHESIZER where the attention weights are not conditioned on any input tokens. Instead, the attention weights are initialized to random values. These values can then either be trainable . . . Let R be a randomly initialized matrix. The Random Synthesizer is defined as: Y = Softmax(R)G(X) . . . The basic idea2 of the Random Synthesizer is to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, where the importance scores, “the attention weights”, are generated only during training, “These values can then . . . be trainable . . . to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, see Lu, Pg. 984, Col. 2, Para. 1, “We evaluate our method on BERT [16], GPT-2 [45], and BART [30]”, where at least the model “BERT” is well-established as a generically trained machine-learning model because it is trained on data that is not task-specific; see also Lu, Pg. 980, Col. 1-2, Para. 4-1, “The weights are dense and fine-tuned to make it easier to form” and Lu, Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks. We use a pruning threshold of 2e-3 for SQuAD and CLOTH, and 2e-2 for the remaining GLUE tasks”)
having a self-attention operation (Lu, Pg. 984, Col. 1, Para. 4, “Implementation Details The first stage of the attention mechanism is to calculate queries”, where the “Sanger” model is implemented with a “self-attention (abbreviated as attention)” operation, see Lu, Pg. 979, Col. 1, Para. 1),
a row or a column (Beltagy, Pg. 3, figure 2(d), where the “Longformer” “attention pattern” includes selected rows and columns, identified by green shading)
in an attention matrix (Lu, Pg. 981, Col. 1, Para. 6, “after obtaining a quantized approximation Sˆ of the attention matrix, we generate a binary attention mask M according to the sparsity pattern it exhibits”, where the “binary attention mask” identifies aspects of the “attention matrix” by masking unidentified values)
with a task-specific importance score (Lu, Pg. 978, Col. 1, Para. 2, “The first step of computing attention is to obtain a score matrix . . . often referred to as the attention matrix”; Lu, Pg. 979, Col. 1, Para. 1, “the score matrix . . . represents the importance of each input token when producing an output element”, which is task specific)
that is above a threshold importance score (Lu, Pg. 985, Col. 1, Para. 3, “We obtain the attention mask S by applying a binary threshold on a low-precision estimation Pˆ of the attention matrix”; Lu, Pg. 981, Equation
PNG
media_image1.png
78
430
media_image1.png
Greyscale
);
including the row or the column in an adaptive attention pattern (Beltagy, Pg. 3, figure 2(d), where the “Longformer” “attention pattern” includes selected rows and columns, identified by green shading)
used with the machine-learning model to limit self-attention operations performed while making an inference (Lu, Pg. 987, Col. 1, Para. 1, “we obtain the attention mask by applying a binary threshold T to the predicted attention matrix. Naturally, the larger T becomes, the more connections are pruned in the attention mechanism”, which is used for an inferencing task, see Lu, Pg. 980, Col. 2, Para. 1, “In the inference phase”),
wherein the adaptive attention pattern is only determined during the training of the task-specific machine-learning model (Lu, Pg. 980, Col. 1, Para. 2, “The resulting attention mask exhibits an unstructured sparsity pattern”; Lu, Pg. 985, Col. 1, Para. 3, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”, where “dynamic patterns” “conditioned on individual input[s]”, is within the broadest reasonable interpretation of adaptive; Lu, Pg. 979, Table 1, “Existing sparse attention patterns”, where the “sparsity pattern” is an attention pattern, which is determined through inclusion of the plurality of aspects as its pattern, “The resulting attention mask exhibits an unstructured sparsity pattern” and “the sparsity patterns we generate”, which, in view of Tay, is only determined during the training of the task-specific machine-learning model, see Tay, Pg. 3-4, Para. 8-1, “We consider another variation of SYNTHESIZER where the attention weights are not conditioned on any input tokens. Instead, the attention weights are initialized to random values. These values can then either be trainable . . . Let R be a randomly initialized matrix. The Random Synthesizer is defined as: Y = Softmax(R)G(X) . . . The basic idea2 of the Random Synthesizer is to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, where the importance scores, “the attention weights”, are generated only during training, “These values can then . . . be trainable . . . to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples”, and, as a result, the attention patterns that are made up of the “attention weights” are generated only during training, see Tay, Pg. 4, Fig. 1, “Our proposed SYNTHESIZER model architecture”, where both “(b) Synthesizer (Dense)” and “Synthesizer (Random)” generate attention patterns; see also Lu, Pg. 984, Col. 2, Para. 2, “We use MKL DNN[23] and CuDNN[10] as operator libraries for Intel CPUs and NVIDIA GPUs. Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”, where a task-specific model is generated through training by “fine-tun[ing]” on “downstream tasks” and Lu, Pg. 989, Col. 1, Para. 5, “Data sets. We evaluate models on three datasets, namely GLUE, SQuAD, and CLOTH. They correspond to three different NLP tasks . . . Our script automatically downloads the GLUE and SQuAD datasets before training”);
storing the adaptive attention pattern (Lu, Pg. 985, Col. 1-2, Para. 3-1, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns can better adapt to data, thereby improving the sparsity accuracy curve. The results on BERT [16] are shown in Table 3. The column headings are NLP tasks. For a given Transformer network, it can be applied to multiple tasks. The row labeled "baseline" corresponds to the original BERT-base (with dense attention) for these tasks”, where the adaptive attention pattern, “the sparsity patterns we generate are conditioned on individual input samples. Such dynamic patterns”, must be stored, at least temporarily, to be used to achieve “results”; see also Lu, Pg. 980, Col. 2, Para. 1, “In the inference phase . . . Sanger generates sparse masks dynamically”, where, similarly, the “spare masks” must be stored, at least temporarily, to be used “[i]n the inference phase”); and
in response to an input, generating a task-specific inference for the input using the machine-learning model (Lu, Pg. 985, Table 3, where the model’s “accuracy”, which requires output in response to an input, for ten specific tasks is displayed; and where the tasks are inferencing tasks, see Lu, Pg. 980, Col. 2, Para. 1, “In the inference phase”)
with the adaptive attention pattern (Lu, Pg. 989, Col. 1, Para. 8, “Experimental workflow . . . Train a model with Sanger sparse attention”).
The reasons of obviousness have been noted in the rejection of Claim 1 above and remain applicable here.
Regarding Claim 15, Lu in view of Beltagy and Tay teach the system of claim 14, wherein the machine-learning model is not retrained on a generic task after adding the adaptive attention pattern to the machine-learning model (Lu, Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”).
Regarding Claim 16, the additional elements of the dependent claim are substantially the same as the limitations of Claim 3, therefore it is rejected under the same rationale.
Regarding Claim 18, Lu in view of Beltagy and Tay teach the system of claim 14, wherein the operations further comprise learning different adaptive attention patterns for different layers of the machine-learning model (Beltagy, Pg. 5, Col. 1, Para. 2-3, “we use differing window sizes across the layers. In particular, we use small window sizes for the lower layers and increase window sizes as we move to higher layers”; Beltagy, Pg. 15, Table 12, “Dilation (small model)”, where the window differs across layers, which is a known component of “attention patterns”, see Beltagy Pg. 3, Figure 2).
The reasons of obviousness have been noted in the rejection of Claim 1 above and remain applicable here.
Claims 2, 8, 13, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Lu in view of Beltagy, Tay, and Chang et al. (hereinafter Chang) (“End-to-End ASR with Adaptive Span Self-Attention”).
Regarding Claim 2, Lu in view of Beltagy and Tay teach the method of claim 1 . . . .
Lu in view of Beltagy and Tay do not teach . . . wherein the adaptive attention pattern is for a single layer of the machine-learning model.
However, Chang teaches [a method] . . . where an adaptive attention pattern is for a single layer of the machine-learning model (Pg. 3595, Abstract, “we propose to use a technique called adaptive span self-attention . . . [which] enables the network to learn an appropriate size and position of the window for each layer and head”, where “each” indicates a given attention pattern is for a single layer).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the identification of an important row or column in an attention matrix, inclusion of the row or column in an adaptive attention pattern and use of the adaptive attention pattern to generate an inference of Lu in view of Beltagy and Tay, with the adaptive attention pattern for a single layer of Chang, in order to account for behavioral differences across model layers, which impacts performance (Chang, Pg.3596, Col. 2, Para. 5-6, “the behavior of each head at every layer is not necessarily the same, and using a single span size hyperparameter W (or Wl and Wr) for all the self-attention computations is not appropriate . . . the motivation is to learn the appropriate span size at each self-attention head and layer during training”; Chang, Pg. 3595, Abstract, “the proposed adaptive span methods consistently improved the performance from the conventional fixed span methods”).
Regarding Claim 8, the additional elements of the dependent claim are substantially the same as the limitations of Claim 2, therefore it is rejected under the same rationale.
Regarding Claim 13, Lu in view of Beltagy and Tay teach the non-transitory computer-readable medium of claim 8, wherein the sparse-attention model is not retrained on a generic task after adding the adaptive attention pattern to the sparse-attention model (Lu, Pg. 984, Col. 2, Para. 2, “Since our method does not require pre-training, we directly fine-tune a pre-trained checkpoint on downstream tasks”, where, as discussed above, the model is a sparse-attention model, see Lu, Pg. 989, Col. 1, Para. 8, “Train a model with Sanger sparse attention”; see also Lu, Pg. 984, Col. 1, Para. 4, “Implementation Details The first stage of the attention mechanism is to calculate queries”, where the “Sanger” model is implemented with a “self-attention (abbreviated as attention)” operation).
Regarding Claim 17, the additional elements of the dependent claim are substantially the same as the limitations of Claim 2, therefore it is rejected under the same rationale.
Claims 5 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lu in view of Beltagy, Tay, and Liu et al. (hereinafter Liu) (“Transformer Acceleration with Dynamic Sparse Attention”).
Regarding Claim 5, Lu in view of Beltagy and Tay teach the method of claim 1 . . .
Lu in view of Beltagy and Tay do not specifically teach . . . wherein the method further comprises controlling a sparsity of the adaptive attention pattern to a sparsity range.
However, Liu teaches . . . wherein the method further comprises controlling a sparsity of the adaptive attention pattern to a sparsity range (Pg. 5, Col. 2, Para. 1, “Different percentage numbers indicate the sparsity ratio that we applied to the DSA models. For instance, DSA-90% means that we only keep 10% of the attention weights in each row of the attention matrix, while masking out all the other 90% of the weights”, where a sparsity ratio contains a range, such as 0.945 – 0.954 for “DSA-95%”, which varies based on level of precision and rounding practices; and where the adaptive attention pattern is based on the attention matrix, see Pg. 981, Col. 1, Para. 6).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the identification of rows in columns in an attention matrix based on an importance threshold, inclusion of the rows and columns in an attention matrix, and use of the attention matrix by a machine learning model of Lu in view of Beltagy and Tay, with the sparsity range controlling for the adaptive attention pattern of Liu in order to balance sparsity ranges with accuracy requirements, which may vary from by task (Liu, Pg. 5, Col. 2, Para. 2, “DSA delivers slightly higher performance with 90% and 95% sparsity ratio. Even with up to 99% of sparsity, DSA still demonstrates promising performance”; Liu, Pg. 5, Figure 3).
Regarding Claim 20, the additional elements of the dependent claim are substantially the same as the limitations of Claim 5, therefore it is rejected under the same rationale.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Lu in view of Beltagy, Tay, Chang, and Liu.
Regarding Claim 9, the additional elements of the dependent claim are substantially the same as the limitations of Claim 5, therefore it is rejected under the same rationale.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Lu in view of Beltagy, Tay, and Merle (“Effortless NLP using pre-trained Hugging Face pipelines (with just 3 lines of code!)”).
Regarding Claim 12, Lu in view of Beltagy and Tay teach the non-transitory computer-readable medium of claim 7, wherein the pre- trained machine-learning model is trained on a . . . task (Lu, pg. 988, col. 2, “Artifact check-list (meta-information)”, “Model”, “Our scripts download pre-trained checkpoints from the Hugging Face Model Hub”).
Lu in view of Beltagy and Tay do not specifically teach the task should be . . . generic . . . .
However, Merle teaches [the] . . . task [is generic] (Pg. 4, para. 2, “Pre-training should be generic”; Pg. 7, Para. 1, “the Hugging Face model hub . . . I will use in this article”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to further combine the pre-trained machine learning model trained on data from the Hugging Face Model Hub of Lu in view of Beltagy and Tay, with training on generic tasks on the Hugging Face Model Hub of Merle in order to obtain models that can be used for multiple tasks (Merle, Pg. 4, Para. 2, “Pre-training should be generic, in order to use the model for a wide range of objectives”).
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Lu in view of Beltagy, Tay, and Xu et al. (hereinafter Xu) (“Transformer Empowered CSI Feedback for Massive MIMO Systems”).
Regarding Claim 19, Lu in view of Beltagy and Tay teach the system of claim 14, wherein the operations further comprise: . . . generat[ing] an importance measure for individual tokens (Lu, Pg. 979, Col. 1, Para. 1, “the score matrix is calculated by multiplying the query matrix and key matrix, which represents the importance of each input token when producing an output element”, where “each input token” indicates the importance measure is for individual tokens); and
and providing the importance measure to a . . . function (Lu, Pg. 979, Col. 1, Para. 1, “We then normalize the score matrix with a row-wise softmax function”)
to generate the task-specific importance score (Lu, Pg. 978, Col. 1, Para. 2, “The first step of computing attention is to obtain a score matrix . . . often referred to as the attention matrix”; Lu, Pg. 979, Col. 1, Para. 1, “the score matrix . . . represents the importance of each input token when producing an output element”, which is task specific)
for the row or the column (Beltagy, Pg. 3, figure 2(d), where the “Longformer” “attention pattern” includes selected rows and columns, identified by green shading).
The reasons of obviousness have been noted in the rejection of Claim 1 above and remain applicable here.
Lu in view of Beltagy and Tay do not teach . . . providing an output from a self-attention layer to a fully-connected layer to . . . sigmoid . . . .
However, Xu teaches . . . providing an output from a self-attention layer to a fully-connected layer to . . . [generate a matrix] (Pg. 158-159, Section “III. Proposed Schemes”, Para. 2, “the input and output of the self-attention layer are added together and subsequently normalized. The normalized data is then fed into a fully-connected layer for linear transformation”)
. . . [providing the output of the transformer layer to a] sigmoid [function] (Pg. 159, Section “III. Proposed Schemes”, Para. 3, “The output of the transformer layer is scaled to [0, 1] by a sigmoid function”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the generating of matrix of importance measures for tokens, providing the matrix to a function to generating task-specific importance scores for rows and columns of Lu in view of Beltagy and Tay, with the generating a matrix by passing self-attention layer output to a fully connected layer and scaling with a sigmoid function of Xu, in order to improve model performance (Xu. Pg. 161, Para. 1, “CsiTransformer . . . achieved significantly better performance than the original CNN-based CsiNet at all compression ratios we tested. In particular, our experiment results have suggested that the proposed CsiTransformer can achieve higher recovery accuracy”, where the above scheme was for “CsiTransformer”, see Pg. 158-159, Section “III. Proposed Schemas, A. CsiTransformer”).
Response to Arguments
Applicant's arguments filed on May 5th, 2026 have been fully considered. Each argument is addressed in detail below.
I. Applicant argues the rejections to the claims, under 35 U.S.C. § 101, should be withdrawn (Applicant’s Remarks, 05/05/2026, Pg. 1-9, Section “Rejections based on 35 U.S.C. § 101”).
Applicant’s amendments to the claims, with reference to Applicant’s arguments, have overcome each and every rejection to the claims, under 35 U.S.C. § 101, as previously communicated in the 02/12/2026 Office Action.
As a result, the rejections to the claims, under 35 U.S.C. § 101, have been withdrawn.
II. Applicant argues the rejections to the claims, under 35 U.S.C. § 103, should be withdrawn (Applicant’s Remarks, 05/05/2026, Pg. 9-14, Section “Rejections based on 35 U.S.C. § 103”).
In response to Applicant’s amendments, the previously communicated rejections under 35 U.S.C. § 103, have been withdrawn. However, Applicants arguments are not persuasive in light of the new grounds for rejection, under 35 U.S.C. § 103, discussed in detail above. The new grounds of rejection rely on new combinations of the existing prior art of record and new prior art of record to teach the new combinations of elements in the amended claims, which were not presented in these arrangements in any of the previously presented claims. As a result, Applicant arguments against the previously communicated rejections under 35 U.S.C. § 103 are rendered moot.
However, in order to expedite prosecution and in the interest of clarity, any arguments still relevant to the new grounds of rejection are discussed below.
Specifically, Applicant argues “a POSITA would not have been motivated to combine Liu and Beltagy because the two references pursue different objectives and use incompatible paradigms” and “Liu expressly criticizes static sparse patterns of Beltagy”, which amounts to “a direct teaching away”, such that “there is no rational motivation to graft Liu’s dynamic, fine-grained inference-time prediction mechanism onto Beltagy’s static architecture; doing so would undermine the advantages of both and contradict Liu’s stated design principles” (Applicant’s Remark’s, Pg. 12-13, Para. 3-1).
Notably, the new grounds for rejection of the independent claims do not rely on Liu. However, dependent claims 5, 9, and 20 continue to rely on a combination that includes both Beltagy and Liu. As a result, Applicant’s assertions that a POSITA would not have been motivated to combine Liu and Beltagy or that Liu teaches away from Beltagy remain relevant. However, as discussed in detail above, a person of ordinary skill in the art would be motivated to modify Lu in a manner that includes teachings from both Liu and Beltagy in order to arrive at the subject matter of claims 5, 9, and 20 (see MPEP 2143(I)).
Additionally, Applicant’s arguments, which are directed to the withdrawn combination of Lu with Beltagy and Liu, as relied upon in the 02/12/2026 Office Action to teach the independent claims, fail to “criticize, discredit, or otherwise discourage” the combination relied upon in the rejections of claims 5, 9, and 20 (see MPEP 2145(X)(D)(1)). Thus, even if Applicant’s assertions were assumed to be true, the argument would not be persuasive. Furthermore, to the contrary of Applicant’s assertions, Liu’s operation for increasing the sparsity of an attention matrix pursues the same objective and is fully compatible with Beltagy’s output of a sparse attention pattern that includes a plurality of rows and columns (compare Beltagy, Pg. 3, Col. 2, Para. 2, “The original Transformer model has a self-attention component with O(n2) time and memory complexity . . . To address this challenge, we sparsify the full self-attention matrix according to an attention pattern specifying pairs of input locations” with Liu, Pg. 5, Col. 2, Para. 1, “Different percentage numbers indicate the sparsity ratio that we applied to the DSA models. For instance, DSA-90% means that we only keep 10% of the attention weights in each row of the attention matrix, while masking out all the other 90% of the weights”).
As a result, the argument is not persuasive.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW BRYCE GOLAN whose telephone number is (571)272-5159. The examiner can normally be reached Monday through Friday, 8:00 AM to 5:00 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW BRYCE GOLAN/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123