DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Title
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. Examiner believes that the title of the invention is imprecise. A descriptive title indicative of the invention will help in proper indexing, classifying, searching, etc. See MPEP 606.01. However, the title of the invention should be limited to 500 characters. Examiner suggests including the aspect(s) of the claims which Applicant believes to be novel or nonobvious over the prior art.
Drawings
The drawings are objected to because of the following informalities.
View numbers must be preceded by the abbreviation "FIG." See 37 CFR 1.84(u)(1)
The following figures have illegible text: 6-8, 10, 16, 17, 19-26
The following figures have text that is too small (37 CFR 1.84(p)(3) Numbers, letters, and reference characters must measure at least .32 cm. (1/8 inch) in height): 6-8, 16, 17, 19-26
It is also noted that several figures use shading. Shading in views is encouraged if it aids in understanding the invention and if it does not reduce legibility.
Additionally, some of the images are of lower quality. Applicant may desire for a clearer image should a patent be issued on this application.
Corrected drawings in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. The replacement sheet(s) should be labeled “Replacement Sheet” in the page header (as per 37 CFR 1.84(c)) so as not to obstruct any portion of the drawing figures. If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Objections
The following claims are objected to because of the following informalities:
Independent claim 14 recites A processing system while dependent claims 15-18 each recite The data processing system. Consistent nomenclature is important for the clarity of the claims.
Appropriate correction is required.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The claims contain numerous clarity issues. Applicant is encouraged to review the claims for clarity issues beyond those listed below.
Claim 1:
the associate self-attention mechanism – lacks proper antecedent basis
at least some set of information – the term some is indefinite
the execution - lacks proper antecedent basis
Claim 8:
the associate self-attention mechanism – lacks proper antecedent basis
Claim 11
the associate self-attention mechanism – lacks proper antecedent basis
at least some set of information – the term some is indefinite
the execution - lacks proper antecedent basis
the schedule - - lacks proper antecedent basis
Claim 12
the plurality of instructions - lacks proper antecedent basis
the one or more processors - lacks proper antecedent basis
the associate self-attention mechanism – lacks proper antecedent basis
on MXM – the MXM was previously introduced, and it is unclear if this is the same element, a different element, or related elements.
at least some set of information – the term some is indefinite
the execution - lacks proper antecedent basis
Claim 13
the plurality of instructions - lacks proper antecedent basis
the one or more processors - lacks proper antecedent basis
the associate self-attention mechanism – lacks proper antecedent basis
on MXM – the MXM was previously introduced, and it is unclear if this is the same element, a different element, or related elements.
at least some set of information – the term some is indefinite
the execution - lacks proper antecedent basis
Claim 14
the associate self-attention mechanism – lacks proper antecedent basis
The following abbreviations are not defined: MEM, SXM, MXM, and ALU
the SXM – lacks proper antecedent basis
the MXM - lacks proper antecedent basis
the output – lacks proper antecedent basis
the output of the ALU – this phrase is indefinite since the first softmax pass is performed by a plurality of ALUs, so it is unclear what the output of the ALU is referring to, whether it is the output of one ALU or a combined output from all the ALUs
the memory - lacks proper antecedent basis
Claim 17
the latency - lacks proper antecedent basis
Claim 18
a transformer – this term was previously introduced in claim 14 on which claim 18 depends so it is unclear if this is the same element, a different element, or related elements.
The term “LLaMA” is a trademark. Since this trademark is used in a claim as a limitation to identify or describe a particular material or product, the claim does not comply with the requirements of the 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. See MPEP 2173.05(u)
The following abbreviations are not defined: XLNet and RoBERTa
Claim 19
the associate self-attention mechanism – lacks proper antecedent basis
The following abbreviations are not defined: MEM, SXM, MXM, ALU, non-GEMM, and GEMMs
the SXM – lacks proper antecedent basis
the MXM - lacks proper antecedent basis
the output – lacks proper antecedent basis
the output of the ALU – this phrase is indefinite since the first softmax pass is performed by a plurality of ALUs, so it is unclear what the output of the ALU is referring to, whether it is the output of one ALU or a combined output from all the ALUs
the memory - lacks proper antecedent basis
batch-1 inference – It is unclear by what is meant by this term. The specification additionally does not clearly link this term to an explanation as to what it references.
the utilization - lacks proper antecedent basis
Claim 20
the layer normalization operation - lacks proper antecedent basis
the idle time - lacks proper antecedent basis
The following abbreviation is not defined: LN
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a program per se (i.e. software per se) which is non-statutory subject matter according to MPEP 2106, which states that a computer program per se is not directed to one of the four categories of statutory subject matter listed in 35 U.S.C. 101. The claimed system comprises no more than a compiler, which when given its broadest reasonable interpretation can be no more than software per se.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3-7, 9-10, 14-17, 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abts et al. (hereinafter Abts) Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads, in view of Vaswani et al. (hereinafter Vaswani) Attention Is All You Need.
Regarding Claim 1, Abts discloses a low latency processing system, comprising
a transformer having an embedding layer; and
a Tensor Streaming Processor (TSP) [“a Tensor Streaming Processor (TSP)” §I ¶3] having a Matrix Multiplication module (MXM) [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and Vector Calculation module (VXM) [“vector execution module (VXM)” §II ¶4; Fig. 1, 2], with the TSP arranged to deterministically process information [“The TSP is designed to exploit parallelism inherent in machine-learning workloads including instruction-level, memory concurrency, data and model parallelism, while guaranteeing determinism by eliminating all reactive elements in the hardware (e.g. arbiters, and caches).” Abstract] arranged by the embedding layer and an encoder layer with the associated self- attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM) mapped directly on the MXM [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and associated accumulator [“matrix execution module (MXM) consists of four (4) independent 2D MACC (multiply-accumulate) arrays” §II ¶5], and wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes [“Data parallelism for each slice’s SIMD execution is provided via a programming abstraction called parallel lanes. These parallel lanes correspond to elements of data vectors, an abstraction common to many ML frameworks like TensorFlow [2]. In the TSP model, instructions flow Northward from the ICUs to the functional slices, while data (operands and results) flow East and West between functional slices.” §I.B ¶1; Fig. 1b].
However, Abts fails to explicitly disclose a transformer having an embedding layer; and
a Tensor Streaming Processor (TSP) having a Matrix Multiplication module (MXM) and Vector Calculation module (VXM), with the TSP arranged to deterministically process information arranged by the embedding layer and an encoder layer with the associated self- attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM) mapped directly on the MXM and associated accumulator, and wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes.
Vaswani discloses a transformer [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1] having an embedding layer [“all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1]; and
a Tensor Streaming Processor (TSP) having a Matrix Multiplication module (MXM) and Vector Calculation module (VXM), with the TSP arranged to deterministically process information arranged by the embedding layer and an encoder layer with the associated self- attention mechanism, the information being further modified according to the transformer [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1] using a general matrix multiply (GEMM) [“dot-product attention is much faster and more space-efficient in practice, since it can be implemented using highly optimized matrix multiplication code” §3.2.1 ¶3; Equation 1] mapped directly on the MXM and associated accumulator, and wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes.
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the processing system of Abts to incorporate the transformer model of Vaswani.
Given the advantage of transformers which allow for significantly more parallelization and can reach a new state of the art in translation quality in shorter time due to eschewing recurrence, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 3, Abts and Vaswani disclose the low latency data processing system of claim 1.
However, Abts fails to explicitly disclose wherein the transformer is part of an encoder-based model that uses self-attention mechanisms to generate contextualized representations for input tokens.
Vaswani discloses wherein the transformer is part of an encoder-based model that uses self-attention mechanisms to generate contextualized representations for input tokens [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “we use learned embeddings to convert the input tokens and output tokens to vectors” §3.4 ¶1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the combination to incorporate encoders which are part of a transformer of Vaswani.
Given the advantage of utilizing an encoding in the transformer to encode data during the process, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 4, Abts and Vaswani disclose the low latency data processing system of claim 1.
However, Abts fails to explicitly disclose wherein the transformer is a part of an encoder-decoder model that uses both encoder and decoder components.
Vaswani discloses wherein the transformer is a part of an encoder-decoder model that uses both encoder and decoder components [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the combination to incorporate encoder-decoder which are part of a transformer of Vaswani.
Given the advantage of utilizing a encoder-decoder model in the transformer to the process data, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 5, Abts and Vaswani disclose the low latency data processing system of claim 1.
However, Abts fails to explicitly disclose wherein the transformer is a part of a decoder model that uses at least one decoder component.
Vaswani discloses wherein the transformer is a part of a decoder model that uses at least one decoder component [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the combination to incorporate decoders which are part of a transformer of Vaswani.
Given the advantage of utilizing an encoding in the transformer to decode data during the process, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 6, Abts and Vaswani disclose the low latency data processing system of claim 1.
However, Abts fails to explicitly disclose wherein the encoder layer has multiple encoders and further accepts positional information.
Vaswani discloses wherein the encoder layer has multiple encoders [“Nx” Fig. 1] and further accepts positional information [“Positional Encoding” Fig. 1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the combination to incorporate multiple encoders which are part of a transformer of Vaswani.
Given the advantage of utilizing multiple encoders in the transformer to encode data during the process and increase the data representation, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 7, Abts and Vaswani the low latency data processing system of claim 3.
However, Abts fails to explicitly disclose wherein the self-attention mechanism further comprises multi-head attention modules associated with multiple encoders.
Vaswani discloses wherein the self-attention mechanism further comprises multi-head attention modules associated with multiple encoders [“Multi-Head Attention” and “Nx” Fig. 1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the combination to incorporate multi-head attention of Vaswani.
Given the advantage of preventing narrow scope by relying on only a single attention, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 9, Abts and Vaswani disclose the low latency data processing system of claim 1. Abts further discloses wherein the TSP further comprises memory modules (MEM) and data path switching modules (SXM) [“MEM” and “SXM” Fig. 2], and wherein a vector can be read from the MEM [“Load vector at address a onto stream s” Table I], reordered on the SXM [“Rearrange or replicate data within a superlane” Table I], passed to the MXM for multiplication operation [“Activation buffer control (ABC) to initiate and coordinate arriving activations” Table I], sent to the VXM [“A vector execution module (VXM) consists of a 4×4 mesh of ALUs in each lane for point-wise arithmetic operations” §II ¶4], modified by a softmax pass [“allows for efficient parallel implementations of algorithms for batch normalization” §III.C ¶1] and results written to MEM [“on-chip memory module (MEM) is composed of 44 parallel slices of SRAM and provides the memory concurrency necessary to fully utilize the 32 streams in each direction.” §II ¶7].
Regarding Claim 10, Abts and Vaswani disclose the low latency data processing system of claim 1. Abts further discloses wherein the TSP is software scheduled [“The tensor stream processor architecture makes several deliberate tradeoffs on the hardware-software interface, pushing the complexities associated with scheduling into the compiler.” §II ¶1].
Regarding Claim 14, Abts discloses a processing system, comprising
a transformer having an embedding layer and an encoder layer with an associated self- attention block; and
a processor having a memory module [“MEM” §II ¶7], a matrix multiplication module [“MXM” §II ¶5], a data path switching module [“SXM” §II ¶6] and a vector calculation module [“VXM” §II ¶4] arranged to process information arranged by the embedding layer and the encoder layer with the associated self-attention mechanism [Fig. 1, 2, 5], the information being further modified according to the transformer using a pipeline [Fig. 1] that 1) reads a vector from MEM [“Load vector at address a onto stream s” Table I], 2) reorders the vector on the SXM [“Rearrange or replicate data within a superlane” Table I], 3) multiplies the reordered vector on the MXM [“matrix execution module (MXM) consists of four (4) independent 2D MACC (multiply-accumulate) arrays that operate on int8 or fp16 data types” §II ¶5], 4) sending the MXM result to a plurality of ALUs to perform a first softmax pass [“A vector execution module (VXM) consists of a 4×4 mesh of ALUs in each lane for point-wise arithmetic operations” §II ¶4”; “allows for efficient parallel implementations of algorithms for batch normalization” §III.C ¶1] and 5) writing the output of the ALU back to the memory [“on-chip memory module (MEM) is composed of 44 parallel slices of SRAM and provides the memory concurrency necessary to fully utilize the 32 streams in each direction.” §II ¶7].
However, Abts fails to explicitly disclose a transformer having an embedding layer and an encoder layer with an associated self- attention block.
Vaswani discloses a transformer having an embedding layer and an encoder layer with an associated self- attention block [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1; Fig. 1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify Abts to incorporate the transformer model of Vaswani.
Given the advantage of transformers which allow for significantly more parallelization and can reach a new state of the art in translation quality in shorter time due to eschewing recurrence, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 15, Abts and Vaswani discloses the data processing system of claim 14. Abts further discloses wherein the processing system provides low tail latency when the processor executes instructions of the transformer [“A sequence of instructions performed on different functional slices can be chained to create more complex actions without the need to writeback intermediate results to memory. This allows us to efficiently process streams at full bandwidth and lowest latency.” §II ¶9; Fig. 1].
Regarding Claim 16, Abts and Vaswani disclose the data processing system of claim 15. Abts further discloses wherein the processing system schedules instructions and execution of the instructions by the processor has no variation in latency [“The TSP programming model relies on two critical elements: (1) deterministic data paths in hardware, and (2) exposing temporal information about an instruction’s execution latency through the ISA, the compiler’s back-end can precisely track the position and time-of-use of any stream on-chip.” §III 1].
Regarding Claim 17, Abts and Vaswani disclose the data processing system of claim 14. Abts further discloses wherein the processing system controls scheduling of execution of processor instructions to hide the latency attributable to non-matrix multiply operations [“A sequence of instructions performed on different functional slices can be chained to create more complex actions without the need to writeback intermediate results to memory. This allows us to efficiently process streams at full bandwidth and lowest latency.” §II ¶9; “instruction informs the compiler how to schedule the operand arrival times with the instruction dispatch time in order to get them to properly intersect in time and space.” §III ¶1; “the model implementation seeks to maximize functional slice utilization, and minimize latency.” §IV ¶2; Fig. 1; Examiner Note: A matrix multiplication operation will be the most time consuming operation, so to minimize latency the latency of the non-matrix multiplication must not add to the overall latency.].
Regarding Claim 19, Abts discloses a processing system, comprising
a transformer having an embedding layer and an encoder layer with an associated self- attention block; and
a processor having a memory module [“MEM” §II ¶7], a data path switching module [“SXM” §II ¶6], a matrix multiplication module [“MXM” §II ¶5], and a vector calculation module [“VXM” §II ¶4] arranged to process information arranged by the embedding layer and the encoder layer with the associated self-attention mechanism [Fig. 1, 2, 5], the information being further modified according to the transformer using a pipeline [Fig. 1] that 1) reads a vector from MEM [“Load vector at address a onto stream s” Table I], 2) reorders the vector on the SXM [“Rearrange or replicate data within a superlane” Table I], 3) multiplies the reordered vector on the MXM [“matrix execution module (MXM) consists of four (4) independent 2D MACC (multiply-accumulate) arrays that operate on int8 or fp16 data types” §II ¶5], 4) sending the MXM result to a plurality of ALUs to perform a first softmax pass [“A vector execution module (VXM) consists of a 4×4 mesh of ALUs in each lane for point-wise arithmetic operations” §II ¶4”; “allows for efficient parallel implementations of algorithms for batch normalization” §III.C ¶1] and 5) writing the output of the ALU back to the memory [“on-chip memory module (MEM) is composed of 44 parallel slices of SRAM and provides the memory concurrency necessary to fully utilize the 32 streams in each direction.” §II ¶7] wherein the processing system enables low latency on batch-1 inference by pipelining non-GEMM operations with GEMMs to hide latency induced by executing non-GEMM operations and to hide latency and increase the utilization of the MXM [“A sequence of instructions performed on different functional slices can be chained to create more complex actions without the need to writeback intermediate results to memory. This allows us to efficiently process streams at full bandwidth and lowest latency.” §II ¶9; “instruction informs the compiler how to schedule the operand arrival times with the instruction dispatch time in order to get them to properly intersect in time and space.” §III ¶1; “the model implementation seeks to maximize functional slice utilization, and minimize latency.” §IV ¶2; Fig. 1; Examiner Note: A matrix multiplication operation will be the most time consuming operation, so to minimize latency the latency of the non-matrix multiplication must not add to the overall latency.].
However, Abts fails to explicitly disclose a transformer having an embedding layer and an encoder layer with an associated self- attention block.
Vaswani discloses a transformer having an embedding layer and an encoder layer with an associated self- attention block [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1; Fig. 1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify Abts to incorporate the transformer model of Vaswani.
Given the advantage of transformers which allow for significantly more parallelization and can reach a new state of the art in translation quality in shorter time due to eschewing recurrence, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 20, Abts and Vaswani disclose the processing system of claim 19. Abts further discloses wherein the processor optimizes the layer normalization operation to reduce the idle time during which the MXM is waiting for LN results [“This allows for efficient parallel implementations of algorithms for batch normalization quantization, or more complex activation functions like leaky ReLU activation function, for instance.” §III.C ¶1].
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abts and Vaswani in view of Polamarasetti, Enhancing CRM Accuracy Using Large Language Models (LLMs) in Salesforce Einstein GPT.
Regarding Claim 2, Abts and Vaswani disclose the low latency data processing system of claim 1.
However, Abts fails to explicitly disclose wherein the transformer is a part of a language representation model (LLM).
Polamarasetti discloses wherein the transformer is a part of a language representation model (LLM) [“Large Language Models (LLMs), especially of the transformer type, such as GPT” Abstract].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts, Vaswani, and Polamarasetti before him before the effective filing date of the claimed invention, to modify the combination to incorporate the LLMs of Polamarasetti.
Given the advantage of utilizing an effective LLM which is transformer based, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abts and Vaswani in view of Hendrycks et al. (hereinafter Hendrycks), GAUSSIAN ERROR LINEAR UNITS (GELUS).
Regarding Claim 8, Abts and Vaswani disclose the low latency data processing system of claim 7.
However, Abts fails to explicitly disclose wherein output from the associated self-attention mechanism is passed to a feed-forward layer and modified using a Gaussian Error Linear Unit (GELU) that can be mapped onto the VXM.
Vaswani discloses wherein output from the associated self-attention mechanism is passed to a feed-forward layer [“Feed Forward” Fig. 1] and modified using a Gaussian Error Linear Unit (GELU) that can be mapped onto the VXM.
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the combination to incorporate output from the associated self-attention mechanism is passed to a feed-forward layer of Vaswani.
Given the advantage of transforming the output of the attention layer, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Abts fails to explicitly disclose wherein output from the associated self-attention mechanism is passed to a feed-forward layer and modified using a Gaussian Error Linear Unit (GELU) that can be mapped onto the VXM.
Hendrycks discloses wherein output from the associated self-attention mechanism is passed to a feed-forward layer and modified using a Gaussian Error Linear Unit (GELU) that can be mapped onto the VXM [“the Gaussian Error Linear Unit (GELU), a high-performing neural network activation function.” Abstract].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts, Vaswani, and Hendrycks before him before the effective filing date of the claimed invention, to modify the combination to incorporate the GELU of Hendrycks.
Given the advantage of nonlinearity matching or exceeding models with ReLUs or ELUs, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claim(s) 11-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abts and Vaswani in view of Turley, What Is a Compiler, Anyway?
Regarding Claim 11, Abts discloses a non-transitory computer-readable storage medium comprising stored computer executable instructions, the instructions which when executed by a compiler operating on at least one computer processor to:
execute a transformer having an embedding layer and an encoder layer with an associated self-attention block; and
wherein the at least one computer processor is a Tensor Streaming Processor (TSP) [“a Tensor Streaming Processor (TSP)” §I ¶3] having a Matrix Multiplication module (MXM) [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and Vector Calculation module (VXM) [“vector execution module (VXM)” §II ¶4; Fig. 1, 2], with the TSP arranged to deterministically process information [“The TSP is designed to exploit parallelism inherent in machine-learning workloads including instruction-level, memory concurrency, data and model parallelism, while guaranteeing determinism by eliminating all reactive elements in the hardware (e.g. arbiters, and caches).” Abstract] arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM) mapped directly on the MXM [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and associated accumulator [“matrix execution module (MXM) consists of four (4) independent 2D MACC (multiply-accumulate) arrays” §II ¶5], and
wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes [“Data parallelism for each slice’s SIMD execution is provided via a programming abstraction called parallel lanes. These parallel lanes correspond to elements of data vectors, an abstraction common to many ML frameworks like TensorFlow [2]. In the TSP model, instructions flow Northward from the ICUs to the functional slices, while data (operands and results) flow East and West between functional slices.” §I.B ¶1; Fig. 1b], and wherein the instructions can be compiled into a binary for execution at the one or more processors, the binary indicating the schedule of execution of the plurality of instructions.
However, Abts fails to explicitly disclose execute a transformer having an embedding layer and an encoder layer with an associated self-attention block; and
wherein the at least one computer processor is a Tensor Streaming Processor (TSP) having a Matrix Multiplication module (MXM) and Vector Calculation module (VXM), with the TSP arranged to deterministically process information arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM) mapped directly on the MXM and associated accumulator.
Vaswani discloses execute a transformer having an embedding layer and an encoder layer with an associated self-attention block [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1; Fig. 1]; and
wherein the at least one computer processor is a Tensor Streaming Processor (TSP) having a Matrix Multiplication module (MXM) and Vector Calculation module (VXM), with the TSP arranged to deterministically process information arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1] using a general matrix multiply (GEMM) [“dot-product attention is much faster and more space-efficient in practice, since it can be implemented using highly optimized matrix multiplication code” §3.2.1 ¶3; Equation 1] mapped directly on the MXM and associated accumulator.
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify Abts to incorporate the transformer model of Vaswani.
Given the advantage of transformers which allow for significantly more parallelization and can reach a new state of the art in translation quality in shorter time due to eschewing recurrence, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Abts fails to explicitly disclose wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes, and wherein the instructions can be compiled into a binary for execution at the one or more processors, the binary indicating the schedule of execution of the plurality of instructions.
Turley discloses wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes, and wherein the instructions can be compiled into a binary for execution at the one or more processors [“A compiler is a translator. It translates the almost-English commands that we saw above into the ones and zeros that the computer understands.” pg. 2], the binary indicating the schedule of execution of the plurality of instructions [“these ones and zeros - in a different order will make the computer do different things” pg. 1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts, Vaswani, and Turley before him before the effective filing date of the claimed invention, to modify the combination to incorporate the compiling of instructions into executable data of Turley.
Given the advantage of implementing the combination on a computer which utilizes binary, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 12, Abts discloses a system, comprising:
a compiler configured to determine a schedule of execution of the plurality of instructions for execution by the one or more processors that can execute a transformer having an embedding layer and an encoder layer with an associated self-attention block [“The compiler precisely tracks the chip’s architectural state and uses that knowledge to ensure that instructions correctly intercept its stream operand(s).” §I.B ¶6; “The tensor stream processor architecture makes several deliberate tradeoffs on the hardware-software interface, pushing the complexities associated with scheduling into the compiler. Specifically, it falls on the compiler to precisely schedule instructions so as to use the hardware correctly and efficiently.” §II ¶1]; and wherein at least one computer processor is a Tensor Streaming Processor (TSP) [“a Tensor Streaming Processor (TSP)” §I ¶3] having a Matrix Multiplication module (MXM) [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and Vector Calculation module (VXM) [“vector execution module (VXM)” §II ¶4; Fig. 1, 2], with the TSP arranged to deterministically process information [“The TSP is designed to exploit parallelism inherent in machine-learning workloads including instruction-level, memory concurrency, data and model parallelism, while guaranteeing determinism by eliminating all reactive elements in the hardware (e.g. arbiters, and caches).” Abstract] arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM) mapped directly on MXM [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and associated accumulator [“matrix execution module (MXM) consists of four (4) independent 2D MACC (multiply-accumulate) arrays” §II ¶5], and wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes [“Data parallelism for each slice’s SIMD execution is provided via a programming abstraction called parallel lanes. These parallel lanes correspond to elements of data vectors, an abstraction common to many ML frameworks like TensorFlow [2]. In the TSP model, instructions flow Northward from the ICUs to the functional slices, while data (operands and results) flow East and West between functional slices.” §I.B ¶1; Fig. 1b].
However, Abts fails to explicitly disclose arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM).
Vaswani discloses arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1] using a general matrix multiply (GEMM) [“dot-product attention is much faster and more space-efficient in practice, since it can be implemented using highly optimized matrix multiplication code” §3.2.1 ¶3; Equation 1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the system of Abts to incorporate the transformer of Vaswani.
Given the advantage of transformers which allow for significantly more parallelization and can reach a new state of the art in translation quality in shorter time due to eschewing recurrence, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Abts fails to explicitly disclose wherein the compiler can compile the plurality of instructions into a binary, the binary indicating the schedule of execution of the plurality of instructions; and
allow one or more processors to be configured to execute the binary.
Turley discloses wherein the compiler can compile the plurality of instructions into a binary [“A compiler is a translator. It translates the almost-English commands that we saw above into the ones and zeros that the computer understands.” pg. 2], the binary indicating the schedule of execution of the plurality of instructions [“these ones and zeros - in a different order will make the computer do different things” pg. 1]; and
allow one or more processors to be configured to execute the binary [“run the program” pg. 2].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts, Vaswani, and Turley before him before the effective filing date of the claimed invention, to modify the combination to incorporate the compiling of instructions into executable data of Turley.
Given the advantage of implementing the combination on a computer which utilizes binary, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 13, Abts discloses a method, comprising:
providing a compiler configured to determine a schedule of execution of the plurality of instructions for execution by the one or more processors that can execute a transformer having an embedding layer and an encoder layer with an associated self-attention block [“The compiler precisely tracks the chip’s architectural state and uses that knowledge to ensure that instructions correctly intercept its stream operand(s).” §I.B ¶6; “The tensor stream processor architecture makes several deliberate tradeoffs on the hardware-software interface, pushing the complexities associated with scheduling into the compiler. Specifically, it falls on the compiler to precisely schedule instructions so as to use the hardware correctly and efficiently.” §II ¶1]; and wherein at least one computer processor is a Tensor Streaming Processor (TSP) [“a Tensor Streaming Processor (TSP)” §I ¶3] having a Matrix Multiplication module (MXM) [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and Vector Calculation module (VXM) [“vector execution module (VXM)” §II ¶4; Fig. 1, 2], with the TSP arranged to deterministically process information [“The TSP is designed to exploit parallelism inherent in machine-learning workloads including instruction-level, memory concurrency, data and model parallelism, while guaranteeing determinism by eliminating all reactive elements in the hardware (e.g. arbiters, and caches).” Abstract] arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM) mapped directly on MXM [“matrix execution module (MXM)” §II ¶5; Fig. 1, 2] and associated accumulator [“matrix execution module (MXM) consists of four (4) independent 2D MACC (multiply-accumulate) arrays” §II ¶5], and wherein at least some set of information is processed to parallelize the execution of GEMMs across all MXM planes [“Data parallelism for each slice’s SIMD execution is provided via a programming abstraction called parallel lanes. These parallel lanes correspond to elements of data vectors, an abstraction common to many ML frameworks like TensorFlow [2]. In the TSP model, instructions flow Northward from the ICUs to the functional slices, while data (operands and results) flow East and West between functional slices.” §I.B ¶1; Fig. 1b], and
wherein compiling the plurality of instructions into a binary, the binary indicating the schedule of execution of the plurality of instructions; and
allowing one or more processors to be configured to execute the binary.
However, Abts fails to explicitly disclose arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer using a general matrix multiply (GEMM).
Vaswani discloses arranged by the embedding layer and the encoder layer with the associated self-attention mechanism, the information being further modified according to the transformer [“The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder” §3 ¶1; “all sub-layers in the model, as well as the embedding layers, produce outputs” §3.1 ¶1] using a general matrix multiply (GEMM) [“dot-product attention is much faster and more space-efficient in practice, since it can be implemented using highly optimized matrix multiplication code” §3.2.1 ¶3; Equation 1].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts and Vaswani before him before the effective filing date of the claimed invention, to modify the system of Abts to incorporate the transformer of Vaswani.
Given the advantage of transformers which allow for significantly more parallelization and can reach a new state of the art in translation quality in shorter time due to eschewing recurrence, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Abts fails to explicitly disclose wherein compiling the plurality of instructions into a binary, the binary indicating the schedule of execution of the plurality of instructions; and
allowing one or more processors to be configured to execute the binary .
Turley discloses wherein compiling the plurality of instructions into a binary [“A compiler is a translator. It translates the almost-English commands that we saw above into the ones and zeros that the computer understands.” pg. 2], the binary indicating the schedule of execution of the plurality of instructions [“these ones and zeros - in a different order will make the computer do different things” pg. 1]; and
allowing one or more processors to be configured to execute the binary [“run the program” pg. 2].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts, Vaswani, and Turley before him before the effective filing date of the claimed invention, to modify the combination to incorporate the compiling of instructions into executable data of Turley.
Given the advantage of implementing the combination on a computer which utilizes binary, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abts and Vaswani in view of Devlin et al. (hereinafter Devlin), BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.
Regarding Claim 18, Abts and Vaswani disclose the data processing system of claim 17.
However, Abts fails to explicitly disclose wherein the processor instructions comprise a transformer selected from the group comprising Generative Pre-trained Transformer (GPT), GPT-2, GPT-3, GPT-4, Large Language Model Meta AI (LLaMA), Bidirectional Encoder Representations from Transformers (BERT), XLNet, or RoBERTa .
Devlin discloses wherein the processor instructions comprise a transformer selected from the group comprising Generative Pre-trained Transformer (GPT), GPT-2, GPT-3, GPT-4, Large Language Model Meta AI (LLaMA), Bidirectional Encoder Representations from Transformers (BERT), XLNet, or RoBERTa [“We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers.” Abstract].
It would have been obvious to one having ordinary skill in the art, having the teachings of Abts, Vaswani, and Devlin before him before the effective filing date of the claimed invention, to modify the combination to incorporate Bidirectional Encoder Representations from Transformers of Devlin.
Given the advantage of being conceptually simple and empirically powerful, one having ordinary skill in the art would have been motivated to make this obvious modification.
Examiner’s Note
The Examiner respectfully requests of the Applicant in preparing responses, to fully consider the entirety of the reference(s) as potentially teaching all or part of the claimed invention. It is noted, REFERENCES ARE RELEVANT AS PRIOR ART FOR ALL THEY CONTAIN. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). A reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art, including non-preferred embodiments (see MPEP 2123). The Examiner has cited particular locations in the reference(s) as applied to the claim(s) above for the convenience of the Applicant. Although the specified citations are representative of the teachings of the art and are applied to the specific limitations within the individual claim(s), typically other passages and figures will apply as well.
Additionally, any claim amendments for any reason should include remarks indicating clear support in the originally filed specification.
Conclusion
Any prior art made of record and not relied upon is considered pertinent to Applicant's disclosure. Applicant is reminded that in amending in response to a rejection of claims, the patentable novelty must be clearly shown in view of the state of the art disclosed by the references cited and the objections made. Applicant must also show how the amendments avoid such references and objections. See 37 CFR §1.111(c). Additionally when amending, in their remarks Applicant should particularly cite to the supporting paragraphs in the original disclosure for the amendments.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT H BEJCEK II whose telephone number is (571)270-3610. The examiner can normally be reached Monday - Friday: 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T. Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/R.B./ Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/ Supervisory Patent Examiner, Art Unit 2148