Prosecution Insights
Last updated: October 02, 2026
Application No. 18/358,297

32-BIT CHANNEL-ALIGNED INTEGER MULTIPLICATION VIA MULTIPLE MULTIPLIERS PER-CHANNEL

Non-Final OA §103
Filed
Jul 25, 2023
Examiner
HAKALA, ALAN GREGORY
Art Unit
Tech Center
Assignee
Intel Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
23 currently pending
Career history
20
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 11, are rejected under 35 U.S.C. 103 as being unpatentable over Siu (US 20060101243 A1) in view of Kershaw (US 20050273485 A1). Regarding claims 11, 1, A graphics processing system comprising: a memory device; a graphics processor including a memory interface coupled with the memory device and a graphics processing cluster coupled with the memory interface, (Siu ¶35 “Graphics processing subsystem 112 includes a graphics processing unit (GPU) 114 and a graphics memory 116, which may be implemented, e.g., using one or more integrated circuit devices such as programmable processors, application specific integrated circuits (ASICs), and memory devices. GPU 114 includes a rendering module 120, a memory interface module 122,”) the graphics processing cluster including a plurality of processing resources,(Siu ¶53 cited above, teaches the GPU may be implemented with multiple integrated circuit devices such as programmable processors, teaching a plurality of processing resources.) a processing resource of the plurality of processing resources including: a plurality of processing resources including functional units configured to provide an integer pipeline to execute instructions to perform operations on integer data elements ( Siu ¶47 “MMAD unit 220 implements a multiply-add (MAD) pipeline for computing A*B+C for integer or floating-point operands, and various circuits within this pipeline are leveraged to perform numerous other integer and floating-point operations.”) the integer pipeline including a plurality of parallel processor lanes associated with a plurality of data channels,(Siu ¶6 “Real-time computer animation places extreme demands on processors. To meet these demands, dedicated graphics processing units typically implement a highly parallel architecture in which a number (e.g., 16) of cores operate in parallel, with each core including multiple (e.g., 8) parallel pipelines containing functional units for performing the operations supported by the processing unit. These operations generally include various integer and floating point arithmetic operations (add, multiply, etc.), bitwise logic operations, comparison operations, format conversion operations, and so on. The pipelines are generally of identical design so that any supported instruction can be processed by any pipeline; accordingly, each pipeline requires a complete set of functional units.” Note: Siu teaches parallel cores each with parallel pipelines that contain functional units. As these functional units can perform integer operations, Siu teaches an integer pipeline with a plurality of parallel processor lanes associated with a plurality of data channels. It is implicit that different cores would have different data channels for their pipelines.) While Siu teaches a GPU with a memory interface that interacts with a plurality of parallel processor lanes that contain a first multiplier Siu does not specifically detail that each of the plurality of parallel processor lanes includes a first and second multiplier. A plurality of parallel processor lanes that each include a first and second multiplier is taught by Kershaw each of the plurality of parallel processor lanes including a first multiplier and a second multiplier. (Kershaw ¶51 “The integer multiplier (NM) unit is implemented as 2 32×16 multiply arrays. Each array is capable of performing two 16×16 operations or four 8×8 operations in a single pass. Each array can also be used to perform a 32×16 operation, allowing 32×32 operations in two passes. This means the NIM is capable of performing eight lanes of 8×8 operations or four lanes of 16×16 operations in a single pass, and two lanes of 32×32 operations in two passes.” Note: Kershaw describes its integer multiplier as two 32x16 multiplier arrays, which could be configured as two lanes to perform a 32x32 operation in two passes, or four lanes of 16x16 operations in one pass. In both configurations each lane has either a 32x32 or 16x16 multiplier performed per lane, thus teaching that there is at least a first and second multiplier in each of the lanes.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Kershaw where the integer pipeline of a GPU that comprises parallel processor lanes have a first and second multiplier in each of the plurality of parallel processor lanes. There are several reasons that would motivate one to do so, Siu and the present application are taking advantage of parallel integer pipelines, which provide a much faster way to transmit multiple pieces of data. This efficiency will be wasted if fast transported the data has to wait in a queue for a multiplier, to prevent this each lane having at least two multipliers helps avoid having not enough multipliers for the amount of data to handle. Regarding claim 2, Siu teaches: The graphics processor of claim 1, comprising a decode unit configured to decode a first instruction,(Siu ¶43 “During operation of execution core 200, fetch and dispatch unit 202 obtains instructions from an instruction store (not shown), decodes them, and dispatches them as opcodes with associated operand references or operand data to issue unit 204.” Note: Siu teaches a specific unit which will accept instructions and decode them to dispatch as opcodes.) the first instruction to configure the integer pipeline to multiply a first plurality of integer data elements to generate a first result that includes a first 32-bit data element. (Siu ¶59 “In this embodiment, MMAD unit 220 implements an eight-stage pipeline that is used for all operations. On each processor cycle, MMAD unit 220 can receive (e.g., from issue circuit 204 of FIG. 2) three new operands (A0, B0, C0) via operand input paths 402, 404, 406 and an opcode indicating the operation to be performed via opcode path 408” ¶60 “MMAD unit 220 processes each operation through all of the pipeline stages 0-7 and produces a 32-bit result value (OUT)” Note: Siu teaches that the previously described operands and opcodes which are decoded instructions are provided to the MMAD unit, which can produce a 32-but result. As the MMAD unit contains multipliers, Siu teaches that a first instruction and integers for the instructions (opcodes and operands) can be input to a multiplier to generate a 32-bit element.) Regarding claim 3, Siu teaches: The graphics processor of claim 2, the first instruction to configure the integer pipeline to load first data elements associated with a first operand and second data elements associated with a second operand, (Siu ¶11 “The input section is configured to receive first, second, and third operands and an opcode designating one of a number of supported operations to be performed and is further configured to generate control signals in response to the opcode.” ¶62 “A. MMAD Pipeline [0063] An initial understanding of the pipeline can be had with reference to how the circuit blocks of stages 0-2 are used during an FMAD operation. Stage 0 is an operand formatting stage that may optionally be implemented in issue unit 204 or in MMAD unit 220 to align and represent operands (which may have fewer than 32 bits) in a consistent manner. Stages 1-3 perform the multiplication (A*B=P) portion of the FMAD operation,”) the first data elements and the second data elements to load into a functional unit of the integer pipeline. (Siu ¶47, cited in the rejection of claim 1, clarifies that MMAD unit 220 is an integer pipeline. Thus, when Siu ¶62 above teaches that multiple “data elements”, the integers, of the operands are provided to the MMAD unit 220 it teaches providing first and second data elements to a function unit of the integer pipeline.) Claims 4-7, are rejected under 35 U.S.C. 103 as being unpatentable over Siu (US 20060101243 A1) in view of Kershaw (US 20050273485 A1) and further in view of Berkeman (US 7318080 B2) Regarding claim 4, Siu teaches: The graphics processor of claim 3, the integer pipeline, in response to the first instruction, is configured to: left shift the result by 16 bits. (Siu ¶55 “the operation on corresponding bits of operands A and B. Left shift (SHL) and right shift (SHR) operations are also supported, with operand A being used to supply the bit field to be shifted and operand B being used to specify the shift amount.” ¶256 “for 16-bit integer formats which are right-aligned in the 32-bit field, the first sixteen bits of the 32-bit field are dropped, and the next eight MSBs are extracted.” Note: Siu teaches that a left shift operation for operands is supported for any operand A can be shifted by B number of bits. As the shift amount is variable, and Siu specifically teaches the use of 32-bit and 16-bit integers, Siu teaches the ability to left shift a result by 16 bits.) While Siu teaches an integer pipeline with multipliers it does not teach how the low/high bits of each input integer can be multiplied together to produce the final result. This is found in Berkeman which teaches multiply low 16 bits of a data element of the first data elements by the low 16 bits of a corresponding data element of the second data elements to generate a first intermediate result; (Berkeman Col. 10 Line 40 “ai L, ai H denote the lower and higher W/2 bits of ai, respectively, and similar notation is used for the b variable.” Col. 7 Line 30 “Assume that the lengths of a and b are each integer multiples of a word length W” Col. 11 Line 50 “The gain from using the new split-radix multiplication method compared to the conventional technique can also be computed. Assume the word length is W bits … Example 1: 16x16 bit partial product. Reduction is 21.9% Example 2: 32x32 bit partial product. Reduction is 23.4%” Col. 8 Line 18 PNG media_image1.png 444 696 media_image1.png Greyscale Note: Berkeman two integers a and b of word length W, where W is taught to be measured in bits. Possible bit sizes used in the operations are specifically taught to be 16 and 32 bits. Thus, in the screenshot above from Berkeman when the low and high bits of integer a/b are obtained each a/bH and a/bL are 16 bits when W is 32 bits. As seen from the screenshot above, the low bits of the first element are multiplied by the low bits of the second element in the last term of the second form of the equation, shown as aL x bL.) multiply high 16 bits of the data element of the first data elements by the low 16 bits of the data element of the second data elements to generate a second intermediate result; (As seen from Berkeman Col. 8 Line 18, screenshotted above, the first data element, a, has its high 16 bits multiplied by the low bits of second data element b in the third listed term, the one placed first inside the parenthesis) multiply the low 16 bits of the data element of the first data elements by the high 16 bits of the data element of the second data elements to generate a third intermediate result;(As seen from Berkeman Col. 8 Line 18, screenshotted above, the low 16 bits of the first data element a and the high 16 bits of second data element b are multiplied together. This can be seen in the second form of the listed equation where the third overall term in the expression, and the second listed term in the parenthesis teaches this portion of the claim.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Berkeman where an operation performed on two data elements involves multiplying the combinations of their high/low halves, and left shifting the third intermediate result. There are several reasons that would motivate one to do so, performing an operation on two data elements in an integer pipeline can be inefficient, as the result of the operation may be too large to fit within the width of the pipeline. To avoid additional overhead in handling large outputs the operation could be handled piecemeal by splitting integers into their high/low halves and performing intermediary operations on smaller parts of the total data elements, like multiplying and bit shifting. While the previously described multiplication combinations of the high and low halves of two data integer elements have been shown to be taught, it has not been directly specified that the operations can be performed by a first and second multiplier. The presence of two differing multipliers which could each carryout the operations is taught by Kershaw which teaches multiply via the first multiplier, and multiply via the second multiplier (Kershaw ¶51 “The integer multiplier (NM) unit is implemented as 2 32×16 multiply arrays. Each array is capable of performing two 16×16 operations or four 8×8 operations in a single pass. Each array can also be used to perform a 32×16 operation, allowing 32×32 operations in two passes. This means the NIM is capable of performing eight lanes of 8×8 operations or four lanes of 16×16 operations in a single pass, and two lanes of 32×32 operations in two passes.” Note: Kershaw describes its integer multiplier as two 32x16 multiplier arrays, which could be configured as two lanes to perform a 32x32 operation in two passes, or four lanes of 16x16 operations in one pass. In both configurations each lane has either a 32x32 or 16x16 multiplier performed per lane, thus teaching that there is at least a first and second multiplier in each of the lanes.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Kershaw where: when a first and second data element of 32 bits are broken into high and low 16-bit halves to perform multiplication on combinations of the data elements halves, different multiplication operations can be performed by a first and second multiplier. There are several reasons that would motivate one to do so, Siu and the present application are using parallel integer pipelines to perform operations. As we are already performing a single instruction or operation by carrying out multiple intermediate operations to obtain intermediate values, this process could be made significantly more efficient by taking advantage of the existing parallel architecture. Leveraging a first and second multiplier could allow two pipelines to perform a multiplication operation in parallel, which would significantly decrease the overall time needed for the operation. Regarding claim 5, Siu teaches: The graphics processor of claim 4, the integer pipeline, in response to the first instruction, is configured to: Siu does not directly teach the summing of intermediate values, this is taught by Berkeman which teaches add the first, second, and third intermediate results to generate a sum; (Berkeman Col. 10 Line 40 “ai L, ai H denote the lower and higher W/2 bits of ai, respectively, and similar notation is used for the b variable.” Col. 8 Line 18 PNG media_image1.png 444 696 media_image1.png Greyscale Note: It was previously established in the rejection of claim 4 that Berkeman Col. 8 Line 18, cited above, teaches the multiplying of the two data elements high and low halves to obtain the three described intermediate values. As seen from the second form of the equation above, all of these intermediate values are summed together to produce a result. While Berkeman details that additional intermediate values are included in the overall sum, it nonetheless teaches summing the three intermediate values to generate a sum.) and output low 32 bits of the sum. (Berkeman Col. 17 Line 8 “The first W-bit wide number and the second W-bit wide number are used to generate a set of sub-partial products. Combinations of the sub-partial products are formed such that each combination is representable by a W-bit wide lower partial product and a carry out term that has fewer than W bits.” Note: In the above provided citation Berkeman describes the previously discussed operation in Col. 8 Line 18, where sub-partial products of two input numbers (the claims intermediate results) can have the combination of them represented by a W-bit wide lower partial output. As can be seen from the equation provided in Col. 8 Line 18, the ‘combination’ Berkeman performs is summing the products. Berkeman Col. 11 Line 50 clarifies that W can be 32 bits, thus the two input numbers that are W-bit wide may produce an output with a width larger than W-bits, but Berkeman specifically teaches it will only be the lower partial 32-bits that will be output. Thus, Berkeman teaches the low 32 bits of the sum are output.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Berkeman where the three intermediate results of an overall operation are summed together to output the low-32 bits output of the sum. There are several reasons that would motivate one to do so, it is implicit that once a plurality of intermediate values from intermediate steps are obtained, there must be some final operation that produces the output by using the intermediates in some manner. If one wishes to perform multiplication operations in this manner, which both Siu and the present invention mention, summing the products of the piecemeal operation provides a straightforward, simply way to obtain the proper final output. Regarding claim 6, Siu teaches: The graphics processor of claim 5, comprising a decode unit configured to decode a second instruction, (Siu ¶43 “During operation of execution core 200, fetch and dispatch unit 202 obtains instructions from an instruction store (not shown), decodes them, and dispatches them as opcodes with associated operand references or operand data to issue unit 204.” Note: Siu teaches the dispatch unit which decodes instructions to send them out as opcodes does this for multiple opcodes and instructions, meaning it is possible to have a first and second decoded instruction.) the second instruction to configure the integer pipeline to multiply a second plurality of integer data elements to generate an intermediate result that includes a second 32-bit data element and add an additional value to the intermediate result. (Siu ¶47 “MMAD unit 220 implements a multiply-add (MAD) pipeline for computing A*B+C for integer or floating-point operands … Operation of MMAD unit 220 is controlled by issue circuit 204, which supplies operands and opcodes to MMAD unit 220 as described above. The opcodes supplied with each set of operands by issue circuit 204 control the behavior of MMAD unit 220, selectively enabling one of its operations to be performed on that set of operands” ¶246 “In one embodiment, the target integer format can be 16 or 32 bits, signed or unsigned, with the target format being specified via the opcode.” ¶11 “The multiplication pipeline is coupled to the input section and is configurable, in response to the control signals, to compute a product of the first and second operands and to select the computed product as a first intermediate result. The test pipeline is coupled to the input section and is configurable, in response to the control signals, to perform a comparison on one or more of the first, second, and third operands and to select a result of the comparison as a second intermediate result. The addition pipeline is coupled to the multiplication section and the test pipeline and is configurable, in response to the control signals, to compute a sum of the first and second intermediate results and to select the computed sum as an operation result.” Note: Siu teaches an instruction that when decoded details a multiply and add operation for the pipeline. Siu ¶11 details the multiply and add operation as carried out by the pipeline where the multiplication pipeline produces a first intermediate result of a product from a first and second operand, which is then provided as input to an addition pipeline with a third operand (aka the claims additional value). It is known that the second intermediate result of Siu can be 32 bits as specified by the claim as Siu specifically details its integers can be in the target format of 16 or 32 bits.) Regarding claim 7, Siu teaches: The graphics processor of claim 6, wherein the additional value is a 32-bit data element. (Siu ¶180 “For integer MAD (IMAD) operations, MMAD unit 220 uses mantissa path 413 to compute A*B+C. Although some integer formats may be unsigned, MMAD unit 220 advantageously treats all formats as being signed 32-bit twos-complement representations; this inherently produces the correct results regardless of actual format. [0181] In stage 0, the operands A, B, and C are extended to 32 bits if necessary using blocks 504-506 (FIG. 5) for 8-bit input formats or 508-510 (for 16-bit formats).” Note: The claim states that the additional value which is added to the product in the multiply-add operation can be a 32-bit data element. Siu teaches that in some cases, if necessary, the additional value C can be 32 bits.) Claims 9, 17, are rejected under 35 U.S.C. 103 as being unpatentable over Siu (US 20060101243 A1) in view of Kershaw (US 20050273485 A1), and further in view of Boutaund (US 6263418 B1). Regarding claim 9, Siu teaches: The graphics processor of claim 1, While Siu details the sizes of its multipliers, it does not specifically detail a set up where the first and second multiplier are of 32x16 and 16x16 bits respectively. While not teaching this exact configuration, Kershaw does teach a first and second multiplier where the first is a 32x16 bit multiplier, and the second performs a 16x16 but multiplication, specifically Kershaw teaches wherein the first multiplier is a 32-bit x 16-bit multiplier and the second multiplier acts as a 16-bit x 16-bit multiplier. (Kershaw ¶51 “The integer multiplier (NM) unit is implemented as 2 32×16 multiply arrays. Each array is capable of performing two 16×16 operations or four 8×8 operations in a single pass. Each array can also be used to perform a 32×16 operation, allowing 32×32 operations in two passes. This means the NIM is capable of performing eight lanes of 8×8 operations or four lanes of 16×16 operations in a single pass, and two lanes of 32×32 operations in two passes.” Note: Kershaw teaches two multipliers that are both 32x16 multipliers. Kershaw specifically taches that these multipliers can be configured to act as smaller multipliers, and perform operations like 16x16 operations. Thus, Kershaw teaches a first and second multiplier where the first is a 32x16 bit multiplier and the second can act as a 16x16 bit multiplier, but is really another 32x16 bit multiplier.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Kershaw where the first multiplier is a 32x16 bit and the second performs a 16x16 bit multiplication. There are several reasons that would motivate one to do so, Siu and the present invention will often have to do multiplication on the GPU integer pipeline with 32-bit integers. Directly handling this multiplication with a 32x32 multiplier would require wider pipelines, leveraging a 32x16 and 16x16 multiplier allow for the low 32 bits of a 32x32 to be produced without needing a wider pipeline. Leveraging a true 16x16 bit multiplier instead of a larger one that performs the smaller operation is taught by Boutaund which teaches a 16-bit x 16-bit multiplier. (Boutaund Col. 17 Line 1 “The 16×16-bit hardware multiplier 27 computes a signed or unsigned 32-bit product in a single machine cycle. All multiply instructions, except MPYU (multiply unsigned) instruction perform a signed multiply operation in the multiplier”) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Boutaund where a parallel integer pipeline that has a first and second multiplier where the first is a 32x16 multiplier and the second performs a 16x16 bit multiplier uses a 16x16 bit multiplier as its second multiplier. There are several reasons that would motivate one to do so, depending on the size of the operands a multiplication operation may result in a product too large to fit within the integer pipeline’s desired width, by using a smaller multiplier for the second multiplier we can guarantee that the product will not be above a certain width. Regarding claim 17, Siu teaches: A method comprising: decoding a first instruction or a second instruction at a graphics processor; (Siu ¶43 “During operation of execution core 200, fetch and dispatch unit 202 obtains instructions from an instruction store (not shown), decodes them, and dispatches them as opcodes with associated operand references or operand data to issue unit 204.” Note: Siu teaches the dispatch unit which decodes instructions to send them out as opcodes does this for multiple opcodes and instructions, meaning it is possible to have a first and second decoded instruction.) and executing the first instruction or the second instruction at an integer pipeline of the graphics processor, wherein executing the first instruction includes multiplying a first plurality of integer data elements (Siu ¶6 “These operations generally include various integer and floating point arithmetic operations (add, multiply, etc.), bitwise logic operations, comparison operations, format conversion operations, and so on.” Note: Siu previously details how multiple instructions can be fetched, decoded, and dispatched as opcodes with operands to perform the instructions. Here, Siu specifically teaches these operations include multiplication.) to generate a first result that includes a first 32-bit data element, (Siu ¶51 “Integer formats are specified herein by an initial “s” or “u” indicating whether the format is signed or unsigned and a number denoting the total number of bits (e.g., 8, 16, 32); thus, s32 refers to signed 32-bit integers” Note: Siu has already been shown to teach multiplication operations, here it is specified that there are a variety of possible integer formats, including 16 and 32 bit. Two 16-bit integers will always produce a 32-bit result, and two 32-bit integers can produce a 64-bit result, but often only the low 32 bits are needed to represent the result. Thus, Siu teaches that the result of a multiplication step can be a 32-bit data element.) wherein executing the second instruction includes multiplying the first plurality of integer data elements to generate an intermediate result that includes a second 32-bit data element (Siu ¶59 “In this embodiment, MMAD unit 220 implements an eight-stage pipeline that is used for all operations. On each processor cycle, MMAD unit 220 can receive (e.g., from issue circuit 204 of FIG. 2) three new operands (A0, B0, C0) via operand input paths 402, 404, 406 and an opcode indicating the operation to be performed via opcode path 408” ¶60 “MMAD unit 220 processes each operation through all of the pipeline stages 0-7 and produces a 32-bit result value (OUT)” ¶63 “Stage 0 is an operand formatting stage that may optionally be implemented in issue unit 204 or in MMAD unit 220 to align and represent operands (which may have fewer than 32 bits) in a consistent manner. Stages 1-3 perform the multiplication (A*B=P) portion of the FMAD operation, while stages 4-6 perform the addition (P+C) portion.” Note: Siu teaches that in response to an instruction, an MMAD unit will perform a multiplying of a first plurality of integer data elements to generate an intermediate result. This is taught by Siu ¶63 which teaches that a first stage of steps produces P from multiplication, and then produces the final result by adding value C to P, making P an intermediate value. Siu teaches that these steps 0-7 will produce a 32-bit result value, as the output is simply the intermediate value summed with C this teaches that P can also be 32 bits.) and adding the intermediate result to an additional value (As taught by Siu ¶63 above, the intermediate result of multiplication “P” is added to additional value “C” to obtain the final result.) While Siu details the sizes of its multipliers, it does not specifically detail a set up where the first and second multiplier are of 32x16 and 16x16 bits respectively. While not teaching this exact configuration, Kershaw does teach a first and second multiplier where the first is a 32x16 bit multiplier, and the second performs a 16x16 but multiplication, specifically Kershaw teaches , and wherein multiplying the first plurality of integer data elements and multiplying the second plurality of integer data elements includes performing a first multiply via a first multiplier and performing a second multiply via a second multiplier,(Kershaw ¶51, cited below, teaches a first and second multiplier) wherein the first multiplier is a 32-bit x 16-bit multiplier and the second multiplier acts as a 16-bit x 16-bit multiplier. (Kershaw ¶51 “The integer multiplier (NM) unit is implemented as 2 32×16 multiply arrays. Each array is capable of performing two 16×16 operations or four 8×8 operations in a single pass. Each array can also be used to perform a 32×16 operation, allowing 32×32 operations in two passes. This means the NIM is capable of performing eight lanes of 8×8 operations or four lanes of 16×16 operations in a single pass, and two lanes of 32×32 operations in two passes.” Note: Kershaw teaches two multipliers that are both 32x16 multipliers. Kershaw specifically taches that these multipliers can be configured to act as smaller multipliers, and perform operations like 16x16 operations. Thus, Kershaw teaches a first and second multiplier where the first is a 32x16 bit multiplier and the second can act as a 16x16 bit multiplier, but is really another 32x16 bit multiplier.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Kershaw where multiplying in a GPU integer pipeline can be performed by a first and second multiplier, where the first multiplier is 32x16 and the second performs a 16x16 multiplication. There are several reasons that would motivate one to do so, Siu and the present invention will often have to do multiplication on the GPU integer pipeline with 32-bit integers. Directly handling this multiplication with a 32x32 multiplier would require wider pipelines, leveraging a 32x16 and 16x16 multiplier allow for the low 32 bits of a 32x32 to be produced without needing a wider pipeline. Similarly, by leveraging two multipliers we can better take advantage of the integer pipeline’s parallel architecture. Leveraging a true 16x16 bit multiplier instead of a larger one that performs the smaller operation is not in Siu, and is found in by Boutaund which teaches a 16-bit x 16-bit multiplier. (Boutaund Col. 17 Line 1 “The 16×16-bit hardware multiplier 27 computes a signed or unsigned 32-bit product in a single machine cycle. All multiply instructions, except MPYU (multiply unsigned) instruction perform a signed multiply operation in the multiplier”) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Boutaund where a parallel integer pipeline has a first and second multiplier where the first is a 32x16 multiplier and the second performs a 16x16 bit multiplier uses a 16x16 bit multiplier as its second multiplier. There are several reasons that would motivate one to do so, depending on the size of input operands a multiplication operation may result in a product too large to fit within the integer pipeline’s desired width, by using a smaller multiplier for the second multiplier we can guarantee that the product will not be above a certain width. Claims 8, 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Siu (US 20060101243 A1) in view of Kershaw (US 20050273485 A1), further in view of Berkeman (US 7318080 B2), and further in view of Boutaund (US 6263418 B1). Regarding claim 8, Siu teaches: The graphics processor of claim 6, wherein the additional value is a 16-bit value. (Siu ¶180 cited above teaches that sometimes it is necessary to make the additional value 32 bits, but not always. If this is not done, Siu teaches the accepted normal inputs are 8-bits and 16-bits for A, B, and C, where C is the additional value. Thus, Siu teaches the additional value can be directly input as a 16-bit value.) While Siu teaches the additional value can be 16-bit it does not specifically detail that it is an “immediate value”, a value embedded in the instruction it is used in. The use of a 16-bit immediate value is taught by Boutaund which teaches a 16-bit immediate value (Boutaund Col. 17 Line 31 “An LT (load TREGO) instruction normally loads the TREGO 49 to provide one operand (from the data bus), and the MPY (multiply) instruction provides the second operand (also from the data bus). A multiplication can also be performed with an immediate operand using the MPYK instruction” Col. 17 Line 1 “The 16×16-bit hardware multiplier 27 computes a signed or unsigned 32-bit product in a single machine cycle. All multiply instructions, except MPYU (multiply unsigned) instruction perform a signed multiply operation in the multiplier” Note: Boutaund teaches that an immediate value/operand can be leveraged. It is known that this value is a 16-bit immediate value, as Boutand teaches the multiplier which accepts the immediate value is a 16x16 bit multiplier.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Boutaund where the additional value added to a product is not just a 16-bit value, but a 16-bit immediate value specifically. There are several reasons that would motivate one to do so, Siu and the present application take efforts to make their integer pipelines more efficient, the system could be made even more efficient by reducing memory transit time by making the additional value an immediate value included in the instruction. Regarding claim 18, Siu teaches: The method of claim 17, wherein executing the first instruction includes: loading first data elements associated with a first operand and second data elements associated with a second operand into a functional unit of the integer pipeline of a graphics processor; As this first portion of the claim is the exact same content as the content of claim 3, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. multiplying low 16 bits of a data element of the first data elements by the low 16 bits of a corresponding data element of the second data element to generate a first intermediate result; multiplying high 16 bits of the data element of the first data elements by the low 16 bits of the data element of the second data elements to generate a second intermediate result; multiplying the low 16 bits of the data element of the first data elements by the high 16 bits of the data element of the second data elements cto generate a third intermediate result As this second portion of the claim is the exact same content as the content of claim 4, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. adding the first, second, and third intermediate results to generate a sum; and outputting low 32-bits of the sum As this first half of the claim is the exact same content as the content of claim 5, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. Regarding claim 19, Siu teaches: The method of claim 17, wherein executing the second instruction includes: loading first data elements associated with a first operand, second data elements associated with a second operand, and third data elements associated with a third operand (Siu ¶11 and ¶62, cited previously in the rejection of claim 3, teaches accepting three operands and three data elements into a functional unit of the pipeline. Below, the claim states the third operand will be used to generate a sum.) into a functional unit of the integer pipeline; (Siu ¶47 and ¶62, cited previously, teach the loading of data elements associated with their respective operands into the functional unit of the integer pipeline. ) the third operand to generate a sum (Siu ¶47 and ¶62 specifically teach three operands can be accepted, and names an example operating which involves the reading in of the operands to the functional unit of the pipeline as an MMAD operation, which would make one of the operands an addition operand to perform the addition part of the multiply-add operation. Thus, Siu teaches that there can be a third operand used to generate a sum, or in other words that the third operand is addition.) Siu does not however detail handling the multiplying combinations of low 16 bits and high 16 bits of the two data elements together to produce intermediate results. Similarly, adding the intermediate results to generate a sum is not taught by Siu. This entire process as detailed in the claims is taught by Berkeman which teaches multiplying low 16 bits of the first data elements by the low 16 bits of the second data elements to generate a first intermediate result; (Berkeman Col. 10 Line 40, Col. 7 Line 30, Col. 11 Line 50, and Col. 8 Line 18, cited above in the rejection of claim 4, teach this. Particularly Col. 8 Line 18 teaches this entire portion of the claim that specifies how the high/low 16 bits of the first/second data elements will be multiplied together to produce the three intermediate results.) multiplying high 16 bits of the first data elements by the low 16 bits of the second data elements to generate a second intermediate result; (The second form of the equation provided in Berkeman Col. 8 Line 18 teaches this specific multiplication step, the high 16 bits of the first multiplied by the low 16 bits of the second, to obtain an intermediate result.) multiplying the low 16 bits of the first data elements by the high 16 bits of the second data elements to generate a third intermediate result and left shift the third intermediate result by 16 bits; ;(As stated previously, all steps in calculating the three intermediate results as specified by the claims are all taught by the Berkeman citations in the rejection of claim 4 referenced previously, particularly Col. 8 Line 18 which details this specific intermediate value of multiplying the low bits of the first element by the high bits of the second.) (Berkeman Col. 10 Line 40, Col. 8 Line 18, cited previously in the rejection of claim 5, teaches that the first, second, and third intermediate results are summed together to produce a sum. As seen from the equation present there are at least 3 operands present, where the ‘third’ operand which produces the sum is the addition operand seen in the equation.) and output low 32 bits of the sum. (Berkeman Col. 17 Line 8, cited previously in the rejection of claim 5, teaches to output the low 32 bits of the sum.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Berkeman where three intermediate results of an overall operation are summed together to output the low-32 bits output of the sum, and the third intermediate result is left shifted by 16 bits. There are several reasons that would motivate one to do so, it is implicit that once a plurality of intermediate values from intermediate steps are obtained, there must be some final operation that produces the output by using the intermediates in some manner. If one wishes to perform multiplication operations in this manner, which both Siu and the present invention mention, summing the products of the piecemeal operation provides a straightforward, simply way to obtain the proper final output. While Berkeman does teach the claims calculation of the three intermediate results in the same manner described by the claims, and similarly generates the sum in the same way as well, it does not directly specify that the different intermediate results multiplication operations will be handled by a first and second multiplier. This can be found in Kershaw which teaches multiply via the first multiplier, and multiply via the second multiplier (Kershaw ¶51, cited previously in the rejection of claim 4, teaches the leveraging of a first and second multiplier) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Kershaw where: a first and second data element of 32 bits are broken into high and low 16-bit halves to perform multiplication on three combinations of the first and second data element’s halves can be done by two different multipliers. There are several reasons that would motivate one to do so, Siu and the present application are using parallel integer pipelines to perform operations. As we are already performing a single instruction or operation by carrying out multiple intermediate operations to obtain intermediate values, this process could be made significantly more efficient by taking advantage of the existing parallel architecture. Leveraging a first and second multiplier could allow two pipelines to perform a multiplication operation in parallel, which would significantly decrease the overall time needed for the operation. Regarding claim 20, Siu teaches: The method of claim 19, further comprising summing the first, second, and third intermediate results with the third operand via a 48-bit adder. (Siu ¶97 “accordingly, IP adder 804 may be implemented as a 48-bit adder, and path RP may be 49 bits wide to accommodate carries” Note: The summing of the first, second, and third intermediate results has been previously established; here Siu specifically teaches that a 48-bit adder can be used for operands. Thus, it is possible to perform the summing via the 48-bit adder.) Claims 10, 12, are rejected under 35 U.S.C. 103 as being unpatentable over Siu (US 20060101243 A1) in view of Kershaw (US 20050273485 A1) and further in view of Liu (US 7834881 B2) Regarding claims 12, 10, Siu teaches: The graphics processing system of claim 11, the processing resource of the plurality of processing resources including: a register file (Siu ¶43 “For each instruction, issue unit 204 obtains any referenced operands, e.g., from register file 224.”) Siu does not however detail a plurality of reuse buffers coupled to the register files which for each processor, this is found in Liu which teaches: each of the plurality of processing resources including a plurality of register files and a plurality of reuse buffers coupled with the plurality of register files, (Liu Col. 3 Line 50 “Graphics memory may include portions of Host Memory 112, Local Memory 140, register files coupled to the components within Graphics Processor 105, and the like.” Col. 6 Line 58 “Register File Unit 250 includes two or more memory banks, Banks 320 that are configured to simulate a single multi-ported memory. Each Bank 320 includes several locations which function as registers that are configured to store operands. Each Collector Unit 330 receives the requests and the corresponding program instruction from Register Address Unit 240 and determines if the program instruction is an instruction for execution by the particular Execution Unit 365 coupled to the Collector Unit 330. If the program instruction is an instruction for execution by the particular Execution Unit 365 coupled to the Collector Unit 330, the Collector Unit 330 accepts the program instruction and requests for processing. In some embodiments of the present invention, each Execution Unit 365 is identical and a priority scheme is used to determine which Execution Unit 365 will execute the program instruction.” PNG media_image2.png 1221 804 media_image2.png Greyscale Note: Liu Fig. 3 and the accompanying description provided by Col. 6 Line 58 teach a plurality of register files. Liu teaches that Register File Unit 250 has multiple memory banks, and teaches that there are multiple register files. Each Bank 320 used my Liu can store multiple register files in its multiple registers, and each bank is attached to its own collector unit as seen in Fig. 3. The Collector Units accept the register files, which store operands, and determines if they should be then output to the Execution Unit. The Collector Unit which accepts, stores, and outputs register files to the Execution Unit is analogous to the claims reuse buffer. As there is a collector unit for each memory bank, Liu teaches a plurality of reuse buffers for a plurality of register files. The Execution Unit is taught to accept the operands and info provided from the Collector Unit, the reuse buffer, and execute the instruction with the operands. As seen from Fig. 3 there are multiple execution units to execute the operations, each has its own reuse buffer/Collector Unit with its own memory bank, teaching a plurality of processing resources each with a plurality of register files and reuse buffer.) the plurality of reuse buffers coupled with the integer pipeline,( PNG media_image3.png 1291 871 media_image3.png Greyscale Col. 14 Line 11“Execution Unit A 765 and Execution Unit B 775 are configured to process sixteen samples in parallel in a single-instruction multiple-data (SIMD) manner. Multiple threads are executed as a thread group to process N samples in parallel, where N may be 16, 32, or another integer value” Note: Liu Fig. 2 is provided above, and is the first half of the complete pipeline, where the second half is shown above in Fig. 3. This clarifies that the displayed series of units and operations is a graphics processing pipeline. Liu Col. 14 Line 11 further clarifies that the input values are integers, making the pipeline all described components are coupled in, including the reuse buffer, an integer pipeline.) the plurality of reuse buffers configured to cache operand data read from the plurality of register files from the integer pipeline. (Liu Col. 6 Line 58, cited above, specifically teaches the operand data from the register files will be read into the Collector Unit. As seen from Liu Fig. 3 the Collector Unit is closer to the processors, the Execution Units, in the pipeline and handles distributing the operands to the different execution units as it is the closest unit to them. This process of reading the data out of memory to be stored closer to the processor for more immediate use is the caching of the operand data from the register files.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Liu where an integer pipeline with a plurality of processors and register files uses a plurality of reuse buffers in the pipeline coupled to the register files and processors to cache the operand data from the register files. There are several reasons that would motivate one to do so, when performing many basic integer operations like Siu and the present application are it is likely that many operations share the same operand. As both Siu and Liu take efforts to make their pipelines efficient, they could be made more efficient by leveraging reuse buffers to cache operands for the processors to avoid having to re obtain the same operand. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Siu (US 20060101243 A1) in view of Kershaw (US 20050273485 A1), further in view of Boutaund (US 6263418 B1), and further in view of Liu (US 7834881 B2) Regarding claim 13, Siu teaches: The graphics processing system of claim 12, wherein the first multiplier is a 32-bit x 16-bit multiplier and the second multiplier is a 16-bit x 16-bit multiplier. As the content of claim 13 is identical to the content of 9, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. Claims 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Siu (US 20060101243 A1) in view of Kershaw (US 20050273485 A1), further in view of Berkeman (US 7318080 B2), further in view of Boutaund (US 6263418 B1), and further in view of Liu (US 7834881 B2). Regarding claim 14, Siu teaches: The graphics processing system of claim 13, comprising a decode unit configured to decode a first instruction and a second instruction, the first instruction to cause the integer pipeline to multiply a first plurality of integer data elements to generate a first result that includes a first 32-bit data element, As this first half of the claim is identical to claim 2, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. and a second instruction to cause the integer pipeline to multiply a second plurality of integer data elements to generate an intermediate result that includes a second 32-bit data element and add an additional value to the intermediate result. As this remaining half of the claim is identical to claim 6, other than the preamble which has already been rejected, it is rejected under the same rationale. Regarding claims 15, Siu teaches: The graphics processing system of claim 14, the integer pipeline, in response to the first instruction, is configured to: load first data elements associated with a first operand and second data elements associated with a second operand into a functional unit of the integer pipeline; As this first portion of the claim is the exact same content as the content of claim 3, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. multiply low 16 bits of a data element of the first data elements by the low 16 bits of a corresponding data element of the second data elements via the first multiplier to generate a first intermediate result; multiply high 16 bits of the data element of the first data elements by the low 16 bits of the data element of the second data elements via the first multiplier to generate a second intermediate result; multiply the low 16 bits of the data element of the first data elements by the high 16 bits of the data element of the second data elements via the second multiplier to generate a third intermediate result and left shift the third intermediate result by 16 bits; As this second portion of the claim is the exact same content as the content of claim 4, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. add the first, second, and third intermediate results to generate a sum; and output low 32 bits of the sum. As this first half of the claim is the exact same content as the content of claim 5, other than the preamble which has already been shown to be rejected, it is rejected under the same rationale. Regarding claim 16, Siu teaches: The graphics processing system of claim 14, the integer pipeline, in response to the second instruction, is configured to: load first data elements associated with a first operand, second data elements associated with a second operand, and third data elements associated with a third operand (Siu ¶11 and ¶62, cited previously in the rejection of claim 3, teaches accepting three operands and three data elements into a functional unit of the pipeline. Below, the claim states the third operand will be used to generate a sum.) into a functional unit of the integer pipeline; (Siu ¶47 and ¶62, cited previously, teach the loading of data elements associated with their respective operands into the functional unit of the integer pipeline. ) the third operand to generate a sum (Siu ¶47 and ¶62 specifically teach three operands can be accepted, and names an example operating which involves the reading in of the operands to the functional unit of the pipeline as an MMAD operation, which would make one of the operands an addition operand to perform the addition part of the multiply-add operation. Thus, Siu teaches that there can be a third operand used to generate a sum, or in other words that the third operand is addition.) and left shift the result by 16 bits. (Siu ¶55 “the operation on corresponding bits of operands A and B. Left shift (SHL) and right shift (SHR) operations are also supported, with operand A being used to supply the bit field to be shifted and operand B being used to specify the shift amount.” ¶256 “for 16-bit integer formats which are right-aligned in the 32-bit field, the first sixteen bits of the 32-bit field are dropped, and the next eight MSBs are extracted.” Note: Siu teaches that a left shift operation for operands is supported for any operand A can be shifted by B number of bits. As the shift amount is variable, and Siu specifically teaches the use of 32-bit and 16-bit integers, Siu teaches the ability to left shift a result by 16 bits.) Siu does not however detail handling the multiplying combinations of low 16 bits and high 16 bits of the two data elements together to produce intermediate results. Similarly, adding the intermediate results to generate a sum is not taught by Siu. This entire process as detailed in the claims is taught by Berkeman which teaches multiply low 16 bits of a data element of the first data elements by the low 16 bits of a corresponding data element of the second data elements to generate a first intermediate result (Berkeman Col. 10 Line 40, Col. 7 Line 30, Col. 11 Line 50, and Col. 8 Line 18, cited above in the rejection of claim 4, teach this. Particularly Col. 8 Line 18 provides an equation that teaches this entire claim portion, specifically it does state: the low 16 bits of the first and second data elements are multiplied.) multiply high 16 bits of the data element of the first data elements by the low 16 bits of the data element of the second data elements to generate a second intermediate result;(The second form of the equation provided in Berkeman Col. 8 Line 18 teaches this specific multiplication step, the high 16 bits of the first multiplied by the low 16 bits of the second, to obtain an intermediate result.) multiply the low 16 bits of the data element of the first data elements by the high 16 bits of the data element of the second data elements to generate a third intermediate result (As stated previously, all steps in calculating the three intermediate results as specified by the claims are all taught by the Berkeman citations in the rejection of claim 4 referenced previously, particularly Col. 8 Line 18 which details this specific intermediate value of multiplying the low bits of the first element by the high bits of the second.) add the first, second, and third intermediate results with a third operand to generate a sum; (Berkeman Col. 10 Line 40, Col. 8 Line 18, cited previously in the rejection of claim 5, teaches that the first, second, and third intermediate results are summed together to produce a sum. As seen from the equation present there are at least 3 operands present, where the ‘third’ operand which produces the sum is the addition operand seen in the equation.) and output low 32 bits of the sum.(Berkeman Col. 17 Line 8, cited previously in the rejection of claim 5, teaches to output the low 32 bits of the sum.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Berkeman where the three intermediate results of an overall operation are summed together to output the low-32 bits output of the sum and the third intermediate result is left shifted by 16 bits. There are several reasons that would motivate one to do so, it is implicit that once a plurality of intermediate values from intermediate steps are obtained, there must be some final operation that produces the output by using the intermediates in some manner. If one wishes to perform multiplication operations in this manner, which both Siu and the present invention mention, summing the products of the piecemeal operation provides a straightforward, simply way to obtain the proper final output. While Berkeman does teach the claims calculation of the three intermediate results in the same manner described by the claims, and similarly generates the sum in the same way as well, it does not directly specify that the different intermediate results multiplication operations will be handled by a first and second multiplier. This can be found in Kershaw which teaches multiply via the first multiplier, and multiply via the second multiplier (Kershaw ¶51, cited previously in the rejection of claim 4, teaches the leveraging of a first and second multiplier) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Siu with Kershaw where: a first and second data element of 32 bits are broken into high and low 16-bit halves to perform multiplication on three combinations of the first and second data element’s halves can be done by two different multipliers. There are several reasons that would motivate one to do so, Siu and the present application are using parallel integer pipelines to perform operations. As we are already performing a single instruction or operation by carrying out multiple intermediate operations to obtain intermediate values, this process could be made significantly more efficient by taking advantage of the existing parallel architecture. Leveraging a first and second multiplier could allow two pipelines to perform a multiplication operation in parallel, which would significantly decrease the overall time needed for the operation. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALAN GREGORY HAKALA whose telephone number is (571)272-7863. The examiner can normally be reached 8:00am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at (571) 270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALAN GREGORY HAKALA/ Examiner, Art Unit 2617 /KING Y POON/Supervisory Patent Examiner, Art Unit 2617
Read full office action

Prosecution Timeline

Jul 25, 2023
Application Filed
Aug 31, 2023
Response after Non-Final Action
Sep 24, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month