DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 5, 7-9, 11-14, 16-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Abts et al. (US 12,340,300 B1), hereinafter Abts.
Regarding claim 5, Abts discloses
A Memory [MEM, See figure 19]
A Vector Processing Tile [VXM] that has a plurality of vector compute Lanes (i.e. Vector Compute Channels), each lane including a plurality of arithmetic logic unit (ALU) circuits coupled in series and wherein the plurality of lanes is operable to generate outputs in parallel [“VXM contains 16 [ALUs] per lane”, “supports chaining together two or more Vector ALUs within each lane”, and “efficient parallel implementations of algorithms for batch normalization…” see Col.12, The Vector Processing Tiles; “Each stream automatically progresses in its designated direction on every cycle” col.10-11, Streams];
And a transpose circuit, and wherein the transpose circuit is operable to receive an input tensor, transpose the input tensor, and output a transposed tensor [“The switch units [SXM”/”NET] execute functions for the transposition, permutation, shifting and rotation of data elements… these operations are used for performing tensor reshape operations common to machine learning algorithms… the SXM can rotate or transpose a stream of data across the lanes.” Col.13, The Switching Processing Tiles];
While Abts does not explicitly disclose the plurality of vector compute channels coupled to the transpose circuit, and output a transposed tensor to the plurality of vector compute channels, and wherein the transpose circuit is communicatively coupled between the memory and the plurality of vector compute channels,
It would have been obvious to one of ordinary skill in the art, to look within the specification of Abts, and take note of an embodiment wherein “additional slices of VXMs (not shown) may be allocated to a set. The additional VXM slices may be physically or logically located next to one of the MXM slices” in Col.9, ll.40-44. Therefore, it would have been obvious to one of ordinary skill in the art that given the normal structure of MXM-SXM-MEM-VXM (i.e. figure 15/14) of Abts, to implement an embodiment of MXM-VXM-SXM-MEM-VXM of Abts, this allows for more independent functional modules of the system see Col.9, ll.34-44. This embodiment of Abts discloses wherein the transpose circuit (SXM) is communicatively coupled between the memory (MEM) and the plurality of vector compute channels (VXM), which in turn also discloses the plurality of vector compute channels (Lanes of the VXM) coupled to the transpose circuit (SXM), and output a transposed tensor to the plurality of compute channels.
Regarding claim 7, Abts disclose the invention substantially as claimed. See the discussion of claim 5 above.
Abts discloses wherein the transpose circuit is operable to transpose the input tensor by receiving input elements corresponding to a column of the input tensor in parallel, and providing the input elements in series as a vector of the transposed tensor to a vector compute channel [“SXM can be used to arbitrarily remap the 16 lanes within each Superlane… A transpose operation takes 16 incoming streams and produces 16 output streams with the rows and columns exchanged” Col.13, The Switching Processing Tiles].
Regarding claim 8, Abts disclose the invention substantially as claimed. See the discussion of claim 5 above.
Abts discloses wherein the output of a vector compute channel is an output vector generated by applying an elementwise operation to each element of the vector of the tensor inputted into the vector compute channel [“The VXM also performs common normalization functions… supports chaining… vector ALUs… for batch normalization, quantization, or more complex activation functions” Col.12, The Vector Processing Tiles, discloses various element-wise operations; See Col.10-11, Stream, talks about vectors in the lanes].
While Abts does not explicitly disclose an elementwise operation to each element of the vector of the transposed tensor inputted into the vector compute channel.
It would have been obvious to one of ordinary skill in the art, to look within the specification of Abts, and take note of an embodiment wherein “additional slices of VXMs (not shown) may be allocated to a set. The additional VXM slices may be physically or logically located next to one of the MXM slices” in Col.9, ll.40-44. Therefore, it would have been obvious to one of ordinary skill in the art that given the normal structure of MXM-SXM-MEM-VXM (i.e. figure 15/14) of Abts, to implement an embodiment of MXM-VXM-SXM-MEM-VXM of Abts, this allows for more independent functional modules of the system see Col.9, ll.34-44. This embodiment of Abts discloses an elementwise operation to each element of the vector of the transposed tensor inputted into the vector compute channel.
Regarding claim 9, Abts disclose the invention substantially as claimed. See the discussion of claim 5 above.
Abts discloses wherein the output of a vector compute channel includes an output value generated by performing a computation on elements of the vector of the tensor inputted into the vector compute channel [“The VXM also performs common normalization functions… supports chaining… vector ALUs… for batch normalization, quantization, or more complex activation functions” Col.12, The Vector Processing Tiles, discloses various element-wise operations; See Col.10-11, Stream, talks about vectors in the lanes].
While Abts does not explicitly disclose performing a computation on elements of the vector of the transposed tensor inputted into the vector compute channel.
It would have been obvious to one of ordinary skill in the art, to look within the specification of Abts, and take note of an embodiment wherein “additional slices of VXMs (not shown) may be allocated to a set. The additional VXM slices may be physically or logically located next to one of the MXM slices” in Col.9, ll.40-44. Therefore, it would have been obvious to one of ordinary skill in the art that given the normal structure of MXM-SXM-MEM-VXM (i.e. figure 15/14) of Abts, to implement an embodiment of MXM-VXM-SXM-MEM-VXM of Abts, this allows for more independent functional modules of the system see Col.9, ll.34-44. This embodiment of Abts discloses performing a computation on elements of the vector of the transposed tensor inputted into the vector compute channel.
Regarding claim 11, Abts disclose the invention substantially as claimed. See the discussion of claim 5 above.
Abts discloses wherein the plurality of vector compute channels and the transpose circuit are part of one of a plurality of vector compute banks [Superlanes] of the integrated circuit device [Fig.15B, Superlanes; “A Superlane processes streams of data in 16 lanes” Col.8-10, Superlanes].
Regarding claim 12, Abts disclose the invention substantially as claimed. See the discussion of claim 11 above.
Abts discloses wherein the plurality of vector compute banks is operable to perform transpose operations in parallel using the transpose circuit in each of the vector compute banks [“Each slice (for example, slice 1502) in a TSP performs any of a variety of functions under the control of instructions… All of the tiles in a slice execute the same set of instructions …each Superlane, and indeed the entire TSP, executes a single set of instructions” Col.8-9,ll.45-36; “The compiler coordinates the control and data flow of the program and specifies any instruction-level parallelism by explicitly bundling instructions that can and should execute concurrently so that they are dispatched together” Col.19-20, The Compiler; see fig.21 for concurrency and col.14-15].
Regarding claim 13, Abts disclose the invention substantially as claimed. See the discussion of claim 5 above.
Abts discloses wherein the transpose circuit includes a bypass mode of operation to output data to the plurality of vector compute channels without transposing the data. [“the results transferred from the VXM through the SXMs through the PCie modules back to the host computer” Col.10,ll.4-19, can merely be used to transfer data; “SXM can rotate or transpose a stream of data across the lanes.” Col. 13, Switching Processing Tiles, can perform other functions not only transposing; “All of the tiles in a slice execute the same set of instructions” Col.8,ll.60-65]
Regarding claims 14, 16-18 and 20, they are directed to claims 5, 7-9 and 13, respectively. A mere change in statutory class is obvious. The claims are rejected for the reasons given in the respective directed claim.
Claims 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Abts, and further in view of Bleiweiss et al. (US 2019/0205737 A1), hereinafter Bleiweiss.
Regarding claim 6, Abts disclose the invention substantially as claimed. See the discussion of claim 5 above.
Abts discloses further comprising a processing engine array operable to perform matrix multiplication computations [Fig.18; wherein MXM can perform multiply-accumulate, and dot products, see Col.12-13 The Matrix Processing Tiles]
However, Abts does not explicitly disclose wherein the transpose circuit is operable to perform transpose operations concurrently with matrix multiplication operations being performed by the processing engine array.
In the analogous art of neural network processing architecture, Bleiweiss discloses perform transpose operations concurrently with matrix multiplication operations being performed [“2D Convolution may perform better on a GPU” par.226; “Accelerator 2708 may be implemented to perform additional data deep layout transformations other than a transpose… data layout conversions may be offloaded to accelerator 2709 to enable parallel execution with other computations performed at the GPU” par.236]
It would have been obvious to one of ordinary skill in the art, having the teachings of Abts and Bleiweiss before him before the effective filing date of the claimed invention to modify the system as taught by Abts, to separate/offload layout transformations as taught by Bleiweiss, to enable parallel execution of various computations [Bleiweiss: par.234-238]. The combination of Abts and Bleiweiss discloses the transpose circuit is operable to perform transpose operations concurrently with matrix multiplication operations being performed by the processing engine array.
Regarding claim 15, it is directed to claim 6. A mere change in statutory class is obvious. The claim is rejected for the reasons given in the directed claim.
Claims 10 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Abts, in view of Bleiweiss, and further in view of Wu et al. (NPL: “Group Normalization”), hereinafter Wu.
Regarding claim 10, Abts disclose the invention substantially as claimed. See the discussion of claim 9 above.
Abts discloses performing batch normalizing or other complex functions using the vector compute channel [“The TSP supports chaining together two or more vector ALUs within each lane… this allows for efficient parallel implementations of algorithms for batch normalization, quantization, or more complex activation functions” Col.12, Vector Processing Tiles]
While Abts does not explicitly disclose the vector of the transposed tensor inputted into the vector compute channel.
It would have been obvious to one of ordinary skill in the art, to look within the specification of Abts, and take note of an embodiment wherein “additional slices of VXMs (not shown) may be allocated to a set. The additional VXM slices may be physically or logically located next to one of the MXM slices” in Col.9, ll.40-44. Therefore, it would have been obvious to one of ordinary skill in the art that given the normal structure of MXM-SXM-MEM-VXM (i.e. figure 15/14) of Abts, to implement an embodiment of MXM-VXM-SXM-MEM-VXM of Abts, this allows for more independent functional modules of the system see Col.9, ll.34-44. This embodiment of Abts discloses the vector of the transposed tensor inputted into the vector compute channel.
However, Abts does not explicitly disclose the output value is a mean, a variance, or a count of the elements of the vector of the tensor
In the analogous art of neural network processing architecture, Bleiweiss discloses computing the variables for normalization before computing the normalization [ “Batch Normalization involves normalization of an input by subtracting a mean and dividing the variance calculated over a mini-batch during training (and inference)” par.244, Fig.29 shows computing the mean and variance prior to the normalization operations; “Once calculated, the results of mean and variance operations 3003 are forwarded to complete computation of normalization operations 3004” par.245]
Additionally, Abts discloses using VXMs for batch normalizations, see col.12, ll.16-29.
It would have been obvious to one of ordinary skill in the art, having the teachings of Abts and Bleiweiss before him before the effective filing date of the claimed invention to modify the Vector Processing Tile as disclosed by Abts, to compute the variance and mean for normalization as taught by Bleiweiss, to ensure to compute the inputs required for normalization, as normalization allows for larger learning rates, decreases sensitivity to initialization of weights, and in some cases aids in increase of accuracy [Bleiweiss: par.244-247]. The combination of Abts and Bleiweiss discloses the output value is a mean, or a variance of the elements of the vector of the tensor.
However, Abts and Bleiweiss does not explicitly disclose a count of the elements of the vector of the tensor.
In the analogous art of Neural Network Normalizations, Wu teaches that normalizations requires a count of the number of elements in order to compute the mean and variance [Fig.2; See Sec.3.1 “m is the size of this set. Many types of feature normalization methods main differ in how the set Si is defined” and equations 1-7].
It would have been obvious to one of ordinary skill in the art, having the teachings of Abts, Bleiweiss, and Wu before him before the effective filing date of the claimed invention to modify the Vector Processing Tile as disclosed by Abts and Bleiweiss, to compute the variance, mean, and count for normalization as taught by Wu, to be able to compute the normalization of each element as it is well-known that normalizing the input data makes training faster [Wu: Sec.2 and 3.1].
Regarding claim 19, it is directed to claim 10. A mere change in statutory class is obvious. The claim is rejected for the reasons given in the directed claim.
Response to Arguments
Applicant’s arguments, see page 10, filed 07/09/2026, with respect to Objections to the Specification have been fully considered and are persuasive. The Objections to the Specification of the Office Action mailed 04/09/2026 (hereinafter Prior Office Action) has been withdrawn.
Applicant's arguments, see pages 10-11, filed 07/09/2026, with respect to Rejections under 35 U.S.C. 103 have been fully considered but they are not persuasive.
On page 11, applicant argument with respect Abts, is not persuasive, as now in view of the new rejection above, relies on an embodiment wherein “the additional VXM slices may be physically or logically located next to one of the MXM slices” in Col.9, ll.40-44. As such, the switching tiles are communicatively coupled between the memory and vector processing tiles.
On page 11, applicant’s arguments directed to Zejda are now moot as the new rejection does not rely on Zejda.
The examiner respectfully disagrees with the applicant’s assertion to the contrary for at least the reasons given above.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Kenny K. Bui whose telephone number is (571)270-0604. The examiner can normally be reached 8:00 am to 3:00 pm on Monday, 8:00 am to 4:00 pm on Tuesday to Friday ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew T Caldwell can be reached at (571)272-3702. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KENNY K. BUI/Patent Examiner, Art Unit 2182 (571)270-0604
/ANDREW CALDWELL/Supervisory Patent Examiner, Art Unit 2182