DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1, 10, 12, 14, 16, 19, and 21 have been amended.
Claims 1, 3-10, 12-14, 16-19, and 21-23 have been examined.
The § 112 rejections in the previous Office Action have been addressed and are withdrawn.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on March 24, 2026 has been entered.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 3-10, 12-14, 16-19, and 21-23 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the Applicant regards as the invention.
Claim 1 recites, “an ending address of a preceding tensor of the any two adjacent tensors in the second tensor is the same as a starting address of a succeeding tensor of the any two adjacent tensors in the second tensor.” This language contradicts portions of the specification. For example, the specification states, at paragraph 81, “the plurality of first tensors obtained by the first processor and the plurality of first tensors included in the second tensor have completely same content and occupy same space, but have different addresses.” This indicates that no data is lost or overwritten due to the copying. This is supported by Figure 4, which shows the 6 bytes previously stored at 0x0030-0x0035 are moved to 0x0050-0x0055. The next 8 bytes, previously stored at 0x0040-0x0047, are moved to 0x0056-0x005d. In this interpretation the ending address of the preceding tensor is 0x0055 and the starting address of the succeeding tensor is 0x0056. Thus, contrary to what is claimed, the addresses are not the same, but are consecutive. If the addresses were the same, then both the preceding and succeeding would write to the same address, e.g., 0x0056, and data would be lost due to being overwritten. This contradiction renders the scope of the claims indefinite. For purposes of examination, this limitation is interpreted as, “an ending address of a preceding tensor of the any two adjacent tensors in the second tensor is consecutive to a starting address of a succeeding tensor of the any two adjacent tensors in the second tensor.”
Claims 3-9, 12, 13, 16-18, and 21-23 are rejected as depending from rejected base claims and failing to cure the indefiniteness of those base claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 4, 8-10, 12-14, 16, 17, 19, 21, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over US Publication No. 2019/0042092 by Wu et al. (hereinafter referred to as “Wu”) in view of US Publication No. 2018/0165381 by Venkatesh et al. (hereinafter referred to as “Venkatesh”).
Regarding claims 1 and 14, taking claim 1 as representative, Wu discloses:
a tensor processing method, comprising: …a plurality of copying instructions …a second processor, wherein the plurality of copying instructions corresponds to a first plurality of tensors, wherein any copying instruction in the plurality of copying instructions comprises a second identifier, and wherein the second identifier indicates a first tensor corresponding to the any copying instruction (Wu discloses, at ¶ [0024], a CPU, i.e., a second processor, and executing, by another processor, a plurality of copy operations corresponding to a first plurality of tensors, which discloses the copy operations identifying the source and target data locations.);
obtaining, using corresponding second identifiers from each of the plurality of copying instructions, the first plurality of tensors (Wu discloses, at Figure 1 and related description, processor circuitry obtaining tensor data, which discloses using the second identifiers to do so.);
copying the first plurality of tensors into a second tensor, wherein the first plurality of tensors occupies consecutive spaces in the second tensor and wherein for any two adjacent tensors of the first plurality of tensors in the second tensor, an ending address of a preceding tensor of the any two adjacent tensors in the second tensor is the same as a starting address of a succeeding tensor of the any two adjacent tensors in the second tensor (Wu discloses, at Figure 4 and related description, gathering tensor data from multiple strides to contiguous locations. Contiguous discloses consecutive ending and starting addresses for adjacent tensors. As disclosed at ¶ [0019], strides can be columns or rows, which discloses tensors.);
receiving a first processing instruction from the second processor, …[that] indicates the second tensor, and …indicates a first processing operation (Wu discloses, at Figure 1 and related description, the DPU, i.e., the first processor, receiving tensor operations from a CPU, i.e., the second processor, which discloses indicating a tensor and an operation to be performed using the tensor.); and
processing based on the first processing operation, the second tensor (Wu discloses, at Figure 1 and related description, the DPU receiving tensor operations from a CPU, which discloses performing the operations indicated using the data indicated.).
Wu does not explicitly disclose the aforementioned plurality of copying instructions are received from the aforementioned second processor, the aforementioned first processing instruction comprises a first identifier and a first processing identifier, wherein the first identifier indicates the aforementioned second tensor and the first processing identifier indicates the aforementioned processing operation.
However, Wu discloses the DPU receiving instructions from the CPU, as discussed above. It would have been obvious to transmit the copy instructions from the CPU to the DPU because doing so is an obvious design choice representing tradeoffs that are well-known in the art. For example, Wu discloses, e.g., at ¶ [0025], that the copy instructions can be generated at the DPU instead of transmitted by the CPU, and doing so has the benefit of using less bandwidth. However, it is evident that doing so also has costs, such as requiring additional circuit complexity and processing power on the part of the DPU. Accordingly, it would have been obvious to a person having ordinary skill in the art, based on design considerations, to have transmitted the copy instructions from the CPU.
Also in the same field of endeavor (e.g., processors) Venkatesh discloses:
an instruction that indicates the operation and corresponding data (Venkatesh discloses, at ¶ [0045], instructions include opcodes to indicate operation and identify the data on which the operation is to be performed.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Wu to include the format disclosed by Venkatesh in order to provide a relatively concise mechanism for enabling correct operation.
Regarding claims 3 and 16, taking claim 3 as representative, Wu, as modified, discloses the elements of claim 1, as discussed above. Wu also discloses:
wherein the any copying instruction further comprises first address information, wherein the first address information indicates a third address, and wherein copying the first plurality of tensors into the second tensor comprises: copying any first tensor into the third address in the second tensor (Wu discloses, at Figure 4 and related description, copying first tensors into a second tensor, which implicitly discloses address information indicating an address in the second tensor to which the first tensors are to be copied.).
Regarding claims 4, 17, and 22, taking claim 4 as representative, Wu, as modified, discloses the elements of claim 1, as discussed above. Wu also discloses:
further receiving the plurality of copying instructions from the second processor in a target sequence, wherein copying the first plurality of tensors into the second tensor comprises: sequentially copying the first plurality of tensors into the second tensor in the target sequence (Wu discloses, at Figure 4 and related description, copying first tensors into a second tensor, which implicitly discloses specifying a target sequence used to perform the copying.).
Regarding claim 8, Wu, as modified, discloses the elements of claim 1, as discussed above. Wu also discloses:
receiving a second processing instruction from the second processor, …[that] indicates a third tensor, wherein the third tensor comprises a portion of the first plurality of tensors in the second tensor… (Wu discloses, at Figure 1 and related description, the DPU, i.e, the processor, receiving tensor operations from a CPU, which implicitly discloses indicating a tensor and an operation to be performed using the tensor.); and
processing based on the second processing operation, the third tensor (Wu discloses, at Figure 1 and related description, the DPU, i.e, the processor, receiving tensor operations from a CPU, which discloses performing the operations indicated using the data indicated.).
Wu does not explicitly disclose the aforementioned second processing instruction comprises a fourth identifier and a second processing identifier, wherein the fourth identifier indicates the aforementioned third tensor and the second processing identifier indicates the aforementioned second processing operation.
However, in the same field of endeavor (e.g., processors) Venkatesh discloses:
an instruction that indicates the operation and corresponding data (Venkatesh discloses, at ¶ [0045], instructions include opcodes to indicate operation and identify the data on which the operation is to be performed.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Wu to include the format disclosed by Venkatesh in order to provide a relatively concise mechanism for enabling correct operation.
Regarding claim 9, Wu, as modified, discloses the elements of claim 8, as discussed above. Wu also discloses:
the third tensor comprises the first plurality of tensors, and wherein the first plurality of tensors is adjacent first plurality of tensors in the second tensor (Wu discloses, at Figure 4 and related description, gathering tensor data from multiple strides to contiguous locations in a single stride.).
Regarding claims 10 and 19, taking claim 10 as representative, Wu discloses:
a tensor processing method, comprising: determining a first identifier that indicates a second tensor, wherein a first plurality of tensors comprised in the second tensor occupies consecutive space in the second tensor and wherein for any two adjacent tensors of the first plurality of tensors in the second tensor, an ending address of a preceding tensor of the any two adjacent tensors in the second tensor is the same as a starting address of a succeeding tensor of the any two adjacent tensors in the second tensor, and wherein the first plurality of tensors in the second tensor are based on copying the first plurality of tensors (Wu discloses, at Figure 1 and related description, a DPU receiving tensor operations from a CPU, which discloses indicating a tensor and an operation to be performed using the tensor, which encompasses determining a corresponding identifier. Wu discloses, at Figure 4 and related description, gathering tensor data from multiple strides to contiguous locations in a single stride, which discloses sharing part of an address, i.e., being consecutive. Contiguous discloses consecutive ending and starting addresses for adjacent tensors. As disclosed at ¶ [0019], strides can be columns or rows, which discloses tensors.);
determining a first processing identifier that indicates a first processing operation corresponding to the second tensor (Wu discloses, at Figure 1 and related description, a DPU receiving tensor operations from a CPU, which discloses indicating a tensor and an operation to be performed using the tensor, which encompasses determining a corresponding processing identifier.); and
sending to a first processor, a first processing instruction… wherein the first processing instruction instructs the first processor to process, based on the first processing operation, the second tensor (Wu discloses, at Figure 1 and related description, a DPU receiving tensor operations from a CPU, which implicitly discloses indicating a tensor and an operation to be performed using the tensor.);
determining a second plurality of identifiers that indicate the first plurality of tensors, wherein the first plurality of tensors correspond to the second plurality of identifiers; and
… a plurality of copying instructions carrying the second plurality of identifiers, wherein the second plurality of identifiers correspond to the first plurality of tensors, and wherein the plurality of copying instructions instructs the first processor to copy the first plurality of tensors into the second tensor.
Wu does not explicitly disclose the aforementioned first processing instruction carrying the first identifier and the first processing identifier and sending, to the first processor, the aforementioned plurality of copying instructions.
However, Wu discloses the DPU receiving instructions from the CPU, as discussed above. It would have been obvious to transmit the copy instructions from the CPU to the DPU because doing so is an obvious design choice representing tradeoffs that are well-known in the art. For example, Wu discloses, e.g., at ¶ [0025], that the copy instructions can be generated at the DPU instead of transmitted by the CPU, and doing so has the benefit of using less bandwidth. However, it is evident that doing so also has costs, such as requiring additional circuit complexity and processing power on the part of the DPU. Accordingly, it would have been obvious to a person having ordinary skill in the art, based on design considerations, to have transmitted the copy instructions from the CPU.
Also, in the same field of endeavor (e.g., processors) Venkatesh discloses:
an instruction that indicates the operation and corresponding data (Venkatesh discloses, at ¶ [0045], instructions include opcodes to indicate operation and identify the data on which the operation is to be performed.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Wu to include the format disclosed by Venkatesh in order to provide a relatively concise mechanism for enabling correct operation.
Regarding claims 12 and 21, Wu, as modified, discloses the elements of claim 10, as discussed above. Wu also discloses:
…the tensor processing method further comprises: determining first address information that indicates a third address corresponding to any first tensor; and … carry the first address information… instructs the first processor to copy the any first tensor into the second tensor (Wu discloses, at Figure 4 and related description, operations for copying first tensors into a second tensor, which implicitly discloses address information indicating an address in the second tensor to which the first tensors are to be copied.)..
Wu does not explicitly disclose the aforementioned determining is before the sending to the first processor the copying instructions carrying the plurality of second identifiers and the aforementioned copying instruction is used to carry the first address information. However, Wu discloses the DPU receiving instructions from the CPU, as discussed above. It follows that the determining is before the sending. It would have been obvious to transmit the copy instructions from the CPU to the DPU because doing so is an obvious design choice representing tradeoffs that are well-known in the art. For example, Wu discloses, e.g., at ¶ [0025], that the copy instructions can be generated at the DPU instead of transmitted by the CPU, and doing so has the benefit of using less bandwidth. However, it is evident that doing so also has costs, such as requiring additional circuit complexity and processing power on the part of the DPU. Accordingly, it would have been obvious to a person having ordinary skill in the art, based on design considerations, to have transmitted the copy instructions from the CPU.
Regarding claim 13, Wu, as modified, discloses the elements of claim 10, as discussed above. Wu also discloses:
wherein the sending to the first processor, the plurality of copying instructions carrying the second plurality of identifiers comprises: sending the plurality of plurality of copying instructions to the first processor in a target sequence, wherein the plurality of copying instructions instructs the first processor to sequentially copy the first tensors into the second tensor in the target sequence (Wu discloses, at Figure 4 and related description, copying first tensors into a second tensor, which implicitly discloses specifying a target sequence used to perform the copying.).
Claims 5 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Venkatesh in view of US Publication No. 2019/0370631 by Fais et al. (hereinafter referred to as “Fais”).
Regarding claims 5 and 18, taking claim 5 as representative, Wu, as modified, discloses the elements of claim 1, as discussed above. Wu also discloses:
…the tensor processing method further comprises: receiving, by the first processor, a creation instruction from the second processor… (Wu discloses, at Figure 1 and related description, the DPU, i.e, the processor, receiving tensor operations from a CPU, which implicitly discloses a creation instruction.); and
creating the second tensor based on the amount of occupied space, wherein space occupied by the second tensor is the same as the amount of occupied space (Wu discloses, at Figure 4 and related description, gathering tensor data from first tensors to a second tensor, which discloses creating the second tensor and the second tensor occupying the same amount of space as the sum of its parts, i.e., the first tensors.).
Wu does not explicitly disclose the aforementioned receiving is before the obtaining the aforementioned first plurality of tensors, and the creation instruction comprises space information, wherein the space information indicates an amount of occupied space, and wherein the amount of occupied space is based on a sum of space occupied by the first tensors.
However, in the same field of endeavor (e.g., tensors) Fais discloses:
prior to obtaining tensor data, specifying space information, wherein the space information indicates an amount of occupied space wherein the amount of occupied space is based on a sum of space occupied by the first tensors (Fais discloses, at ¶ [0031], allocating space for input and output tensors prior to producing the tensors. The space needed for a tensor is understood to be the sum of its parts.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Wu to include Fais’s allocation in order to improve performance by ensuring efficient access to data.
Claims 6 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Venkatesh in view of Fais in view of US Publication No. 2019/0278507 by Du et al. (hereinafter referred to as “Du”).
Regarding claims 6 and 23, taking claim 6 as representative, Wu, as modified, discloses the elements of claim 5, as discussed above. Wu also discloses:
…second address information, wherein the second address information indicates a third address, and wherein the creating comprises: creating based on the amount of occupied space, the second tensor at the third address (Wu discloses, at Figure 4 and related description, creating the second tensor and the second tensor occupying the same amount of space as the sum of its parts, i.e., the first tensors. Doing so at a second address is implicit.).
Wu does not explicitly disclose the creation instruction further comprises the aforementioned second address information.
However, in the same field of endeavor (e.g., data storage) Du discloses:
instructions include destination addresses (Du discloses, at ¶ [0091], instructions include destination addresses.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Wu to include instructions specifying addresses, as disclosed by Du, in order to facilitate efficient storing of data.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Venkatesh in view of US Publication No. 2022/0253488 by Li et al. (hereinafter referred to as “Li”).
Regarding claim 7, Wu, as modified, discloses the elements of claim 1, as discussed above. Wu also discloses:
Wu does not explicitly disclose wherein after the obtaining, the tensor processing method further comprises: receiving a deletion instruction from the second processor, wherein the deletion instruction comprises a third identifier, and wherein the third identifier indicates a to-be-deleted first tensor; and deleting the to-be-deleted first tensor.
However, in the same field of endeavor (e.g., tensor processing) Li discloses:
deleting tensors (Li discloses, at ¶ [0064], deleting tensors, which implicitly discloses an instruction to do so and an identification of the tensor to be deleted.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Wu to include deleting tensors, as disclosed by Li, in order to facilitate efficient execution in systems having limited available memory.
Response to Arguments
On pages 15-16of the response filed March 24, 2026 (“response”), the Applicant argues, “Wu's data processing unit is configured to perform gather copy operations, and includes a copy operation 400 that reads tensor data 136 from data strides 402 in a data matrix 204 from the first memory circuitry 128 in non-contiguous memory locations. Wu refers to a stride as contiguous data locations in a tensor or data matrix. Wu's data processing unit 104, through the gather copy tensor operation 400, writes the tensor data 136 as contiguous matrix elements 226 within the data matrix 214 of the second memory circuitry 130. Therefore, Wu's gather operation produces contiguous output but never defines or discloses a shared-boundary-address property as between adjacent sub-tensors. The Examiner argues that "there are only two possibilities, namely the bits are the same, i.e., both bits are zero or both bits are ones, or the bits are different. Either possibility is obvious." See id., p. 5. This argument misapprehends the claimed feature entirely. Independent claim 1 does not recite that a binary digit has one of two possible values but a specific, functional memory layout property: the ending memory address of tensor A equals the beginning memory address of tensor B, confirming that the tensors are packed contiguously with a shared boundary, as illustrated in FIG. 4 and described in 81-82 of the Specification. On the contrary, Wu's gather operation copies tensor data strides to "contiguous matrix elements" within a destination memory (Wu, FIG. 4, 59). However, Wu does not disclose a shared-boundary-address relationship between adjacent copied sub-tensors within a destination. Indeed, the fact that Wu produces contiguous data does not teach or suggest the specific address relationship, as claimed in claim 1. Wu operates at the level of matrix elements and data strides, and does not disclose that adjacent copied tensors share a boundary address in the claimed manner. Further, Wu does not copy the first plurality of tensors into a second tensor, wherein for any two adjacent tensors of the first plurality of tensors in the second tensor, an ending address of a preceding tensor of the any two adjacent tensors in the second tensor is the same as a starting address of a succeeding tensor of the any two adjacent tensors in the second tensor. Thus, claims 1, 10, 14, and 19 are allowable over Wu.”
Though fully considered, the Examiner respectfully disagrees. Contiguous, as used by Wu, means touching, sharing a common border, or immediately adjacent. The claim limitations reciting a same ending and starting address are interpreted as meaning the same thing, i.e., the succeeding tensor immediately follows the preceding tensor.
As discussed above, having the same starting and ending address would result in the first byte of the succeeding tensor overwriting the last byte of the preceding tensor. For example, if the last byte of the preceding tensor had an address of 0x56, and the first byte of the succeeding tensor also had the address 0x0056, there would be data loss since both bytes can’t be stored at the same address, i.e., 0x0056, at the same time. Therefore, the claim limitations are interpreted as disclosing the two tensors are adjacent, or contiguous, as disclosed by Wu. This interpretation is consistent with Applicant’s specification and figures, e.g., at ¶ [0081] and Figure 4. Accordingly, the Applicant’s arguments are deemed unpersuasive.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAWN DOMAN whose telephone number is (571)270-5677. The examiner can normally be reached on Monday through Friday 8:30am-6pm Eastern Time.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached on 571-270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHAWN DOMAN/Primary Examiner, Art Unit 2183