Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending.
The IDSes, filed 4/8/26 and 10/21/25, have been considered.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-4, 9-13, and 18-20 is/are rejected under 35 U.S.C. 102(a)( as being anticipated by Kreinin et al (US20220276964, “Kreinin”).
As to claim 1:
Kreinin teaches a computer-implemented method (0008, 0013, 0020, 0025) for accessing memory banks of a hardware accelerator (memory 118 accessible by SoC 102 components: CPU 104, controller 108, accelerators 112, 114; 0103), the method comprising:
receiving an instruction for a compute tile of the hardware accelerator (instructions for SoC 102; 0101, 0103), the instruction being executable at the compute tile to cause performance of operations comprising:
identifying, in the instruction, an opcode for an operation that involves accessing one or more banks of a first memory of the compute tile (opcode of command and its operands; 0393; DPUs 3330a-m with AGUs 3350a-n typically execute code by accepting control instructions in the form of AGU parameters (e.g., 3320) from control subsystem 3310. The AGUs 3350a-n may use these parameters 3320a-k to generate addresses for the DPUs 3330a-m to write to and from external or other memory; 0449);
generating, based on the opcode, a byte addressing sequence used to access different lengths of contiguous bytes stored at the one or more banks (generate vector addresses to access contiguous data; 0394; AGUs 3350a-n may aid memory access by the DPUs 3330a-m, such as in loading and storing data to and from memories, accessing the data in its various locations in memory may require the generation and use of a large number of memory addresses; 0447);
based on the byte addressing sequence, performing, at the first memory, a plurality of byte-addressable memory accesses to obtain a plurality of input vectors from the one or more banks of the first memory (obtain vectors for read; 0393-0395; a single data request for eight bytes of data may map to eight different addresses in memory that are not part of a contiguous data block and/or are not stored proximately in memory; a single request for eight bytes of vector data may include single byte vector components that are each stored on different banks 2310a-h.; 0399); and
processing, based on the instruction, the plurality of input vectors at the hardware accelerator (execute instructions on SoC 102; 0101, 0103; external memory 118 can be potentially accessed by other components in the SoC 102 (e.g., the CPU 104, peripheral device controller 108, and the accelerators 112 and 114; 0103; accessing each of the eight banks 2310a-h in the same cycle via a “gather” memory operation may present advantages over accessing a single memory bank, especially via a scalar or single vector memory read; 0398; a single request for eight bytes of vector data may include single byte vector components that are each stored on different banks 2310a-h; , a single vector read, as described above, may not be optimized for handling such a request. For example, a single vector read to a single memory bank may take eight cycles to read such a data set, taking one cycle to read from each of the different addresses; 0399).
As to claim 2, 11:
Kreinin teaches the byte addressing sequence is operable to access the different lengths of contiguous bytes starting from any given byte address across the one or more banks of the first memory (a single data request for eight bytes of data may map to eight different addresses in memory that are not part of a contiguous data block and/or are not stored proximately in memory. As an example, it is possible that a single request for eight bytes of vector data may include single byte vector components that are each stored on different banks 2310a-h; single vector read, as described above, may not be optimized for handling such a request. For example, a single vector read to a single memory bank may take eight cycles to read such a data set, taking one cycle to read from each of the different addresses; 0399).
As to claim 3, 12:
Kreinin teaches generating, based on the opcode, a byte access request that provides byte-granularity access for obtaining at least one byte of data stored at a given row of a bank among the one or more banks of the first memory (accessing each of the eight banks 2310a-h in the same cycle via a “gather” memory operation may present advantages over accessing a single memory bank, especially via a scalar or single vector memory read. This potential advantage is best shown by example. Each bank 2310a-h may have an access width of, for example, one byte per bank; 0398-0400; a read mask that can specify the particular byte within a word that is to be read from; 0016; write a byte of data to the memory module, wherein the byte of data is located within a word of data stored on the memory module, 0019).
As to claim 4, 13:
Kreinin teaches generating the byte addressing sequence comprises:
generating a byte addressing sequence corresponding to a plurality of byte access requests, each byte access request being used to access one or more distinct portions of byte- level data stored at a respective address location across the one or more banks of the first memory (Each bank 2310a-h may have an access width of, for example, one byte per bank. In this case, if an operation requires eight bytes of data, each byte stored on a different bank, the architecture 2300 could conceivably provide all eight required bytes of data in the same cycle. Doing so, however, would require implementing a scheme to map addresses of the required bytes to the bank 2310a-h corresponding to the required data and the proper addresses are provided by initiators 2310a-h; 0398-0401).
As to claim 9, 18:
Kreinin teaches processing the plurality of input vectors comprises: processing the plurality of input vectors through a neural network layer of a neural network implemented at the hardware accelerator; and generating an output for the neural network layer in response to processing the plurality of input vectors through the neural network layer (use accelerator in SoC in machine learning applications; 0144-0145, 0149, 0159; to generate output information via I/O devices; 0655).
As to claim 10, 19:
Kreinin teaches a system comprising a hardware accelerator (system with SoC 102; 0026, 0090, 0099); a processing device (CPU 104; Fig .1); and a non-transitory machine-readable storage medium for storing instructions used to access memory banks of a hardware accelerator (computer readable medium storing instructions; 0021, 0024, 0095-0097), the instructions being executable by the processing device to cause performance of operations comprising:
receiving an instruction for a compute tile of the hardware accelerator, the instruction being executable at the compute tile to cause performance of operations comprising:
identifying, in the instruction, an opcode for an operation that involves accessing one or more banks of a first memory of the compute tile (DPUs 3330a-m with AGUs 3350a-n typically execute code by accepting control instructions in the form of AGU parameters (e.g., 3320) from control subsystem 3310. The AGUs 3350a-n may use these parameters 3320a-k to generate addresses for the DPUs 3330a-m to write to and from external or other memory; 0449);
generating, based on the opcode, a byte addressing sequence used to access different lengths of contiguous bytes stored at the one or more banks (AGUs 3350a-n may aid memory access by the DPUs 3330a-m, such as in loading and storing data to and from memories, accessing the data in its various locations in memory may require the generation and use of a large number of memory addresses; 0447);
based on the byte addressing sequence, performing, at the first memory, a plurality of byte-addressable memory accesses to obtain a plurality of input vectors from the one or more banks of the first memory (a single data request for eight bytes of data may map to eight different addresses in memory that are not part of a contiguous data block and/or are not stored proximately in memory; a single request for eight bytes of vector data may include single byte vector components that are each stored on different banks 2310a-h.; 0399); and
processing, based on the instruction, the plurality of input vectors at the hardware accelerator (external memory 118 can be potentially accessed by other components in the SoC 102 (e.g., the CPU 104, peripheral device controller 108, and the accelerators 112 and 114; 0103; accessing each of the eight banks 2310a-h in the same cycle via a “gather” memory operation may present advantages over accessing a single memory bank, especially via a scalar or single vector memory read; 0398; a single request for eight bytes of vector data may include single byte vector components that are each stored on different banks 2310a-h; , a single vector read, as described above, may not be optimized for handling such a request. For example, a single vector read to a single memory bank may take eight cycles to read such a data set, taking one cycle to read from each of the different addresses; 0399).
As to claim 20:
Kreinin teaches the byte addressing sequence is operable to access the different lengths of contiguous bytes starting from any given byte address across the one or more banks of the first memory (a single data request for eight bytes of data may map to eight different addresses in memory that are not part of a contiguous data block and/or are not stored proximately in memory; a single request for eight bytes of vector data may include single byte vector components that are each stored on different banks 2310a-h; single vector read, as described above, may not be optimized for handling such a request. For example, a single vector read to a single memory bank may take eight cycles to read such a data set, taking one cycle to read from each of the different addresses; 0399); and the operations further comprise: generating, based on the opcode, a byte access request that provides byte- granularity access for obtaining at least one byte of data stored at a given row of a bank among the one or more banks of the first memory (accessing each of the eight banks 2310a-h in the same cycle via a “gather” memory operation may present advantages over accessing a single memory bank, especially via a scalar or single vector memory read. Each bank 2310a-h may have an access width of, for example, one byte per bank; 0398-0400; a read mask that can specify the particular byte within a word that is to be read from; 0016; write a byte of data to the memory module, wherein the byte of data is located within a word of data stored on the memory module, 0019).
Allowable Subject Matter
Claims 5-8 and 14-17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
As to claim 5/14, the prior art does not further suggest the invention of claim 4/13, wherein the opcode is for a tensor operation that is executed at the compute tile to traverse respective elements of an input tensor.
Claims 6-8 and 15-17 are also allowable for incorporating the limitations of claim 5/14, and further limitations.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US20130318067 discloses techniques are provided for hardware-accelerated relational joins. A first table comprising one or more rows is processed through a hardware accelerator. At least one join column in at least one of the one or more rows of the first table is hashed to set at least one bit in at least one bit vector. A second table comprising one or more rows is processed through a hardware accelerator. At least one join column in at least one of the one or more rows of the second table is hashed to generate at least one hash value. At least one bit vector is probed using the at least one hash value. A joined row is constructed responsive to the probing step. The row-construction step is performed in the hardware accelerator.
US11257183 discloses method may include determining a set of filter vectors. Each filter vector in the set of filter vectors may include a set of filter weights associated with at least one portion of an output volume of a resampling operation. The method may also include generating, via a clustering algorithm and based on the set of filter vectors, a filter bank for the resampling operation. The filter bank may include an additional set of filter vectors. The method may further include transmitting the filter bank to a memory module included in a hardware accelerator, and directing the hardware accelerator to execute the resampling operation using an input volume and the filter bank.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THAN NGUYEN whose telephone number is (571)272-4198. The examiner can normally be reached M-F 7:00am -4:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tim Vo can be reached at (571)272-3642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THAN NGUYEN/Primary Examiner, Art Unit 2138