DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event a determination of the status of the application as subject to AIA 35 U.S.C. 102, 103, and 112 (or as subject to pre-AIA 35 U.S.C. 102, 103, and 112) is incorrect, any correction of the statutory basis for a rejection will not be considered a new ground of rejection if the prior art relied upon and/or the rationale supporting the rejection, would be the same under either status.
Notice of Claim Interpretation
Claims in this application are not interpreted under 35 U.S.C. 112(f) unless otherwise noted in an office action.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-8 and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Lin et al. (US 2001/0016861) in view of Ko et al. (US 12,518,146).
In regards to claim 1, Lin teaches an apparatus employed in a processing device comprising:
a processor (processor 109, figure 2) configured to process data of a predefined data structure (“In particular, the present invention describes an apparatus for performing arithmetic operations using a single control signal to manipulate multiple data elements. The present invention allows execution of shift operations on packed data types.”, paragraph 0003);
a memory fetch device, coupled to the processor, configured to:
determine a plurality of addresses and sub-word data-item information of packed data, wherein the packed data is stored on a memory device that is coupled to the processor (“In another embodiment of the present invention, any one, or all, of SRC1, SRC2 and DEST, can define a memory location in the addressable memory space of processor 109. For example, SRC1 may identify a memory location in main memory 104 while SRC2 identifies a first register in integer registers 201, and DEST identifies a second register in registers 209. For simplicity of the description herein, references are made to the accesses to the register file 204, however, these accesses could be made to another memory instead.”, paragraph 0041; “Control bits SZ 610, bit twelve and bit thirteen, indicates the length of the data elements in the first and second packed data source registers. If SZ 610 equals 012, then the packed data is formatted as packed byte 501. If SZ 610 equals 102, then the packed data is formatted as packed word 502.”, paragraph 0060);
fetch the packed data from the memory device based on the plurality of addresses (“Decoder 202 accesses the register file 204, or a location in another memory, at step 302. Registers in the register file 204, or memory locations in another memory, are accessed depending on the register address specified in the control signal 207. For example, for an operation on packed data, control signal 207 can include SRC1, SRC2 and DEST register addresses. SRC1 is the address of the first source register. SRC2 is the address of the second source register.”, paragraph 0040); and
provide output data based on the fetched packed data to the processor, wherein the output data is configured according to the predefined data structure (“Where the control signal requires an operation, at step 303, functional unit 203 will be enabled to perform this operation on accessed data from register file 204.”, paragraph 0043; “In particular, the present invention describes an apparatus for performing arithmetic operations using a single control signal to manipulate multiple data elements. The present invention allows execution of shift operations on packed data types.”, paragraph 0003).
This embodiment of Lin fails to teach that the memory fetch device is configured to process the packed data in a predefined order. Another embodiment of Lin teaches that the memory fetch device is configured to process the packed data in a predefined order (“However, in another embodiment, these shifts are performed serially.”, paragraph 0077). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to try configuring the memory fetch device to process the packed data in a predefined order because it is one of three possible solutions identified by Lin in paragraph 0077 and one of ordinary skill in the art could have pursued the three possible solutions with a reasonable expectation of success because the tradeoffs between performing operations like these in serial versus in parallel is well understood.
Lin fails to teach fetch the packed data from the memory device based on the sub-word data-item information. Ko teaches fetch the packed data from the memory device based on the sub-word data-item information (“The block size specifies the length of each block (e.g., in bits, nibbles, etc.), which can be used along with the block address to identify the location of the requested block (starting from the base address). Finally, the offset value is used when the data is not stored in uniform-size blocks (e.g., for weight data that is stored as variable-size encoded filter slices) to identify the beginning of the current block.”, Col. 40, lines 7-14; “FIG. 19 illustrates a row of activation data 1900 with six blocks of data, each of which includes seven 4-bit activation values. These values are stored in three RAM words, starting at word K.”, Col. 41, lines 43-45) thereby “minimizing the number of memory reads that are required” (Col. 38, lines 29-30). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Lin with Ko to include fetch the packed data from the memory device based on the sub-word data-item information thereby “minimizing the number of memory reads that are required” (id.).
In regards to claim 2, Lin further teaches that the memory fetch device is configured to process the packed data sequentially based on the predefined order (“However, in another embodiment, these shifts are performed serially.”, paragraph 0077).
In regards to claim 3, Ko further teaches that the plurality of addresses are generated for a convolution operation performed by the processor on an array of fixed-size packed data in the packed data (“As shown, in some embodiments, the read controller decoder 1805 receives (i) a base memory address, (ii) a block address, (iii) a block size, and (iv) an offset amount. In some embodiments, the base memory address specifies a particular RAM word at which a number of blocks of data begin. This RAM word may be the first RAM word for a row of activation data or a set of weight data for a pass of a convolution layer (e.g., for numerous filter slice buffers) in some embodiments.”, Col. 39, lines 53-61).
In regards to claim 4, Ko further teaches that the predefined order is set based on the convolution operation (“For an image, e.g., these instructions might specify the order in which the pixels should be arranged and streamed to the computation fabric 315, so that the computation fabric stores this data in the appropriate locations of its memory 330 for subsequent operations.”, Col. 14, lines 31-36).
In regards to claim 5, Ko further teaches that the memory fetch device is configured to fetch additional packed data stored on the memory device and generate additional output data based on the additional packed data while the processor performs operations on the output data (“In the third stage 2215, the memory controller reads the RAM word K+4, that stores the next set of data to be loaded into the activation window buffer 2230, from the activation memory 2220, and loads this word into the read cache 2225. In addition, the (0,2) activation values are loaded into the activation window buffer 2230, at the entry points in each shift register. Concurrently, the activation values at coordinates (0,0) and (0,1) are shifted by one register block within the activation window buffer 2230.”, Col. 44, lines 39-47).
In regards to claim 6, Ko further teaches that the packed data comprises stored data and generated data, wherein the plurality of addresses comprises first addresses associated with the stored data and second addresses associated with the generated data, wherein the memory fetch device is configured to selectively fetch the stored data based on the first addresses and selectively suppress a fetch of the generated data based on the second addresses (“The second stage 2210 illustrates that the memory controller reads the RAM word K+2, that stores the next set of data to be loaded into the activation window buffer 2230, from the activation memory 2220, and loads this word into the read cache 2225. That is, the next RAM word read from memory is not the next word sequentially in memory, as word K+1 is skipped”, Col. 44, lines 28-34).
In regards to claim 7, Lin further teaches that the generated data is present in the packed data based on a predefined packing format of the packed data, wherein the processor is configured to program the predefined packing format (“Control bits SZ 610, bit twelve and bit thirteen, indicates the length of the data elements in the first and second packed data source registers. If SZ 610 equals 012, then the packed data is formatted as packed byte 501. If SZ 610 equals 102, then the packed data is formatted as packed word 502. SZ 610 equaling 002 or 112 is reserved, however, in another embodiment, one of these values could be used to indicate that the packed data is to be formatted as a packed doubleword 503.”, paragraph 0060).
In regards to claim 8, Ko further teaches that the plurality of addresses comprises one or more repeated addresses, wherein the memory fetch device is configured to selectively fetch the packed data associated with the one or more repeated addresses a single time from the memory device (“In some embodiments, the read controller decoder 1805 uses this data specifying the requested block of data to provide (i) information to the read cache 1810 enabling the read cache to retrieve this data from either RAM words already stored in the cache or RAM words in memory and (ii) the read controller merge circuit 1815 enabling the merge circuit to merge and align data received from the read cache block 1810 and output this data to the core controller.”, Col. 40, lines 15-22).
In regards to claim 16, Lin teaches a method, comprising:
determining, via a memory fetch device, a plurality of addresses and sub-word data-item information of packed data stored on a memory device (“In another embodiment of the present invention, any one, or all, of SRC1, SRC2 and DEST, can define a memory location in the addressable memory space of processor 109. For example, SRC1 may identify a memory location in main memory 104 while SRC2 identifies a first register in integer registers 201, and DEST identifies a second register in registers 209. For simplicity of the description herein, references are made to the accesses to the register file 204, however, these accesses could be made to another memory instead.”, paragraph 0041; “Control bits SZ 610, bit twelve and bit thirteen, indicates the length of the data elements in the first and second packed data source registers. If SZ 610 equals 012, then the packed data is formatted as packed byte 501. If SZ 610 equals 102, then the packed data is formatted as packed word 502.”, paragraph 0060), wherein the plurality of addresses and the sub-word data-item information are arranged in a predefined order (“FIG. 6a illustrates a general format for a control signal operating on packed data. Operation field OP 601, bit thirty-one through bit twenty-six, provides information about the operation to be performed by processor 109; for example, packed addition, packed subtraction, etc., SRC1 602, bit twenty-five through twenty, provides the source register address of a register in registers 209. This source register contains the first packed data, Source1, to be used in the execution of the control signal. Similarly, SRC2 603, bit nineteen through bit fourteen, contains the address of a register in registers 209. This second source register contains the packed data, Source2, to be used during execution of the operation. DEST 605, bit five through bit zero, contains the address of a register in registers 209. This destination register will store the result packed data, Result, of the packed data operation. Control bits SZ 610, bit twelve and bit thirteen, indicates the length of the data elements in the first and second packed data source registers.”, paragraphs 0059-0060);
fetching, via the memory fetch device, stored packed data from the memory device based on the plurality of addresses (“Decoder 202 accesses the register file 204, or a location in another memory, at step 302. Registers in the register file 204, or memory locations in another memory, are accessed depending on the register address specified in the control signal 207. For example, for an operation on packed data, control signal 207 can include SRC1, SRC2 and DEST register addresses. SRC1 is the address of the first source register. SRC2 is the address of the second source register.”, paragraph 0040); and
generating, via the memory fetch device, formatted output data from the stored packed data based a predefined data structure of a processor (“Where the control signal requires an operation, at step 303, functional unit 203 will be enabled to perform this operation on accessed data from register file 204.”, paragraph 0043), wherein the memory fetch device processes the packed data based on the predefined order (“At step 702, via internal bus 205, decoder 202 accesses integer registers 209 in register file 204 given the SRC1 602 and SRC2 603 addresses. Integer registers 209 provides functional unit 203 with the packed data stored in the SRC1 602 register (Source1), and the scalar shift count stored in SRC2 603 register (Source2).”, paragraph 0072; “At step 710, the size of the data element determines which step is to be executed next. If the size of the data elements is eight bits (byte data), then functional unit 203 performs step 712. However, if the size of the data elements in the packed data is sixteen bits (word data), then functional unit 203 performs step 714.”, paragraph 0074).
This embodiment of Lin fails to teach that the memory fetch device processes the packed data consecutively. Another embodiment of Lin teaches that the memory fetch device processes the packed data consecutively (“However, in another embodiment, these shifts are performed serially.”, paragraph 0077). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to try processing by the memory fetch device the packed data consecutively because it is one of three possible solutions identified by Lin in paragraph 0077 and one of ordinary skill in the art could have pursued the three possible solutions with a reasonable expectation of success because the tradeoffs between performing operations like these in serial versus in parallel is well understood.
Lin fails to teach fetching stored packed data from the memory device based on the sub-word data-item information. Ko teaches fetching stored packed data from the memory device based on the sub-word data-item information (“The block size specifies the length of each block (e.g., in bits, nibbles, etc.), which can be used along with the block address to identify the location of the requested block (starting from the base address). Finally, the offset value is used when the data is not stored in uniform-size blocks (e.g., for weight data that is stored as variable-size encoded filter slices) to identify the beginning of the current block.”, Col. 40, lines 7-14; “FIG. 19 illustrates a row of activation data 1900 with six blocks of data, each of which includes seven 4-bit activation values. These values are stored in three RAM words, starting at word K.”, Col. 41, lines 43-45) thereby “minimizing the number of memory reads that are required” (Col. 38, lines 29-30). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Lin with Ko to include fetching stored packed data from the memory device based on the sub-word data-item information thereby “minimizing the number of memory reads that are required” (id.).
In regards to claim 17, Ko further teaches determining, via the memory fetch device, a first set of addresses in the plurality of addresses associated with the stored packed data and a second set of addresses in the plurality of addresses associated with generated packed data (“The second stage 2210 illustrates that the memory controller reads the RAM word K+2, that stores the next set of data to be loaded into the activation window buffer 2230, from the activation memory 2220, and loads this word into the read cache 2225. That is, the next RAM word read from memory is not the next word sequentially in memory, as word K+1 is skipped”, Col. 44, lines 28-34); and
selectively skipping, via the memory fetch device, a fetch of the generated packed data (id.).
In regards to claim 18, Lin further teaches generating, via the memory fetch device, generated output data based on the second set of addresses and the predefined data structure, wherein the formatted output data comprises the generated output data (“Where the control signal requires an operation, at step 303, functional unit 203 will be enabled to perform this operation on accessed data from register file 204.”, paragraph 0043; “Control bits SZ 610, bit twelve and bit thirteen, indicates the length of the data elements in the first and second packed data source registers. If SZ 610 equals 012, then the packed data is formatted as packed byte 501. If SZ 610 equals 102, then the packed data is formatted as packed word 502. SZ 610 equaling 002 or 112 is reserved, however, in another embodiment, one of these values could be used to indicate that the packed data is to be formatted as a packed doubleword 503.”, paragraph 0060).
In regards to claim 19, Ko further teaches that the memory fetch device fetches the stored packed data sequentially from the memory device based on the predefined order (“As a contrast, FIGS. 23A-B conceptually illustrate, over four stages 2305-2320, the sequential retrieval and loading of weight data into filter slice buffers for a pass of a convolutional layer, so that the weight values can be used to compute dot products for the computation nodes of the pass.”, Col. 44, line 64 - Col. 45, line 1).
In regards to claim 20, Lin further teaches reprogramming, via the processor, format information of the packed data, wherein the memory fetch device switches from processing the packed data based on a first packing format to processing the packed data based on a second packing format while generating the formatted output data, wherein the first packing format is different from the second packing format (“Control bits SZ 610, bit twelve and bit thirteen, indicates the length of the data elements in the first and second packed data source registers. If SZ 610 equals 012, then the packed data is formatted as packed byte 501. If SZ 610 equals 102, then the packed data is formatted as packed word 502. SZ 610 equaling 002 or 112 is reserved, however, in another embodiment, one of these values could be used to indicate that the packed data is to be formatted as a packed doubleword 503.”, paragraph 0060).
Allowable Subject Matter
Claims 9-15 are allowed.
The following is a statement of reasons for the indication of allowable subject matter: The terminal disclaimer has overcome the double patenting rejection.
Response to Arguments
Applicant’s arguments, see sections I and IV-VI, filed 20 July 2026, with respect to the 112 rejection and double patenting rejections have been fully considered and are persuasive. The 112 rejection and double patenting rejections have been withdrawn.
Applicant's arguments, see sections II and III, filed 20 July 2026 have been fully considered but they are not persuasive.
The Examiner notes the Lin does teach sub-word data-item information as explained above. Ko has been incorporated to teach fetching the packed data based on the sub-word data-item information. In regards to the predefined order, Lin teaches a predefined order in the form of a format of bits for specifying the operation. The Examiner disagrees that Lin’s serial processing is not based on the predefined order. The predefined order specifies which bits are used to determine what data is being processed. Not just any data is being processed.
Conclusion
Applicant's amendment necessitated the new grounds of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NATHAN SADLER whose telephone number is (571)270-7699. The examiner can normally be reached Monday - Friday 8am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Reginald Bragdon can be reached at (571)272-4204. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Nathan Sadler/Primary Examiner, Art Unit 2139 9 September 2026