DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to amendments filed February 27th, 2026. The status of the claims is as follows. Claims 1, 8, 11 and 17 are amended. Claims 1-17, 19, and 21-22 are currently pending.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-5, 9-11, 13-17, 19, 22 are rejected under 35 U.S.C. 103 as being anticipated by Ovsiannikov et al. (US20200336272 A1, hereinafter “Ovsiannikov”) in view of Ko et al. (KR20210053791 A, hereinafter “Ko”) further in view of Boesch et al. (US20210397933A1, hereinafter “Boesch”)
Regarding Claim 1,
Ovsiannikov discloses an apparatus comprising at least one memory; instructions in the apparatus; and processor circuitry to execute the instructions; (Ovsiannikov [0055]; “The software may be embodied as a software package, code and/or instruction set or instructions, and the term “hardware,” as used in any implementation described herein, may include, for example, singly or in any combination, hardwired circuitry, programmable circuitry, state machine circuitry, and/or firmware that stores instructions executed by programmable circuitry.”,
Ovsiannikov [0125]; “At 1002, the compressed pixel data for the pixel array 1001 is depicted as stored in, for example, the non-zero value data region 621 of the memory 620 in FIG. 6A. Bit-mask data 1003 and row-pointer data 1004 for the pixel array 1001 may be stored in, for example, the metadata region 622 of the memory 620.”)
determine a slice size based on a memory size difference between the first and second data lanes, the first data lane including first data, the second data lane including second data; (Ovsiannikov [0088]; “More specifically, unpacking packed streams involves first determining lengths of original streams, which is readily available from zero bit mask 303 at the beginning of a packed block, followed by reproducing the calculations performed during packing when stream lengths within pairs (or pairs-of-pairs, or pairs-of-pairs-of-pairs) were compared to determine which head of which stream is to be cropped-and-appended to which stream. At that time, the difference between stream lengths was determined or calculated, divided by two to determine the length of the head to be cropped-and-appended,” wherein the original bit streams read on a first and second data lane; wherein the first larger bit stream identified as the stream to be cropped-and-appended to the identified second smaller steam reads on identification of first and second data lanes based on relative memory sizes of the data lanes,
Ovsiannikov [0006]; “The controller may control the multiplexers so that the multiplexers of the last column output the 2.sup.N channels of bit streams of unpacked bit streams in which each has the bit-stream length that corresponds to unpacked data of the bit stream. In one embodiment, each packed bit stream received by a multiplexer in the first column may include a zero bit mask portion and a non-zero data portion.” Wherein the pairs of bit streams are read on as partitions of bit streams to be repacked),
and the third data lane including third data, the third data lane having size smaller than the first data lane and larger than the second data lane (Ovsiannikov [Figure 2A];
PNG
media_image1.png
314
655
media_image1.png
Greyscale
wherein the disclosed embodiment of partition 201 thus discloses a third data lane (bit stream 2) having size smaller than the first data lane (bit stream 0) and larger than the second data lane (bit stream 1))
and the first and second data lanes being identified based on relative memory sizes of the data lanes in the partition including the first, second, and third data lanes; (Ovsiannikov [0088]; “More specifically, unpacking packed streams involves first determining lengths of original streams, which is readily available from zero bit mask 303 at the beginning of a packed block, followed by reproducing the calculations performed during packing when stream lengths within pairs (or pairs-of-pairs, or pairs-of-pairs-of-pairs) were compared to determine which head of which stream is to be cropped-and-appended to which stream. At that time, the difference between stream lengths was determined or calculated, divided by two to determine the length of the head to be cropped-and-appended,” wherein the first larger bit stream identified as the stream to be cropped-and-appended to the identified second smaller steam reads on identification of first and second data lanes based on relative memory sizes of the data lanes in the pair partitions of bit streams
Ovsiannikov [Figure 2A];
PNG
media_image1.png
314
655
media_image1.png
Greyscale
wherein the partition 201 disclosing bit streams 0, 1, and 2 thus read on the partition including at least a first, second, and third data lane)
remove a portion of the first data from the first data lane based on the slice size; (Ovsiannikov Fig. 2A:
PNG
media_image1.png
314
655
media_image1.png
Greyscale
Ovsiannikov [0088]; “More specifically, unpacking packed streams involves first determining lengths of original streams, which is readily available from zero bit mask 303 at the beginning of a packed block, followed by reproducing the calculations performed during packing when stream lengths within pairs (or pairs-of-pairs, or pairs-of-pairs-of-pairs) were compared to determine which head of which stream is to be cropped-and-appended to which stream. At that time, the difference between stream lengths was determined or calculated, divided by two to determine the length of the head to be cropped-and-appended, with optional padding to avoid having fractional part after division. The calculations provide offsets into packed streams pointing to where each cropped-and-appended head may be located in storage. During unpacking, the butterfly shuffler 101 may be controlled to swap back cropped heads between channels to restore original streams”),
append the portion of the first data to the second data lane to modify the size of the second data lane relative to the first data lane while the size of the third data lane remains intact. (Ovsiannikov [0088]; “More specifically, unpacking packed streams involves first determining lengths of original streams, which is readily available from zero bit mask 303 at the beginning of a packed block, followed by reproducing the calculations performed during packing when stream lengths within pairs (or pairs-of-pairs, or pairs-of-pairs-of-pairs) were compared to determine which head of which stream is to be cropped-and-appended to which stream. At that time, the difference between stream lengths was determined or calculated, divided by two to determine the length of the head to be cropped-and-appended, with optional padding to avoid having fractional part after division. The calculations provide offsets into packed streams pointing to where each cropped-and-appended head may be located in storage. During unpacking, the butterfly shuffler 101 may be controlled to swap back cropped heads between channels to restore original streams” wherein the second stream in which a length of the first stream is appended on the second stream reads on modifying the size of the second data lane relative to the first data lane
Ovsiannikov [Figure 2A];
PNG
media_image1.png
314
655
media_image1.png
Greyscale
wherein the crop-appending action occurring between pairs of bit streams within the partition 201 thus reads on the appending action occurring between the first and second data lanes while the third stream maintains intact),
Ovsiannikov fails to disclose but Ko discloses to determine sizes of data lanes in a partition of neural network weights including a first data lane, a second data lane, and a third data lane; (Ko [Pg. 6, Paragraph 16]; “The processor 410 may obtain a bit representation 6 of neural network data from a memory (420 in FIG. 4). The obtained bit representation 6 may be a bit stream. For example, the neural network data may include any one of activation, weight, and gradient of the neural network (2 in FIG. 2), and the processor 410 is a bit stream of weight data of the neural network (2 in FIG. 2).”,
[Pg. 6, Paragraph 22]; “The processor 410 may generate candidate profiles 61, 62, 63, and 64 by differently applying at least one of partitioning schemes and compression schemes for dividing bit representations into one or more lanes … In profile 1 (61), the number of lanes is 4, and the bit widths corresponding to each of the lanes are 2, 3, 4, and 1.").
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Ovsiannikov’s method of bit-manipulation performed on partitions of data bit streams of multiplexers to instead be performed on Ko’s partitions of multiple data lanes (including a first, second, and third data lane) of neural network weights. The motivation to do so is “reducing the amount of memory traffic of the neural network 2 operation and improving the performance of the neural network 2 operation limited by the memory bound” (Ko [Pg. 5, Section 1-5 Paragraph 3]).
The combination of Ovsiannikov/Ko fails to explicitly disclose but Boesch discloses to
cause a plurality of decompression engines operating at least partially in parallel to decompress the data lanes …, the plurality of decompression engines to provide the decompressed data lanes to a neural network for inference execution (Boesch [0006]; “the one or more vector decompression engines receive encoded kernel data streams comprising one or more kernel data frames, and the one or more kernel data frames include one or more data markers that each indicate a data type of one or more subsequent portions of the encoded kernel data stream. In an embodiment, the indicated data type is a first type signifying compressed kernel data values or a second type signifying kernel decompression tables. In an embodiment, a data marker indicates: a position associated with a next additional data marker within the kernel data frame; a table indicator associated with the data marker; or combinations thereof. In an embodiment, during operation in the second mode, a number of the plurality of MAC circuits of a MAC cluster of one of the convolutional accelerators multiply and accumulate received feature data and kernel data in parallel. In an embodiment, each of the one or more convolution accelerators comprises a vector decompression engine that, during operation in the second mode, provides kernel decompression tables to the feature line buffer for storage and decompressed kernel data to the MAC cluster of the convolutional accelerator” wherein each of a plurality of parallel decompression engines are receiving and decompressing kernel data streams,
Boesch [0024]; “FIG. 9 is a block diagram depicting an embodiment of a vector decompression engine within a convolution accelerator configured in accordance with one or more techniques described herein.”,
Boesch [0032]; " Convolutional Neural Networks (CNN) are types of Deep Neural Networks (DNN) with one or multiple layers, each of which perform a convolution on a 3-dimensional (3D) feature data tensor (expressed as width×height×depth). Typically, the convolution operation is associated with a majority of the processing workload, commonly performing a large number of multiply-accumulate (MAC) operations per inference. Dedicated convolution accelerators are designed to process convolution operations more efficiently, such as by exploiting a higher level of data parallelism than standard processor cores. Many CNNs also include Fully Connected (FC) layers, in which the classical 3D convolution is deformed into a Vector by Matrix operation on a feature data tensor of 1×1×Depth. These FC layers may typically be associated with a far lower level of data reuse than typical convolution operations, and may be associated with a much higher kernel data bandwidth per MAC operation compared to the classical 3D convolution” wherein the convolution accelerators in which the vector decompression engine works to process data for reads on a the plurality of decompression engines providing the decompressed data to a neural network for inference)
Boesch discloses to cause a plurality of decompression engines operating at least partially in parallel to decompress the data lanes …, the plurality of decompression engines to provide the decompressed data lanes to a neural network for inference execution. Boesch fails to explicitly disclose to cause a plurality of decompression engines operating at least partially in parallel to decompress the data lanes after the appending of the portion of the first data, the plurality of decompression engines to provide the decompressed data lanes to a neural network for inference execution. By using Boesch’s plurality of decompression engines in Ovsiannikov/Ko’s method of data lane appending, the combination of Ovsiannikov/Ko/Boesch reads on causing a plurality of decompression engines operating at least partially in parallel to decompress the data lanes after the appending of the portion of the first data, the plurality of decompression engines to provide the decompressed data lanes to a neural network for inference execution.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Ovsiannikov/Ko’s method of bit-manipulation performed on neural network weight partitions of data bit streams of multiplexers to be input into Boesch’s plurality of decompression engines for neural network inference. The motivation to do so is “utilizing a feature line buffer as decompression kernel table storage in conjunction with a vector decompression engine or circuit 657 in order to increase kernel data bandwidth” (Boesch [0049]).
Regarding Claim 2,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 1 (and thus the rejection of Claim 1 is incorporated). Ovsiannikov/Ko/Boesch further discloses to record a value corresponding to the slice size in the second data lane prior to decompression. (Ovsiannikov [0088]; “More specifically, unpacking packed streams involves first determining lengths of original streams, which is readily available from zero bit mask 303 at the beginning of a packed block, followed by reproducing the calculations performed during packing when stream lengths within pairs (or pairs-of-pairs, or pairs-of-pairs-of-pairs) were compared to determine which head of which stream is to be cropped-and-appended to which stream. At that time, the difference between stream lengths was determined or calculated, divided by two to determine the length of the head to be cropped-and-appended”, wherein the head reads on a value corresponding to the slice size; wherein the value is to be recorded before the “cropped-and-appended” decompression is to occur,
Ovsiannikov [0011]; “At 202 in FIG. 2A, a leading portion, or head, of the longer bit-stream length of each pair is relocated, or redirected, through the butterfly shuffler 101 to be part of the shorter bit stream of the pair by controlling the multiplexers of the pair so that the pair of bit streams has equal bit-stream lengths. For example, a portion of bit stream 0 is redirected by the multiplexers of column C=0 to become part of bit stream 1. Similarly, a portion of bit stream 2 is redirected to become part of bit stream 3. A portion of bit stream 4 is redirected to bit stream 5, and a portion of bit stream 7 is directed to be part of bit stream 6. In situations in which the difference in bit-stream lengths of a pair of bit streams is an odd number of bits, a dummy, or filler, bit may be added to the shorter of the two bit streams”, wherein the head of the larger bit stream recorded into the smaller bit stream reads on a value corresponding to the slice size in the second lane)
Regarding Claim 3,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 2 (and thus the rejection of Claim 2 is incorporated). Ovsiannikov/Ko/Boesch further discloses wherein the value corresponding to the slice size is indicative of a size of the portion of the first data in the second data lane. (Ovsiannikov [0011]; “At 202 in FIG. 2A, a leading portion, or head, of the longer bit-stream length of each pair is relocated, or redirected, through the butterfly shuffler 101 to be part of the shorter bit stream of the pair by controlling the multiplexers of the pair so that the pair of bit streams has equal bit-stream lengths. For example, a portion of bit stream 0 is redirected by the multiplexers of column C=0 to become part of bit stream 1. Similarly, a portion of bit stream 2 is redirected to become part of bit stream 3. A portion of bit stream 4 is redirected to bit stream 5, and a portion of bit stream 7 is directed to be part of bit stream 6. In situations in which the difference in bit-stream lengths of a pair of bit streams is an odd number of bits, a dummy, or filler, bit may be added to the shorter of the two bit streams”)
Regarding Claim 4,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 1 (and thus the rejection of Claim 1 is incorporated). Ovsiannikov/Ko/Boesch further discloses wherein the slice size is within 50 bytes of half of the size difference between the first data lane and the second data lane. (Ovsiannikov [0088]; “At that time, the difference between stream lengths was determined or calculated, divided by two to determine the length of the head to be cropped-and-appended” thus the slice size is exactly half the size difference)
Regarding Claim 5,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 1 (and thus the rejection of Claim 1 is incorporated). Ovsiannikov/Ko/Boesch further discloses wherein the portion of the first data is positioned after the second data in the second data lane. (Ovsiannikov Fig. 2A, displaying pairs of bit streams in which the smaller bit stream has the first data head appended after its second data)
Regarding Claim 9,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 1 (and thus the rejection of Claim 1 is incorporated). Ovsiannikov/Ko/Boesch further discloses wherein the portion of the first data is cut from an end of the first data lane. (Ovsiannikov Fig. 2A, displaying pairs of bit streams in which the head (portion of the first data) of the larger bit stream is removed from its end)
Regarding Claim 10,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 1 (and thus the rejection of Claim 1 is incorporated). Ovsiannikov/Ko/Boesch further discloses wherein the first data lane is positioned adjacent the second data lane in the partition. (Ovsiannikov Fig. 2A Element 203, displaying both the new first bit stream with the head portion cut and the second bit stream with the head portion of the first data appended, with both streams positioned adjacently in the partition pair)
Regarding Claim 11,
Claim 11 recites at least one non-transitory computer readable medium having the same instructions as the apparatus of Claim 1, and is thus rejected for reasons set forth in the rejection of Claim 1.
Regarding Claim 13,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 11 (and thus the rejection of Claim 11 is incorporated). Ovsiannikov/Ko/Boesch further discloses to cause the one or more processors to write a value corresponding to the slice size in the second data lane. (Ovsiannikov [0011]; “At 202 in FIG. 2A, a leading portion, or head, of the longer bit-stream length of each pair is relocated, or redirected, through the butterfly shuffler 101 to be part of the shorter bit stream of the pair by controlling the multiplexers of the pair so that the pair of bit streams has equal bit-stream lengths. For example, a portion of bit stream 0 is redirected by the multiplexers of column C=0 to become part of bit stream 1. Similarly, a portion of bit stream 2 is redirected to become part of bit stream 3. A portion of bit stream 4 is redirected to bit stream 5, and a portion of bit stream 7 is directed to be part of bit stream 6. In situations in which the difference in bit-stream lengths of a pair of bit streams is an odd number of bits, a dummy, or filler, bit may be added to the shorter of the two bit streams”)
Regarding Claim 14,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 11 (and thus the rejection of Claim 11 is incorporated). Ovsiannikov/Ko/Boesch further discloses to cause the one or more processors to cut the portion of the first data from a first end of the first data lane. (Ovsiannikov Fig. 2A, displaying pairs of bit streams in which the head of the larger bit stream is removed from its end)
Regarding Claim 15,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 14 (and thus the rejection of Claim 14 is incorporated). Ovsiannikov/Ko/Boesch further discloses to cause the one or more processors to append the portion of the first data to a second end of the second data lane. (Ovsiannikov Fig. 2A Element 203, displaying both the new first bit stream with the head portion cut and the second data lane with the head portion of the first data appended at its end)
Regarding Claim 16,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 11 (and thus the rejection of Claim 11 is incorporated). Ovsiannikov/Ko/Boesch further discloses wherein in response to appending the portion of the first data to the second data lane, the first data lane includes a first size and the second data lane includes a second size within 200 bytes of the first size. (Ovsiannikov [0012]; “FIGS. 2A-2C conceptually depict eight example bit streams of different bit-stream lengths being recursively packed to become eight bit streams each having equal bit-stream lengths according to the subject matter disclosed herein”, wherein equal bit-stream lengths reads on a first and second data lane of sizes within 200 bytes)
Regarding Claim 17,
Claim 17 recites a method comprising the same instructions performed by the apparatus of Claim 1, and is thus rejected for reasons set forth in the rejection of Claim 1.
Regarding Claim 19,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 17 (and thus the rejection of Claim 17 is incorporated). Ovsiannikov/Ko/Boesch further discloses identifying the first data lane in response to determining the first data has a smallest size in the partition. (Ovsiannikov [0078]; “At 202 in FIG. 2A, a leading portion, or head, of the longer bit-stream length of each pair is relocated, or redirected, through the butterfly shuffler 101 to be part of the shorter bit stream of the pair by controlling the multiplexers of the pair so that the pair of bit streams has equal bit-stream lengths. For example, a portion of bit stream 0 is redirected by the multiplexers of column C=0 to become part of bit stream 1.”, in which one stream is of a smaller size than the other stream, thus the smaller bit stream is read on as the identification of a first data lane)
and identifying the second data lane in response to determining the second data has a largest size in the partition. (Ovsiannikov [0078]; “At 202 in FIG. 2A, a leading portion, or head, of the longer bit-stream length of each pair is relocated, or redirected, through the butterfly shuffler 101 to be part of the shorter bit stream of the pair by controlling the multiplexers of the pair so that the pair of bit streams has equal bit-stream lengths. For example, a portion of bit stream 0 is redirected by the multiplexers of column C=0 to become part of bit stream 1.”, in which one stream is of a larger size than the other stream, thus the larger bit stream is read on as the identification of a second data lane)
Regarding Claim 22,
Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 17 (and thus the rejection of Claim 17 is incorporated). Ovsiannikov/Ko/Boesch further discloses recording the slice size or a value corresponding to the slice size in the second data lane. (Ovsiannikov [0088]; “More specifically, unpacking packed streams involves first determining lengths of original streams, which is readily available from zero bit mask 303 at the beginning of a packed block, followed by reproducing the calculations performed during packing when stream lengths within pairs (or pairs-of-pairs, or pairs-of-pairs-of-pairs) were compared to determine which head of which stream is to be cropped-and-appended to which stream. At that time, the difference between stream lengths was determined or calculated, divided by two to determine the length of the head to be cropped-and-appended … The calculations provide offsets into packed streams pointing to where each cropped-and-appended head may be located in storage.”, wherein the head reads on a value corresponding to the slice size, and the cropped-and-appended head located in storage reads on recording the slice size
Claims 6-8, 12 and 21 are rejected under 35 U.S.C. 103 as being anticipated by Ovsiannikov et al. (US20200336272 A1, hereinafter “Ovsiannikov”) in view of Ko et al. (KR20210053791 A, hereinafter “Ko”) further in view of Boesch et al. (US20210397933A1, hereinafter “Boesch”) and further in view of Hans et al. (US20110307659 A1, hereinafter “Hans”)
Regarding Claim 6,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 1 (and thus the rejection of Claim 1 is incorporated). The combination of Ovsiannikov/Ko/Boesch fails to disclose but Hans discloses wherein the processor circuitry is to execute the instructions to: assign a first identifier to the first data lane; (Hans [0134]; “In at least some example embodiments, multiple hash codes are concurrently generated in parallel using data within multiple windows over different portions of the incoming chunk data stream. FIG. 12 shows an example using two moving windows of three bytes each, each window defining a data lane within chunk byte stream 1200. For a minimum code word size of two bytes, three bytes is the minimum window size that can be used that produces a compression of the data (i.e., a reduction of at least one byte).”,
[0158]; “The data for each lane is identified by the least recent byte of the three bytes within the lane window”)
and assign a second identifier to the second data lane, the second identifier different from the first identifier. (Hans [0136]; one window defines lane 0 (Data0), which includes data bytes B.sub.8 (the first byte of the 3-byte sequence that includes bits 0 through 23) through B.sub.10 (the last byte of the 3-byte sequence). Similarly, a second window defines lane 1 (Data 1), which includes data bytes B.sub.9 through B.sub.11 (bits 8-31). Because both lanes are processed concurrently in parallel, for each processing cycle the processed byte stream is shifted by two bytes, and two new bytes are loaded. Thus, in the next cycle after that shown in FIG. 12 data lanes 0 and 1 will include data bytes B.sub.10-B.sub.12 and B.sub.11-B.sub.13 respectively.”,
[0158]; “The data for each lane is identified by the least recent byte of the three bytes within the lane window”, wherein a second window reads on a different corresponding least recent byte, thus a second identifier differing from the first identifier)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the combination of Ovsiannikov/Ko/Boesch, which performs bit-manipulation across neural network data lanes, to implement Hans’ method of additionally assigning identifiers to each data lane. The motivation to do so is to clearly distinguish between data lanes in data operations and “to uniquely identify each chunk” (Hans [0099]).
Regarding Claim 7,
The combination of Ovsiannikov/Ko/Boesch/Hans teaches the apparatus of Claim 6 (and thus the rejection of Claim 6 is incorporated). The combination of Ovsiannikov/Ko/Boesch/Hans already inherently discloses wherein the first identifier is a first header byte recorded in the first data lane and the second identifier is a second header byte recorded in the second data lane. (Hans [0134]; “In at least some example embodiments, multiple hash codes are concurrently generated in parallel using data within multiple windows over different portions of the incoming chunk data stream. FIG. 12 shows an example using two moving windows of three bytes each, each window defining a data lane within chunk byte stream 1200. For a minimum code word size of two bytes, three bytes is the minimum window size that can be used that produces a compression of the data (i.e., a reduction of at least one byte).”,
[0136]; one window defines lane 0 (Data0), which includes data bytes B.sub.8 (the first byte of the 3-byte sequence that includes bits 0 through 23) through B.sub.10 (the last byte of the 3-byte sequence). Similarly, a second window defines lane 1 (Data 1), which includes data bytes B.sub.9 through B.sub.11 (bits 8-31). Because both lanes are processed concurrently in parallel, for each processing cycle the processed byte stream is shifted by two bytes, and two new bytes are loaded. Thus, in the next cycle after that shown in FIG. 12 data lanes 0 and 1 will include data bytes B.sub.10-B.sub.12 and B.sub.11-B.sub.13 respectively.”,
[0158]; “The data for each lane is identified by the least recent byte of the three bytes within the lane window”)
Regarding Claim 8,
The combination of Ovsiannikov/Ko/Boesch/Hans teaches the apparatus of Claim 6 (and thus the rejection of Claim 6 is incorporated). The combination of Ovsiannikov/Ko/Boesch/Hans already inherently discloses wherein the processor circuitry is to execute the instructions to assign a third identifier to the third data lane, the third data of a smaller size than the first data and a larger size than the second data. (Ovsiannikov Fig. 2A, which discloses a third data lane,
Hans [0134]; “In at least some example embodiments, multiple hash codes are concurrently generated in parallel using data within multiple windows over different portions of the incoming chunk data stream. FIG. 12 shows an example using two moving windows of three bytes each, each window defining a data lane within chunk byte stream 1200. For a minimum code word size of two bytes, three bytes is the minimum window size that can be used that produces a compression of the data (i.e., a reduction of at least one byte).”, which discloses windows read on as identifiers for each data lane,
[0136]; one window defines lane 0 (Data0), which includes data bytes B.sub.8 (the first byte of the 3-byte sequence that includes bits 0 through 23) through B.sub.10 (the last byte of the 3-byte sequence). Similarly, a second window defines lane 1 (Data 1), which includes data bytes B.sub.9 through B.sub.11 (bits 8-31). Because both lanes are processed concurrently in parallel, for each processing cycle the processed byte stream is shifted by two bytes, and two new bytes are loaded. Thus, in the next cycle after that shown in FIG. 12 data lanes 0 and 1 will include data bytes B.sub.10-B.sub.12 and B.sub.11-B.sub.13 respectively.”, wherein the disclosure of three bytes being the minimum data lane size reads on data lanes of different sizes including a third data of a smaller size than the first data lane and a larger size than a second data lane.)
Regarding Claim 12,
The combination of Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 11 (and thus the rejection of Claim 11 is incorporated). The combination of Ovsiannikov/Ko/Boesch fails to disclose but Hans discloses to write a first identifier at a first end of the first data lane, the first identifier indicative of data to be removed from the first data lane. (Hans [0136]; one window defines lane 0 (Data0), which includes data bytes B.sub.8 (the first byte of the 3-byte sequence that includes bits 0 through 23) through B.sub.10 (the last byte of the 3-byte sequence). Similarly, a second window defines lane 1 (Data 1), which includes data bytes B.sub.9 through B.sub.11 (bits 8-31). Because both lanes are processed concurrently in parallel, for each processing cycle the processed byte stream is shifted by two bytes, and two new bytes are loaded. Thus, in the next cycle after that shown in FIG. 12 data lanes 0 and 1 will include data bytes B.sub.10-B.sub.12 and B.sub.11-B.sub.13 respectively.”,
[0158]; “The data for each lane is identified by the least recent byte of the three bytes within the lane window”,
[0183]; ”Once the data frames begin to arrive at an HAA-1 module input port, hardware within the HAA-1 module subdivides the incoming frames into chunks, calculates chunk identifiers on the fly for each chunk, and compresses and stores the chunks in memory for later retrieval. As the processing of each chunk is completed, information for each corresponding chunk, including the chunk identifier generated by the HAA-1 module, is forwarded to the HAA-2 module for further processing.”, wherein the information corresponding to each chunk reads on an indicator of data to be removed)
write a second identifier at a second end of the second data lane, the second identifier indicative of data to be added to the second data lane. (Hans [0158]; “The data for each lane is identified by the least recent byte of the three bytes within the lane window”,
[0136]; one window defines lane 0 (Data0), which includes data bytes B.sub.8 (the first byte of the 3-byte sequence that includes bits 0 through 23) through B.sub.10 (the last byte of the 3-byte sequence). Similarly, a second window defines lane 1 (Data 1), which includes data bytes B.sub.9 through B.sub.11 (bits 8-31). Because both lanes are processed concurrently in parallel, for each processing cycle the processed byte stream is shifted by two bytes, and two new bytes are loaded. Thus, in the next cycle after that shown in FIG. 12 data lanes 0 and 1 will include data bytes B.sub.10-B.sub.12 and B.sub.11-B.sub.13 respectively.”, wherein a second window reads on a different corresponding least recent byte, thus a second identifier differing from the first identifier,
[0183]; ”Once the data frames begin to arrive at an HAA-1 module input port, hardware within the HAA-1 module subdivides the incoming frames into chunks, calculates chunk identifiers on the fly for each chunk, and compresses and stores the chunks in memory for later retrieval. As the processing of each chunk is completed, information for each corresponding chunk, including the chunk identifier generated by the HAA-1 module, is forwarded to the HAA-2 module for further processing.”, wherein the information corresponding to each chunk reads on an indicator of data to be appended)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the combination of Ovsiannikov/Ko/Boesch, which performs bit-manipulation across neural network data lanes, to implement Hans’ method of additionally assigning semantic identifiers indicative of data status to each data lane. The motivation to do so is to clearly distinguish between data lanes in data operations and “to uniquely identify each chunk” (Hans [0099]).
Regarding Claim 21,
Ovsiannikov/Ko/Boesch teaches the apparatus of Claim 17 (and thus the rejection of Claim 17 is incorporated). Ovsiannikov/Ko/Boesch fails to disclose but Hans discloses assigning a first identifier to a first end of the first data lane, the first identifier indicative of data to be removed from the first data lane; (Hans [0136]; one window defines lane 0 (Data0), which includes data bytes B.sub.8 (the first byte of the 3-byte sequence that includes bits 0 through 23) through B.sub.10 (the last byte of the 3-byte sequence). Similarly, a second window defines lane 1 (Data 1), which includes data bytes B.sub.9 through B.sub.11 (bits 8-31). Because both lanes are processed concurrently in parallel, for each processing cycle the processed byte stream is shifted by two bytes, and two new bytes are loaded. Thus, in the next cycle after that shown in FIG. 12 data lanes 0 and 1 will include data bytes B.sub.10-B.sub.12 and B.sub.11-B.sub.13 respectively.”,
[0158]; “The data for each lane is identified by the least recent byte of the three bytes within the lane window”,
[0183]; ”Once the data frames begin to arrive at an HAA-1 module input port, hardware within the HAA-1 module subdivides the incoming frames into chunks, calculates chunk identifiers on the fly for each chunk, and compresses and stores the chunks in memory for later retrieval. As the processing of each chunk is completed, information for each corresponding chunk, including the chunk identifier generated by the HAA-1 module, is forwarded to the HAA-2 module for further processing.”, wherein the information corresponding to each chunk reads on an indicator of data to be removed)
and assigning a second identifier to a second end of the second data lane, the second identifier indicative of data to be appended to the second data lane. (Hans [0158]; “The data for each lane is identified by the least recent byte of the three bytes within the lane window”,
[0136]; one window defines lane 0 (Data0), which includes data bytes B.sub.8 (the first byte of the 3-byte sequence that includes bits 0 through 23) through B.sub.10 (the last byte of the 3-byte sequence). Similarly, a second window defines lane 1 (Data 1), which includes data bytes B.sub.9 through B.sub.11 (bits 8-31). Because both lanes are processed concurrently in parallel, for each processing cycle the processed byte stream is shifted by two bytes, and two new bytes are loaded. Thus, in the next cycle after that shown in FIG. 12 data lanes 0 and 1 will include data bytes B.sub.10-B.sub.12 and B.sub.11-B.sub.13 respectively.”, wherein a second window reads on a different corresponding least recent byte, thus a second identifier differing from the first identifier,
[0183]; ”Once the data frames begin to arrive at an HAA-1 module input port, hardware within the HAA-1 module subdivides the incoming frames into chunks, calculates chunk identifiers on the fly for each chunk, and compresses and stores the chunks in memory for later retrieval. As the processing of each chunk is completed, information for each corresponding chunk, including the chunk identifier generated by the HAA-1 module, is forwarded to the HAA-2 module for further processing.”, wherein the information corresponding to each chunk reads on an indicator of data to be appended)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of Ovsiannikov/Ko/Boesch, which performs bit-manipulation across data bit streams, to implement Hans’ method of additionally assigning semantic identifiers indicative of data status to each data lane. The motivation to do so is to clearly distinguish between data lanes in data operations and “to uniquely identify each chunk” (Hans [0099]).
Response to Arguments
The Examiner acknowledges the Applicant’s amendments to Claims 1, 8, 11 and 17.
Applicant’s amendments filed February 27th, 2026, considering the claim objection to Claim 17, have been fully considered and are fully persuasive.
Applicant’s arguments filed February 27th, 2026, traversing the rejection of claims 1-17, 19, 21-22 under 35 U.S.C. § 101 have been fully considered and are fully persuasive.
Applicant’s arguments regarding the 35 U.S.C. § 103 rejection of claims 1-17, 19, 21-22 of the previous office action have been considered, have been fully considered, but are not fully persuasive.
Applicant alleges, on Pages 10-12 of Remarks, that since Ovsiannikov identifies the lengths of earlier streams instead of current streams to be processed, and the proposed claim element recites determining a slice size based on present data lanes with different memory sizes, the cited Ovsiannikov portion can’t teach or suggest the claim element of “determine a slice size based on a memory size difference between the first and second data lanes”. Additionally, applicant alleges that Ovsiannikov describes a comparison between two data streams within a pair, and thus can’t teach or suggest these size comparisons among at least three data lanes.
Examiner respectfully disagrees. Although examiner concedes that applicant’s invention as described in the specification does meaningfully differ from the scope of Ovsiannikov in intent, such differences are not meaningfully conveyed in the broadest reasononable interpretable scope of the currently proposed claim language. Applicant’s claim language does not present any language regarding the first, second, and third data lanes that indicate that the data lanes must be “present” data lanes. There is no indication in applicant’s claim language that Ovisannikov’s relocating of future bits in one bit stream to another cannot be interpreted as applicant’s removal of bits in one stream to another bit stream. By nature, relocating bits from one stream to another is reasonable interpretable as removal of the bits from a current stream location and its placement in another stream location. As such, Ovsiannikov’s disclosure of bitwise comparisons performed between earlier data streams falls under the broadest reasonable interpretation of applicant’s claim language that fails to positively recite the nature of such bit streams.
Regarding applicant’s arguments about Ovsiannikov failing to describe such size comparisons among at least three lanes, examiner maintains that Ovsiannikov’s disclosure of bitwise appending/cropping operations in partition including 4 pairs of bitstreams is sufficient for applicant’s current claim scope. Applicant only describes bitwise operations of “removing” and “appending” and “determining a slice size” between a first and second data lane. Importantly, the third data lane’s only current relation to the first and second data lanes is that its size is smaller than the first data lane and larger than the second data lane, while maintaining its size during aforementioned bitwise operations between the first and second data lanes. Moreso, Ovsiannikov’s bit stream operations being performed in pairwise calculations indicate that such calculations are not performed all at once, but rather updated in pairs within the partition. As such, it is reasonable to interpret such pairwise calculations as teaching bitwise operations between the first and second data lanes while a third lane (of a different pair within the partition of four pairs of bit streams) remains idle and thus maintains its size. Ovsiannikov’s disclosure of bitwise operations between pairs of streams within a partition of 8 bit streams, as such, thus discloses a first, second, and third data lane in line with the broadest reasonable interpretation of the applicant’s claim language due to the very broad scope of the third data lane’s functionality and its functional relationship with the first and second data lanes aside from simple size comparisons.
The rejection of Claim 1 under 35 U.S.C. § 103 has been maintained. Similarly, the rejection of Claims 11 and 17 under 35 U.S.C. § 103 have been maintained.
The rejection of Claims 2-10 under 35 U.S.C. § 103, which depend directly or indirectly from Claim 1, have been maintained.
The rejection of Claims 12-16, which depend directly or indirectly from Claim 11, have been maintained.
The rejection of Claims 19, 21, and 22, which depend directly or indirectly from Claim 17, have been maintained.
Conclusion
Applicant’s amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN J KIM whose telephone number is (571)272-0523. The examiner can normally be reached 8-6.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt El can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONATHAN J KIM/Examiner, Art Unit 2141
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141