Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-6 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Desoli et al. (US 2023/0062910 A1).
As per claim 1, Desoli et al. discloses a computation circuit (101) comprising:
a plurality of first operation circuits; ([0038], [0042], [0043], "Operation circuit" (e.g., convolution, pooling, matrix multiplication) with its own requantization (quantization) circuit is described repeatedly);
a plurality of quantization circuits configured to quantize outputs of the plurality of first operation circuits, respectively (See Fig. 1: multiple operation circuits, each followed by a requantization (quantization) circuit.);
a plurality of second operation circuits configured to perform operations on outputs of the plurality of quantization circuits, respectively; ([0042], [0044], [0050]; describes that operation circuits and requantization circuits may be arranged in sequence or parallel, and data flows from one to the next (e.g., from convolution to pooling, etc.), forming a chain of operations ("second operation circuits")); and
an adder circuit configured to perform element wise addition operation on outputs of the plurality of second operation circuits. ([0044], [0052]; adder/combiner circuits are shown for combining outputs (e.g., for residual/skip connections, element-wise addition of parallel branches). See Figs. 1-4 (diagrams show multiple operation circuits, each with a quantizer, connected to further operation circuits, and outputs routed to an adder/combiner).
As per claim 2, Desoli et al. discloses wherein each the plurality of first operation circuits performs a convolution operation, a bottleneck operation, a max pooling operation, or a matrix multiplication operation. ([0043], [0050], Figs. 2–4; operation circuits can be convolution, pooling, matrix multiplication, etc.; via flexible architecture).
As per claim 3, Desoli et al. discloses wherein each the plurality of second operation circuits performs a linear operation. ([0044], [0050]; second operation circuits can be linear (e.g., convolution, matrix multiplication)).
As per claim 4, Desoli et al. discloses wherein the linear operation is a convolution operation or a matrix multiplication operation. ([0050]; via explicitly describing convolution and matrix multiplication as supported operations).
As per claim 5, Desoli et al. discloses wherein respective kernels of the linear operations performed by the plurality of second operation circuits are the same. ([0052]; supports use of same kernel/weights across multiple operation circuits (e.g., grouped convolution)).
As per claim 6, Desoli et al. discloses wherein each the plurality of first operation circuits corresponds to a layer in a neural network including a plurality of layers. ([0036]–[0042], Figs. 1–3; Each operation circuit (with quantizer) can correspond to a layer in a neural network).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zhang et al., US 20210216871 A1, describes a systems and methods for efficient CNN computation using sparse and quantized weights. A location vector (LV) table records non-zero weight coordinates, and a lookup table recovers real weight values from quantized IDs. Convolution is performed by accumulating products of input activations and quantized weights, skipping zeros. The focus is on memory and computation reduction via sparsity and quantization.
Kim, US 20210264232 A1, describes a methods and systems for bit quantization of neural network parameters (weights, activations, etc.) at the layer or parameter group level. Bit-widths are selected per parameter/layer/group to optimize memory/computation while maintaining accuracy.
Tomida et al., US 2021/0319294 A1, describes a neural network circuit for embedded devices with a looped structure: convolution operation circuit → quantization operation circuit → back to convolution, with memory units in between. Quantization is performed after each convolution operation, and the quantized data is used as the next input. Tomika et al. further supports parallelization by partitioning tensors.
Wen et al., US 11514136 B2, describes a Circuit For Neural Network Convolutional Calculation Of Variable Feature And Kernel Sizes focusing on hardware for parallel convolution with variable feature/kernel sizes, feature/kernel managers, and row convolution processors.
Tomida et al., WO 2021210527 A1, this invention describes a control method and architecture for a neural network circuit, particularly for low-bit (quantized) convolutional neural networks (CNNs) used in edge devices.
Menkhoff et al., US 20150280724 A1, describes a quantization circuit (100) includes a quantizer (102) configured to provide a quantized sample using an input quantity and an error estimator (110) configured to determine a quantization error of the quantized sample. An error corrector (120) is configured to correct the quantized sample by a correction value depending on the quantization error.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Lynda Jasmin whose telephone number is (571)272-6782. The examiner can normally be reached M-F: 8:00AM-6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Lynda Jasmin can be reached at 5712726782. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LYNDA JASMIN/Supervisory Patent Examiner, Art Unit 3629