DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on July 5, 2023; January 3, 2024; February 7, 2024; October 4, 2024; February 26, 2025; May 6, 2026, were considered by the examiner. The submissions are in compliance with the provisions of 37 CFR 1.97.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims 19 - 23 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (math concept) without significantly more.
Claim 19:
Regarding claim 19, in step 1 of the 101-analysis set forth in MPEP2106, the claim recites “A method comprising: performing vector-by-matrix multiplication operations in a first layer of a neural network using a digital computation-in-memory engine; and performing vector-by-matrix multiplication operations in a second layer different than the first layer of the neural network using an analog computation-in-memory engine”, and a method is one of the four statuary categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components:
A method comprising: performing vector-by-matrix multiplication operations in a first layer of a neural network … (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0038, 0040] from the specification
PNG
media_image1.png
517
1016
media_image1.png
Greyscale
PNG
media_image2.png
210
1002
media_image2.png
Greyscale
stating using equations to calculate vector matrix multiplication operations, see MPEP 2106.04(a)(2), subsection I),
and performing vector-by-matrix multiplication operations in a second layer different than the first layer of the neural network … (This recites a mathematical relationship, math formula or equation, or mathematical calculation, see above in paragraphs [0038, 0040] from the specification, stating using equations to calculate vector matrix multiplication operations, see MPEP 2106.04(a)(2), subsection I),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
… using a digital computation-in-memory engine, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
… using an analog computation-in-memory engine, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements iii and iv recite using generic computer components as tools, which are not indicative of significantly more.
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 20:
Regarding claim 20, it is dependent upon claim 19, and thereby incorporates the limitations of, and corresponding analysis applied to claim 19. Further, claim 20 recites the following additional element:
The method of claim 19, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns. (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 21:
Regarding claim 21, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites
“A method comprising: performing vector-by-matrix multiplication operations in a first layer of a neural network using a digital computation-in-memory engine; performing vector-by-matrix multiplication operations in a second layer of the neural network using an analog computation-in-memory engine; and performing vector-by-matrix multiplication operations in a third layer of the neural network using a dynamic weight engine,” and a method is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components:
performing vector-by-matrix multiplication operations in a first layer of a neural network … (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0038, 0040] from the specification
PNG
media_image1.png
517
1016
media_image1.png
Greyscale
PNG
media_image2.png
210
1002
media_image2.png
Greyscale
stating using equations to calculate vector matrix multiplication operation, see MPEP 2106.04(a)(2), subsection I),
performing vector-by-matrix multiplication operations in a second layer of the neural network … (This recites a mathematical relationship, math formula or equation, or math calculation, see above in paragraphs [0038, 0040] from the specification using equations to calculate vector matrix multiplication operations, see MPEP 2106.04(a)(2), subsection I),
and performing vector-by-matrix multiplication operations in a third layer of the neural network … (This recites a mathematical relationship, math formula or equation, or mathematical calculation, see above in paragraphs [0038, 0040] from the specification stating using equations to calculate vector matrix multiplication operations, see MPEP 2106.04(a)(2), subsection I),
If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
… using a digital computation-in-memory engine, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
… using an analog computation-in-memory engine, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
… using a dynamic weight engine, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements iv, v, and vi recite using generic computer components as tools, which are not indicative of significantly more.
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 22:
Regarding claim 22, it is dependent upon claim 21, and thereby incorporates the limitations of, and corresponding analysis applied to claim 21. Further, claim 22 recites the following additional element:
The method of claim 21, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 23:
Regarding claim 23, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites
“A method comprising: performing vector-by-matrix multiplication operations in a first layer of a neural network using one of an analog computation-in-memory engine, a digital computation-in-memory engine, and a dynamic weight engine; and performing vector-by-matrix multiplication operations in a second layer of a neural network using another of the analog computation-in-memory engine, the digital computation-in-memory engine, and the dynamic weight engine,” and a method is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components:
A method comprising: performing vector-by-matrix multiplication operations in a first layer of a neural network…; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0038, 0040] from the specification
PNG
media_image1.png
517
1016
media_image1.png
Greyscale
PNG
media_image2.png
210
1002
media_image2.png
Greyscale
stating using equations to calculate vector matrix multiplication operations, see MPEP 2106.04(a)(2), subsection I),
and performing vector-by-matrix multiplication operations in a second layer of a neural network …, (This recites a mathematical relationship, math formula or equation, or math calculation, see above in paragraphs [0038, 0040] from the specification using equations to calculate vector matrix multiplication operations, see MPEP 2106.04(a)(2), subsection I),
If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
… using one of an analog computation-in-memory engine, a digital computation-in-memory engine, and a dynamic weight engine, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
… using another of the analog computation-in-memory engine, the digital computation-in-memory engine, and the dynamic weight engine, (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements iii and iv recite using generic computer components as a tool, which are not indicative of significantly more.
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claims 1 and 9 are rejected under 35 U.S.C. 103 over Chen, J., et al. in “A charge-digital hybrid compute-in-memory macro with full precision 8-bit multiply-accumulation for edge computing devices,” published on December 19-22, 2022 for a conference, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10008468 , (hereafter, Chen) in view of Badaroglu, M. et al., in PG Pub. No. US20230025068, published on January 26, 2023, (hereafter, Badaroglu).
Claim 1:
Regarding claim 1, Chen teaches “1. A system comprising: an analog computation-in-memory engine to perform operations in a first layer in a neural network;”
See Chen in page 153, from abstract “Compute-in-memory (CIM) is emerging as a new computing architecture to overcome the high energy consumption of edge-side AI and IoT devices. When performing high-precision neural network calculations, analog CIM and digital CIM have their own advantages and disadvantages. In this paper, we combine the advantages of high energy efficiency of analog CIM and high accuracy of digital CIM to propose a charge-digital hybrid CIM (CDH-CIM) macro.... Simulation shows that the macro achieves 6.98~11.0 TOPS/W at 0.8V and 71.92% inference accuracy when performing CIFAR-100 dataset.” Chen shows using a hybrid of using both analog and digital CIM systems into the same system for performing calculations of neural network models, includes both analog CIM and digital CIM systems.
Further, see Chen in page 157, section B. Network Mapping describe “Figure 7 shows the mapping relationship of the MAC operations in the convolutional layer implemented in the CDH-CIM macro. For CNN operations, the vast majority of the computation is in the MAC operations. The figure shows the computation of the convolution of a 56∗56∗64 feature map of with 64 3∗3∗64 filters. The convolution operation of each filter with the input feature map gets the outputs of one layer and the convolution operation of 64 filters gets the output of 64 layers. The size and depth of the feature map is kept constant by setting the number of fi[l]ters and using padding for the input feature map.” Here, Chen shows that each hardware CIM component of the CDH-CIM macro calculates for one layer of the neural network model.
Further, Chen teaches “and a digital computation-in-memory engine to perform operations in a second layer different than the first layer in the neural network,”
See Chen in page 153, from abstract “Compute-in-memory (CIM) is emerging as a new computing architecture to overcome the high energy consumption of edge-side AI and IoT devices. When performing high-precision neural network calculations, analog CIM and digital CIM have their own advantages and disadvantages. In this paper, we combine the advantages of high energy efficiency of analog CIM and high accuracy of digital CIM to propose a charge-digital hybrid CIM (CDH-CIM) macro.... Simulation shows that the macro achieves 6.98~11.0 TOPS/W at 0.8V and 71.92% inference accuracy when performing CIFAR-100 dataset.” Chen shows using a hybrid of using both analog and digital CIM systems into the same system for performing calculations of neural network models, includes both analog CIM and digital CIM systems.
However, Chen did not teach “and a digital computation-in-memory engine to perform operations in a second layer different than the first layer in the neural network.”
In an analogous system, Badaroglu teaches “and a digital computation-in-memory engine to perform operations in a second layer different than the first layer in the neural network.”
See Badaroglu in [0028] mention “Some aspects provide a hybrid neural network architecture using both compute-in-memory (CIM) and neural processing unit (NPU) processing elements (PEs), where the CIM PEs and the NPU PEs can share resources (e.g., memory), can concurrently operate, and can transfer data from one type of PE to another type of PE within the same neural network layer or in different neural network layers (e.g., adjacent layers).” Badaroglu also mentions for each CIM system, the performance operations work in a second (or another) layer different from a first neural network layer within the same neural network model.
Further, see Badaroglu in [0059] describe "This processing may be repeated for each layer of the image data, and the outputs (e.g., output activations) may be stored in the output buffers and then used by the mobile device for an ML/AI task, such as facial recognition." In addition, Badaroglu mentions that the ‘processing may be repeated for each layer’ relates to performing calculations in subsequent layers like first, second, third or additional layers of a neural network model.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Chen and incorporate into the teachings of Badaroglu because both references teach using a type of computation-in-memory system to perform operations within a neural network layer.
One of ordinary skill in the art would be motivated to do so because “CIM units are generally better than NPUs in terms of energy efficiency, whereas NPUs are generally better than CIM units for depth-wise convolution. Due to bit-serial operation, digital compute-in-memory (DCIM) units may have lower TOPS, but comparable or better performance for a given area (e.g., in terms of TOPS/mm2) than NPUs,” (see Badaroglu in [0082]).
Claim 9:
Regarding claim 9, Chen in view of Badaroglu, teach the limitations of claim 1.
Further, Badaroglu teaches “The system of claim 7, wherein the digital computation-in-memory engine comprises a shift and adder tree coupled to the plurality of CIM digital cells.
See Badaroglu in [0065 - 0066] describe “As shown, the DCIM circuit 400 may include a bit-column adder tree 409, which may include eight adder trees 4100 to 4107 (collectively referred to as “adder trees 410”), each adder tree being implemented for a respective one of the columns 406. Each of the adder trees 410 adds the output signals from the CIM cells 402 on the respective one of the columns 406, and the adder trees 410 may operate in parallel (e.g., concurrently). The outputs of the adder trees 410 may be coupled to a weight-shift adder tree circuit 412, as shown. The weight-shift adder tree circuit 412 includes multiple weight-shift adders 414, each including a bit-shift-and-add circuit to facilitate the performance of a bit-shifting-and-addition operation. In other words, the CIM cells on column 4060 may store the most-significant bits (MSBs) for respective weights on each word-line 404, and the CIM cells on column 4067 may store the least-significant bits (LSBs) for respective weights on each word-line. Therefore, when performing the addition across the columns 406, a bit-shift operation is performed to shift the bits to account for the significance of the bits on the associated column. [0066] The output of the weight-shift adder tree circuit 412 is provided to an activation-shift accumulator circuit 416. The activation-shift accumulator circuit 416 includes a bit-shift circuit 418, a serial accumulator 420, and a flip-flop (FF) array 422. For example, the FF array 422 may be used to implement a register.” Here, Badaroglu describes using a weight shift adder tree circuit 412 that is part of the digital CIM or DCIM circuit unit 400.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Chen and incorporate into the teachings of Badaroglu because both references teach using a type of computation-in-memory system to perform operations within a neural network layer using shift and adder trees.
One of ordinary skill in the art would be motivated to do so because “CIM units are generally better than NPUs in terms of energy efficiency, whereas NPUs are generally better than CIM units for depth-wise convolution. Due to bit-serial operation, digital compute-in-memory (DCIM) units may have lower TOPS, but comparable or better performance for a given area (e.g., in terms of TOPS/mm2) than NPUs,” (see Badaroglu in [0082]).
Claims 2, 6, and 7 are rejected under 35 U.S.C. 103 over Chen in view of Badaroglu, further in view of Kumar R. et al., in US PG Pub. US20190042160 A1, published on February 7, 2019, (hereafter, Kumar).
Claim 2:
Regarding claim 2, Chen in view of Badaroglu, teaches the limitations of claim 1.
However, Chen in view of Badaroglu, did not teach “2. The system of claim 1, a system bus coupled to the analog computation-in-memory engine and the digital computation-in-memory engine.”
In an analogous system, Kumar teaches “2. The system of claim 1, a system bus coupled to the analog computation-in-memory engine and the digital computation-in- memory engine,”
See Kumar in [0167] describe " it will be understood that system 1200 can include one or more buses or bus systems between devices, such as a memory bus, a graphics bus, interface buses, or others. Buses or other signal lines can communicatively or electrically couple components together, or both communicatively and electrically couple the components. Buses can include physical communication lines, point-to-point connections, bridges, adapters, controllers, or other circuitry or a combination. Buses can include, for example, one or more of a system bus, a Peripheral Component Interconnect (PCI) bus, a HyperTransport or industry standard architecture (ISA) bus, a small computer system interface (SCSI) bus, a universal serial bus (USB), or an Institute of Electrical and Electronics Engineers (IEEE) standard 1394 bus." Here, Kumar mentions that the bus or bus system (i.e. system bus) connects between different components, including connecting an analog CIM with a digital CIM or other hardware components within the system.
Further, see Kumar in [0028], mention “CIM accelerators based on analog operations allow for lower cost computation and higher effective memory bandwidth from multibit data readout per column access.” Kumar also uses an analog CIM in the system as one embodiment.
Later, see Kumar in [0031] mention "Employing such a time to digital technique in a CIM circuit offers multiple advantages over more traditional voltage or current based CIM techniques." Kumar mentions using a digital CIM also being a part of the system.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Chen with the reference of Badaroglu, and incorporate with the teachings of Kumar by using the teachings of Chen and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Kumar’s teaching of a system bus coupled to the analog and digital CIM systems.
One of ordinary skill in the art would be motivated to do so because by integrating Kumar’s framework into the methods of Chen and Badaroglu, one with ordinary skill in the art would achieve the goal of providing “It will be understood that a differential bitline architecture can improve the ability of analog processor 630 to read or sense the bit value of the storage cells that make up array 610, given that the sensing can be performed as a comparison of the two lines that reduces the effects of noise.” (see Kumar in [0116]).
Claim 6:
Regarding claim 6, Chen in view of Badaroglu, teach the limitations of claim 1.
However, Chen in view of Badaroglu, did not teach “6. The system of claim 1, wherein the digital computation-in-memory engine comprises a plurality of static random access memory (SRAM) cells arranged into rows and columns.”
In an analogous art, Kumar teaches “6. The system of claim 1, wherein the digital computation-in-memory engine comprises a plurality of static random access memory (SRAM) cells arranged into rows and columns”
See Kumar in [0035] describe “Bitcell 122 is an example of a memory cell. The memory cell can be a bitcell in accordance with any of a variety of different technologies. The bitcells are at the intersection of a row with a column. In one example, bitcell 122 is a static random access memory (SRAM) cell.” Here, Kumar shows that the CIM system is made of SRAM cells, which are organized into rows and columns.
Further, see Kumar in [0031] "Employing such a time to digital technique in a CIM circuit offers multiple advantages over more traditional voltage or current based CIM techniques." Later, see Kumar in [0034] mention "FIG. 1A is a block diagram of an example of a compute-in memory system that performs computations with time-to-digital computation. System 100 represents an example of a compute-in memory (CIM) block or CIM circuitry." Kumar mentions the time to digital computation means this system can also include a digital CIM system and explicitly mentions using a digital CIM from [0031].
Later, see Kumar in [0113] describe " FIG. 6 is a block diagram of an example of a compute-in memory circuit that performs global charge sharing for multi-row access with a column major memory array and a differential bitline. System 600 is an example of a CIM array in accordance with an embodiment of system 500 of FIG. 5. System 600 illustrates elements of a memory array 610 with CIM circuitry, and it will be understood that the memory array includes more elements than what are shown. In one example, memory array 610 is an SRAM array.” Kumar shows that SRAM is part of the CIM system. See Kumar in paragraphs [0114, 0183] for details.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Chen with the reference of Badaroglu, and incorporate with the teachings of Kumar by using the teachings of Chen and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Kumar’s teaching of SRAM cells arranged into rows and columns.
One of ordinary skill in the art would be motivated to do so because by integrating Kumar’s framework into the methods of Chen and Badaroglu, one with ordinary skill in the art would achieve the goal of providing “It will be understood that a differential bitline architecture can improve the ability of analog processor 630 to read or sense the bit value of the storage cells that make up array 610, given that the sensing can be performed as a comparison of the two lines that reduces the effects of noise.” (see Kumar in [0116]).
Claim 7:
Regarding claim 7, Chen in view of Badaroglu, teaches the limitations of claim 1.
However, Chen in view of Badaroglu, did not teach “7. The system of claim 1, wherein the digital computation-in-memory engine comprises a plurality of CIM digital cells arranged into rows and columns.”
In an analogous art, Kumar teaches “7. The system of claim 1, wherein the digital computation-in-memory engine comprises a plurality of CIM digital cells arranged into rows and columns”
See Kumar in [0035] describe “Bitcell 122 is an example of a memory cell. The memory cell can be a bitcell in accordance with any of a variety of different technologies. The bitcells are at the intersection of a row with a column. In one example, bitcell 122 is a static random access memory (SRAM) cell.” Here, Kumar shows that the memory cells are organized by rows and columns. The examiner construes CIM digital cells to also include similar synonyms, such as bitcells or in memory computing unit or other similar elements.
See Kumar in [0081-0082] mention "In one example, the in-memory processor uses the digital output from the TDC cell directly for further computations such as multiplication, accumulation, or other operations, ..., the CIM circuit can perform direct multiplication through bit-serial operation for higher precision inputs. Bit-serial operation refers to accumulation and shift of the outputs for different wordlines. [0082] FIG. 3 is a block diagram of an example of a compute-in memory circuit with time-to-digital computation. System 300 provides an example of a CIM circuit in accordance with system 100 of FIG. 1. System 300 includes a memory array, which is not specifically identified, but includes rows and columns of storage cells or bitcells. In one example, the memory array is partitioned. System 300 represents four partitions, Partition[3:0], but it will be understood that more or fewer partitions can be used. Partitioning the memory array into multiple subarrays allows control over rows or wordlines by local row decoders, which can access multiple rows simultaneously per subarray." Here, Kumar explicitly describes that the memory array includes rows and columns of bitcells, which are the CIM digital cells mentioned from [0035].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Chen with the reference of Badaroglu, and incorporate with the teachings of Kumar by using the teachings of Chen and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Kumar’s teaching of CIM digital cells arranged into rows and columns.
One of ordinary skill in the art would be motivated to do so because by integrating Kumar’s framework into the methods of Chen and Badaroglu, one with ordinary skill in the art would achieve the goal of providing “It will be understood that a differential bitline architecture can improve the ability of analog processor 630 to read or sense the bit value of the storage cells that make up array 610, given that the sensing can be performed as a comparison of the two lines that reduces the effects of noise.” (see Kumar in [0116]).
Claims 3, 4, and 5 are rejected under 35 U.S.C. 103 over Chen in view of Badaroglu, further in view of Tran, H. et al. in US PG Pub. No. US20210019608-A1, published on January 21, 2021, (hereafter, Tran21).
Claim 3:
Regarding claim 3, Chen in view of Badaroglu, teach the limitations of claim 1.
However, Chen in view of Badaroglu, did not teach “3. The system of claim 1, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns.”
In an analogous art, Tran21 teaches “3. The system of claim 1, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns,”
See Tran21 in [0006] mention “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells,” and see Tran21 in [0002] describe “Testing circuitry and methods are disclosed for use with analog neural memory in deep learning artificial neural networks. The analog neural memory comprises one or more arrays of non-volatile flash memory cells.” Here, Tran21 describes an analog CIM that performs calculations on neural networks. Tran21 specifies that the memory cells do the work, meaning the system computes data right inside the memory device instead of moving it to a separate processor.
Further, see Tran21 in paragraph [0011] describe “One embodiment comprises a method of verifying values programmed into a plurality of non-volatile memory cells in an array of analog neural non-volatile memory cells, wherein the array is arranged in rows and columns, wherein each row is coupled to a word line and each column is coupled to a bit line, and wherein each word line is selectively coupled to a row decoder and each bit line is selectively coupled to a column decoder, the method comprising: asserting, by the row decoder, all word lines in the array; asserting, by the column decoder, a bit line in the array; sensing, by a sense amplifier, a current received from the bit line; and comparing the current to a reference current to determine if the non-volatile memory cells coupled to the bit line contain the desired values.” Here, Tran21 describes that the analog CIM has non-volatile memory cell arrays arranged in rows and columns (i.e. non-volatile memory cells arranged into rows and columns).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen and of Badaroglu, and incorporate with the teachings of Tran21 by using the teachings of Chen and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Tran21’s teaching of non-volatile memory cell arrays arranged in rows and columns.
One of ordinary skill in the art would be motivated to do so because by integrating Tran21’s framework into the methods of Chen and Badaroglu, one with ordinary skill in the art would achieve “by performing the multiplication and addition function, VMM array 33 negates the need for separate multiplication and addition logic circuits and is also power efficient due to its in-situ memory computation,” (see Tran21 in [0094]).
Claim 4:
Regarding claim 4, Chen in view of Badaroglu, further in view of Tran21, teaches the limitations of claim 3. Further, Tran21 teaches “4. The system of claim 3, wherein the non-volatile memory cells are stacked-gate flash memory cells.”
See Tran21 in [0082] describe “FIG. 7 depicts stacked gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is similar to memory cell 210 of FIG. 2, except that floating gate 20 extends over the entire channel region 18, and control gate terminal 22 (which here will be coupled to a word line) extends over floating gate 20, separated by an insulating layer (not shown). The erase, programming, and read operations operate in a similar manner to that described previously for memory cell 210.” Tran21 shows this system includes stacked gate flash memory cells.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen and of Badaroglu, and incorporate with the teachings of Tran21 by using the teachings of Chen and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Tran21’s teaching of stacked gate flash memory cells.
One of ordinary skill in the art would be motivated to do so because by integrating Tran21’s framework into the methods of Chen and Badaroglu, one with ordinary skill in the art would achieve “by performing the multiplication and addition function, VMM array 33 negates the need for separate multiplication and addition logic circuits and is also power efficient due to its in-situ memory computation,” (see Tran21 in [0094]).
Claim 5:
Regarding claim 5, Chen in view of Badaroglu, further in view of Tran21, teach the limitations of claim 3. Further, Tran21 teaches “5. The system of claim 3, wherein the non-volatile memory cells are split-gate flash memory cells”
See Tran21 in [0002] describe “Testing circuitry and methods are disclosed for use with analog neural memory in deep learning artificial neural networks. The analog neural memory comprises one or more arrays of non-volatile flash memory cells.” Tran21 shows using non-volatile memory cells as part of the system.
Further, see Tran21 in figures 2-6, also in [0025-0029] mention "[0025] FIG. 2 depicts a prior art split gate flash memory cell. [0026] FIG. 3 depicts another prior art split gate flash memory cell, [0027] FIG. 4 depicts another prior art split gate flash memory cell. [0028] FIG. 5 depicts another prior art split gate flash memory cell , [0029] FIG. 6 depicts another prior art split gate flash memory cell." From figures 2-6, Tran21 illustrates having split-gate flash memory cells as part of the non-volatile memory cells mentioned in [0002].
PNG
media_image3.png
490
617
media_image3.png
Greyscale
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen and of Badaroglu, and incorporate with the teachings of Tran21 by using the teachings of Chen and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Tran21’s teaching of split-gate flash memory cells.
One of ordinary skill in the art would be motivated to do so because by integrating Tran21’s framework into the methods of Chen and Badaroglu, one with ordinary skill in the art would achieve “By performing the multiplication and addition function, VMM array 33 negates the need for separate multiplication and addition logic circuits and is also power efficient due to its in-situ memory computation,” (see Tran21 in [0094]).
Claim 8 is rejected under 35 U.S.C. 103 over Chen in view of Badaroglu, further in view of Kumar, further in view of Hoang T. et al., in US PG. Pub. No. US20210397974-A1 “Multi-precision digital compute-in-memory deep neural network engine for flexible and energy efficient inferencing”, published on December 23, 2021, (hereafter, Hoang).
Claim 8:
Regarding claim 8, Chen in view of Badaroglu, further in view of Kumar, teaches the limitations of claim 7.
However, Chen in view of Badaroglu, further in view of Kumar, did not teach “8. The system of claim 7, wherein the plurality of CIM digital cells respectively comprise a multiply logic cell and a 2-bit adder logic cell.”
In an analogous art, Hoang teaches “8. The system of claim 7, wherein the plurality of CIM digital cells respectively comprise a multiply logic cell and a 2-bit adder logic cell”
See Hoang in [0028] describe “For the embodiments described below, the in-array multiplication is performed between multi-bit valued inputs, or activations, for a layer of the DNN and multi-bit valued weights of the layer. Each bit of a weight value is stored in a binary valued memory cell of the memory array and each bit of the input is applied as a binary input to a word line of the array for the multiplication of the input with the weight. To perform a multiply and accumulate operation, the results of the multiplications are accumulated by adders connected to sense amplifiers along the bit lines of the array. The adders can be configured to multiple levels of precision, so that the same structure can accommodate weights and activations of 8-bit, 4-bit, and 2-bit precision.” Here, Hoang describes using an adder cell that can calculate with various levels of precision, including precision of 2 - bit (i.e. 2-bit adder logic cell).
Further, see Hoang in [0078] mention “The following presents embodiments of a digital multi-precision Compute-in-Memory Deep Neural Network (CIM-DNN) engine for flexible and energy-efficient inferencing. The described memory array architectures support multi-precisions for both activations (inputs to a network layer) and weights…” Here, Hoang mentions information related to a digital CIM system.
Further, see Hoang in [0076] describe “A common technique for executing the matrix multiplications is by use of a multiplier-accumulator (MAC, or MAC unit). However, this has a number of issues. Referring back to FIG. 9B, the inference phase loads the neural network weights at step 922 before the matrix multiplications are performed by the propagation at step 923. However, as the amount of data involved can be extremely large, use of a multiplier-accumulator for inferencing has several issues related to the loading of weights. One of these is high energy dissipation due to having to use large MAC arrays with the required bit-width. Another is high energy dissipation due to the limited size of MAC arrays, resulting in high data movement between logic and memory and an energy dissipation that can be much higher than used in the logic computations themselves.” Here, Hoang describes a MAC unit is another related term for a multiply logic cell, since this reflects performing multiplication operations inside the memory array of the digital CIM.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen, Badaroglu, and Kumar, and incorporate with the teachings of Hoang by using the teachings of Chen, Badaroglu, and Kumar of using a computation-in-memory system to perform operations within a neural network layer, with Hoang’s teaching of using a CIM digital cells that include a multiply logic cell and a 2-bit adder logic cell.
One of ordinary skill in the art would be motivated to do so because by integrating Hoang’s framework into the methods of Chen, Badaroglu, and Kumar, one with ordinary skill in the art would achieve a method where “each PP can be computed in-memory by using a single binary value (or single level cell, SLC) NAND Flash memory cell or SCM memory cell without the need of reading out data, improving performance and energy efficiency of the inference engine,” (see Hoang in [0081]).
Claim 10 is rejected under 35 U.S.C. 103 over Chen in view of Badaroglu, further in view of Freye, F. et al., in “Memristive Devices for Time Domain Compute-in-Memory,” published on October 25, 2022, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9930136 , (hereafter, Freye).
Claim 10:
Regarding claim 10, Chen teaches “1. A system comprising: an analog computation-in-memory engine to perform operations in a first layer in a neural network;”
See Chen in page 153, from abstract “Compute-in-memory (CIM) is emerging as a new computing architecture to overcome the high energy consumption of edge-side AI and IoT devices. When performing high-precision neural network calculations, analog CIM and digital CIM have their own advantages and disadvantages. In this paper, we combine the advantages of high energy efficiency of analog CIM and high accuracy of digital CIM to propose a charge-digital hybrid CIM (CDH-CIM) macro.... Simulation shows that the macro achieves 6.98~11.0 TOPS/W at 0.8V and 71.92% inference accuracy when performing CIFAR-100 dataset.” Chen shows using a hybrid of using both analog and digital CIM systems into the same system for performing calculations of neural network models, includes both analog CIM and digital CIM systems.
Further, see Chen in page 157, section B. Network Mapping describe “Figure 7 shows the mapping relationship of the MAC operations in the convolutional layer implemented in the CDH-CIM macro. For CNN operations, the vast majority of the computation is in the MAC operations. The figure shows the computation of the convolution of a 56∗56∗64 feature map of with 64 3∗3∗64 filters. The convolution operation of each filter with the input feature map gets the outputs of one layer and the convolution operation of 64 filters gets the output of 64 layers. The size and depth of the feature map is kept constant by setting the number of fi[l]ters and using padding for the input feature map.” Here, Chen shows that each hardware CIM component of the CDH-CIM macro calculates for one layer of the neural network model.
Further, Chen teaches “and a digital computation-in-memory engine to perform operations in a second layer in the neural network,”
See Chen in page 153, from abstract “Compute-in-memory (CIM) is emerging as a new computing architecture to overcome the high energy consumption of edge-side AI and IoT devices. When performing high-precision neural network calculations, analog CIM and digital CIM have their own advantages and disadvantages. In this paper, we combine the advantages of high energy efficiency of analog CIM and high accuracy of digital CIM to propose a charge-digital hybrid CIM (CDH-CIM) macro.... Simulation shows that the macro achieves 6.98~11.0 TOPS/W at 0.8V and 71.92% inference accuracy when performing CIFAR-100 dataset.” Chen shows using a hybrid of using both analog and digital CIM systems into the same system for performing calculations of neural network models, includes both analog CIM and digital CIM systems.
However, Chen did not teach “and a digital computation-in-memory engine to perform operations in a second layer in the neural network” or “and a dynamic weight engine to perform operations in a third layer in the neural network; wherein the first layer, the second layer, and the third layer are different layers in the neural network.”
In an analogous art, Badaroglu teaches “and a digital computation-in-memory engine to perform operations in a second layer in the neural network” and “…engine to perform operations in a third layer in the neural network; wherein the first layer, the second layer, and the third layer are different layers in the neural network”
See Badaroglu in [0028] describe “Some aspects provide a hybrid neural network architecture using both compute-in-memory (CIM) and neural processing unit (NPU) processing elements (PEs), where the CIM PEs and the NPU PEs can share resources (e.g., memory), can concurrently operate, and can transfer data from one type of PE to another type of PE within the same neural network layer or in different neural network layers (e.g., adjacent layers). For example, the CIM and NPU PEs may be coupled to the same tightly coupled memory (TCM) bus for transferring weights, activation inputs, and/or outputs. A hybrid architecture as presented herein may offer the best (or at least better) energy consumption and speed trade-offs than conventional neural network architectures,” Here, Badaroglu also mentions for each CIM system, the performance operations work in a second layer or a third layer, each different from a first neural network layer within the same neural network model.
Further, see Badaroglu in [0059] describe "This processing may be repeated for each layer of the image data, and the outputs (e.g., output activations) may be stored in the output buffers and then used by the mobile device for an ML/AI task, such as facial recognition." In addition, Badaroglu mentions that the ‘processing may be repeated for each layer’ relates to performing calculations in subsequent layers like first, second, third or additional layers of a neural network model.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Chen and incorporate into the teachings of Badaroglu because both references teach using a type of computation-in-memory system to perform operations within a neural network layer.
One of ordinary skill in the art would be motivated to do so because “CIM units are generally better than NPUs in terms of energy efficiency, whereas NPUs are generally better than CIM units for depth-wise convolution. Due to bit-serial operation, digital compute-in-memory (DCIM) units may have lower TOPS, but comparable or better performance for a given area (e.g., in terms of TOPS/mm2) than NPUs,” (see Badaroglu in [0082]).
However, Chen in view of Badaroglu did not teach “a dynamic weight engine to perform operations in a third layer in the neural network;”
In an analogous system, Freye teaches “a dynamic weight engine to perform operations in a third layer in the neural network;”
See Freye on page 120 in Section II. Cascaded TDCIM Architecture describe “The operation of the typical convolution layer is shown in (2), where x is the input activation and w is the weight vector. f is the activation function and usually the binarize function for binary neural networks (BNNs) and the rectified linear unit (RELU) function for CNNs
Z=f(w⋅x).(2)
In TD computing, cascaded variable delay elements can implement an accumulation. Each element realizes the delay to encode one multiplication result, thus realizing the MAC function. Unlike in the traditional digital circuits, the convolution result is therefore presented as an accumulated delay.
For the TDCIM architecture, the weights of one kernel are stored in the memory of one computing chain. For a kernel size of N with M computing chains in parallel, the total area, Atot , is given by M⋅N⋅Acell+ATDC with the cell area, Acell , and the area for time-to-digital converter (TDC), ATDC . After computing a convolution, the activation signals are changed, whereas the weights can remain in memory. The weights are only updated after the complete output feature map is computed, thus reducing the data movement from the main memory. The input activation can be shared by all the computing chains and further reduce the data movement.” Here, Freye mentions a time domain compute in memory or TDCIM that relates to a third separate engine that helps store and update weights within a CNN or a type of neural network, which relates to being a dynamic weight engine. The examiner construes dynamic weight to mean any value or quantity that can be updated, see reference from specification in [0153] stating “dynamic weight engine 3603 is a device whose stored weights can be modified by changing a bias voltage or bias current without performing a separate erase or program operation. Since the weights can be modified without performing a separate erase or program operation, these weights are considered dynamic weights.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen and Badaroglu and incorporate with the teachings of Freye by using the teachings of Chen and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Freye’s teaching of using a dynamic weight engine.
One of ordinary skill in the art would be motivated to do so because by integrating Freye’s framework into the methods of Chen and Badaroglu, one with ordinary skill in the art would achieve “In this domain, analog computing is often considered to decrease power consumption further. ... A different compute scheme is time-domain CIM (TDCIM). In time-domain (TD) computing, the values are encoded as discrete arrival times of signal edges. While signaling is sample discrete, the arrival time is fundamentally continuous. Similar to charge and current, time is inherently additive, allowing for efficient accumulation operations,” (see Freye in page 119, second paragraph of Introduction).
Claims 11, 15, 16, and 18 are rejected under 35 U.S.C. 103 over Chen in view of Badaroglu, further in view of Freye, and further in view of Kumar.
Claim 11:
Regarding claim 11, Chen in view of Badaroglu, further in view of Freye, teaches the limitations of claim 10.
Further, Freye teaches “The system of claim 10, comprising: a system bus coupled to the analog computation-in-memory engine, the digital computation-in-memory engine, and the dynamic weight engine”
See Freye on page 120 in Section II. Cascaded TDCIM Architecture describe “the operation of the typical convolution layer is shown in (2), where x is the input activation and w is the weight vector. f is the activation function and usually the binarize function for binary neural networks (BNNs) … For the TDCIM architecture, the weights of one kernel are stored in the memory of one computing chain. For a kernel size of N with M computing chains in parallel, the total area, Atot , is given by M⋅N⋅Acell+ATDC with the cell area, Acell , and the area for time-to-digital converter (TDC), ATDC . After computing a convolution, the activation signals are changed, whereas the weights can remain in memory. The weights are only updated after the complete output feature map is computed, thus reducing the data movement from the main memory. The input activation can be shared by all the computing chains and further reduce the data movement.” Here, Freye mentions a time domain compute in memory or TDCIM that relates to a third separate engine that performs operations by helping store and update weights within a CNN neural network, which relates to being a dynamic weight engine. The examiner construes dynamic weight to mean any value or quantity that can be updated, see reference from specification in [0153] for details.
However, Chen in view of Badaroglu, further in view of Freye, did not teach “The system of claim 10, comprising: a system bus coupled to the analog computation-in-memory engine, the digital computation-in-memory engine …”
In an analogous system, Kumar teaches “The system of claim 10, comprising: a system bus coupled to the analog computation-in-memory engine, the digital computation-in-memory engine …”
See Kumar in [0167] describe " it will be understood that system 1200 can include one or more buses or bus systems between devices, such as a memory bus, a graphics bus, interface buses, or others. Buses or other signal lines can communicatively or electrically couple components together, or both communicatively and electrically couple the components. Buses can include physical communication lines, point-to-point connections, bridges, adapters, controllers, or other circuitry or a combination. Buses can include, for example, one or more of a system bus, a Peripheral Component Interconnect (PCI) bus, a HyperTransport or industry standard architecture (ISA) bus, a small computer system interface (SCSI) bus, a universal serial bus (USB), or an Institute of Electrical and Electronics Engineers (IEEE) standard 1394 bus." Here, Kumar mentions that the bus or bus system (i.e. system bus) connects between different components, including connecting an analog CIM with a digital CIM or other hardware components within the system.
Further, see Kumar in [0028], mention “CIM accelerators based on analog operations allow for lower cost computation and higher effective memory bandwidth from multibit data readout per column access.” Kumar also uses an analog CIM in the system as one embodiment.
Later, see Kumar in [0031] mention "Employing such a time to digital technique in a CIM circuit offers multiple advantages over more traditional voltage or current based CIM techniques." Kumar mentions using a digital CIM also being a part of the system.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Chen, Badaroglu, and Freye, and incorporate with the teachings of Kumar by using the teachings of Chen, Badaroglu, and Freye of using a computation-in-memory system to perform operations within a neural network layer, with Kumar’s teaching of a system bus coupled to the analog and digital CIM systems.
One of ordinary skill in the art would be motivated to do so because by integrating Kumar’s framework into the methods of Chen, Badaroglu, and Freye, one with ordinary skill in the art would achieve the goal of providing “It will be understood that a differential bitline architecture can improve the ability of analog processor 630 to read or sense the bit value of the storage cells that make up array 610, given that the sensing can be performed as a comparison of the two lines that reduces the effects of noise.” (see Kumar in [0116]).
Claim 15:
Regarding claim 15, Chen in view of Badaroglu, further in view of Freye, teaches the limitations of claim 10. However, Chen in view of Badaroglu, further in view of Freye, did not teach “The system of claim 10, wherein the digital computation-in-memory engine comprises a plurality of static random access memory (SRAM) cells arranged into rows and columns.”
In an analogous art, Kumar teaches “The system of claim 10, wherein the digital computation-in-memory engine comprises a plurality of static random access memory (SRAM) cells arranged into rows and columns”
See Kumar in [0035] describe “Bitcell 122 is an example of a memory cell. The memory cell can be a bitcell in accordance with any of a variety of different technologies. The bitcells are at the intersection of a row with a column. In one example, bitcell 122 is a static random access memory (SRAM) cell.” Here, Kumar shows that the CIM system is made of SRAM cells, which are organized into rows and columns.
Further, see Kumar in [0031] "Employing such a time to digital technique in a CIM circuit offers multiple advantages over more traditional voltage or current based CIM techniques." Later, see Kumar in [0034] mention "FIG. 1A is a block diagram of an example of a compute-in memory system that performs computations with time-to-digital computation. System 100 represents an example of a compute-in memory (CIM) block or CIM circuitry." Kumar mentions the time to digital computation means this system can also include a digital CIM system and explicitly mentions using a digital CIM from [0031].
Later, see Kumar in [0113] describe " FIG. 6 is a block diagram of an example of a compute-in memory circuit that performs global charge sharing for multi-row access with a column major memory array and a differential bitline. System 600 is an example of a CIM array in accordance with an embodiment of system 500 of FIG. 5. System 600 illustrates elements of a memory array 610 with CIM circuitry, and it will be understood that the memory array includes more elements than what are shown. In one example, memory array 610 is an SRAM array.” Kumar shows that SRAM is part of the CIM system. See Kumar in paragraphs [0114, 0183] for details.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen, Badaroglu, and Freye, and incorporate with the teachings of Kumar by using the teachings of Chen, Badaroglu, and Freye of using a computation-in-memory system to perform operations within a neural network layer, with Kumar’s teaching of SRAM cells arranged into rows and columns.
One of ordinary skill in the art would be motivated to do so because by integrating Kumar’s framework into the methods of Chen, Badaroglu, and Freye, one with ordinary skill in the art would achieve the goal of providing “a differential bitline architecture can improve the ability of analog processor 630 to read or sense the bit value of the storage cells that make up array 610, given that the sensing can be performed as a comparison of the two lines that reduces the effects of noise.” (see Kumar in [0116]).
Claim 16:
Regarding claim 16, Chen in view of Badaroglu, further in view of Freye, teaches the limitations of claim 10. However, Chen in view of Badaroglu, further in view of Freye, did not teach “The system of claim 10, wherein the digital computation-in-memory engine comprises a plurality of CIM digital cells.”
However, in an analogous art, Kumar teaches “The system of claim 10, wherein the digital computation-in-memory engine comprises a plurality of CIM digital cells”
See Kumar in [0035] describe “Bitcell 122 is an example of a memory cell. The memory cell can be a bitcell in accordance with any of a variety of different technologies. The bitcells are at the intersection of a row with a column. In one example, bitcell 122 is a static random access memory (SRAM) cell.” Here, Kumar shows that the memory cells are organized by rows and columns. The examiner construes CIM digital cells to also include similar synonyms, such as bitcells or in memory computing unit or other similar elements.
See Kumar in [0081-0082] mention "In one example, the in-memory processor uses the digital output from the TDC cell directly for further computations such as multiplication, accumulation, or other operations, ..., the CIM circuit can perform direct multiplication through bit-serial operation for higher precision inputs. Bit-serial operation refers to accumulation and shift of the outputs for different wordlines. [0082] FIG. 3 is a block diagram of an example of a compute-in memory circuit with time-to-digital computation. System 300 provides an example of a CIM circuit in accordance with system 100 of FIG. 1. System 300 includes a memory array, which is not specifically identified, but includes rows and columns of storage cells or bitcells. In one example, the memory array is partitioned. System 300 represents four partitions, Partition[3:0], but it will be understood that more or fewer partitions can be used. Partitioning the memory array into multiple subarrays allows control over rows or wordlines by local row decoders, which can access multiple rows simultaneously per subarray." Here, Kumar explicitly describes that the memory array includes rows and columns of bitcells, which are the CIM digital cells mentioned from [0035].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen, Badaroglu, and Freye, and incorporate with the teachings of Kumar by using the teachings of Chen, Badaroglu, and Freye of using a computation-in-memory system to perform operations within a neural network layer, with Kumar’s teaching of CIM digital cells arranged into rows and columns.
One of ordinary skill in the art would be motivated to do so because by integrating Kumar’s framework into the methods of Chen, Badaroglu, and Freye, one with ordinary skill in the art would achieve the goal of providing “a differential bitline architecture can improve the ability of analog processor 630 to read or sense the bit value of the storage cells that make up array 610, given that the sensing can be performed as a comparison of the two lines that reduces the effects of noise.” (see Kumar in [0116]).
Claim 18:
Regarding claim 18, Chen in view of Badaroglu, further in view of Freye, and further in view of Kumar, teach the limitations of claim 16.
Further, Badaroglu teaches “The system of claim 16, wherein the digital computation-in-memory engine comprises a shift and adder tree coupled to the plurality of CIM digital cells”
See Badaroglu in [0065 - 0066] describe “As shown, the DCIM circuit 400 may include a bit-column adder tree 409, which may include eight adder trees 4100 to 4107 (collectively referred to as “adder trees 410”), each adder tree being implemented for a respective one of the columns 406. Each of the adder trees 410 adds the output signals from the CIM cells 402 on the respective one of the columns 406, and the adder trees 410 may operate in parallel (e.g., concurrently). The outputs of the adder trees 410 may be coupled to a weight-shift adder tree circuit 412, as shown. The weight-shift adder tree circuit 412 includes multiple weight-shift adders 414, each including a bit-shift-and-add circuit to facilitate the performance of a bit-shifting-and-addition operation. In other words, the CIM cells on column 4060 may store the most-significant bits (MSBs) for respective weights on each word-line 404, and the CIM cells on column 4067 may store the least-significant bits (LSBs) for respective weights on each word-line. Therefore, when performing the addition across the columns 406, a bit-shift operation is performed to shift the bits to account for the significance of the bits on the associated column. [0066] The output of the weight-shift adder tree circuit 412 is provided to an activation-shift accumulator circuit 416. The activation-shift accumulator circuit 416 includes a bit-shift circuit 418, a serial accumulator 420, and a flip-flop (FF) array 422. For example, the FF array 422 may be used to implement a register.” Here, Badaroglu describes using a weight shift adder tree circuit 412 that is part of the digital CIM or DCIM circuit unit 400.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Chen and incorporate into the teachings of Badaroglu because both references teach using a type of computation-in-memory system to perform operations within a neural network layer using shift and adder trees.
One of ordinary skill in the art would be motivated to do so because “CIM units are generally better than NPUs in terms of energy efficiency, whereas NPUs are generally better than CIM units for depth-wise convolution. Due to bit-serial operation, digital compute-in-memory (DCIM) units may have lower TOPS, but comparable or better performance for a given area (e.g., in terms of TOPS/mm2) than NPUs,” (see Badaroglu in [0082]).
Claims 12, 13, and 14 are rejected under 35 U.S.C. 103 over Chen in view of Badaroglu, further in view of Freye, and further in view of Tran21.
Claim 12:
Regarding claim 12, Chen in view of Badaroglu, further in view of Freye, teaches the limitations of claim 10. However, Chen in view of Badaroglu, further in view of Freye, did not teach “The system of claim 10, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns.”
In an analogous art, Tran21 teaches “The system of claim 10, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns”
See Tran21 in [0006] mention “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells,” and see Tran21 in [0002] describe “Testing circuitry and methods are disclosed for use with analog neural memory in deep learning artificial neural networks. The analog neural memory comprises one or more arrays of non-volatile flash memory cells.” Here, Tran21 describes an analog CIM that performs calculations on neural networks. Tran21 specifies that the memory cells do the work, meaning the system computes data right inside the memory device instead of moving it to a separate processor.
Further, see Tran21 in paragraph [0011] describe “One embodiment comprises a method of verifying values programmed into a plurality of non-volatile memory cells in an array of analog neural non-volatile memory cells, wherein the array is arranged in rows and columns, wherein each row is coupled to a word line and each column is coupled to a bit line, and wherein each word line is selectively coupled to a row decoder and each bit line is selectively coupled to a column decoder, the method comprising: asserting, by the row decoder, all word lines in the array; asserting, by the column decoder, a bit line in the array; sensing, by a sense amplifier, a current received from the bit line; and comparing the current to a reference current to determine if the non-volatile memory cells coupled to the bit line contain the desired values.” Here, Tran21 describes that the analog CIM has non-volatile memory cell arrays arranged in rows and columns (i.e. non-volatile memory cells arranged into rows and columns).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen, Badaroglu, and Freye, and incorporate with the teachings of Tran21 by using the teachings of Chen, Badaroglu, and Freye of using a computation-in-memory system to perform operations within a neural network layer, with Tran21’s teaching of non-volatile memory cell arrays arranged in rows and columns.
One of ordinary skill in the art would be motivated to do so because by integrating Tran21’s framework into the methods of Chen, Badaroglu, and Freye, one with ordinary skill in the art would achieve “by performing the multiplication and addition function, VMM array 33 negates the need for separate multiplication and addition logic circuits and is also power efficient due to its in-situ memory computation,” (see Tran21 in [0094]).
Claim 13:
Regarding claim 13, Chen in view of Badaroglu, further in view of Freye, and further in view of Tran21, teach the limitations of claim 12.
Further, Tran21 teaches “The system of claim 12, wherein the non-volatile memory cells are stacked-gate flash memory cells”
See Tran21 in [0082] describe “FIG. 7 depicts stacked gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is similar to memory cell 210 of FIG. 2, except that floating gate 20 extends over the entire channel region 18, and control gate terminal 22 (which here will be coupled to a word line) extends over floating gate 20, separated by an insulating layer (not shown). The erase, programming, and read operations operate in a similar manner to that described previously for memory cell 210.” Tran21 shows this system includes stacked gate flash memory cells.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen, Badaroglu, and Freye, and incorporate with the teachings of Tran21 by using the teachings of Chen, Badaroglu, and Freye of using a computation-in-memory system to perform operations within a neural network layer, with Tran21’s teaching of stacked gate flash memory cells.
One of ordinary skill in the art would be motivated to do so because by integrating Tran21’s framework into the methods of Chen, Badaroglu, and Freye, one with ordinary skill in the art would achieve “by performing the multiplication and addition function, VMM array 33 negates the need for separate multiplication and addition logic circuits and is also power efficient due to its in-situ memory computation,” (see Tran21 in [0094]).
Claim 14:
Regarding claim 14, Chen in view of Badaroglu, further in view of Freye, and further in view of Tran21, teach the limitations of claim 12.
Further, Tran21 teaches “The system of claim 12, wherein the non-volatile memory cells are split-gate flash memory cells”
See Tran21 in [0002] describe “Testing circuitry and methods are disclosed for use with analog neural memory in deep learning artificial neural networks. The analog neural memory comprises one or more arrays of non-volatile flash memory cells.” Tran21 shows using non-volatile memory cells as part of the system.
Further, see Tran21 in figures 2-6, also in [0025-0029] mention "[0025] FIG. 2 depicts a prior art split gate flash memory cell. [0026] FIG. 3 depicts another prior art split gate flash memory cell, [0027] FIG. 4 depicts another prior art split gate flash memory cell. [0028] FIG. 5 depicts another prior art split gate flash memory cell , [0029] FIG. 6 depicts another prior art split gate flash memory cell." From figures 2-6, Tran21 illustrates having split-gate flash memory cells as part of the non-volatile memory cells mentioned in [0002].
PNG
media_image3.png
490
617
media_image3.png
Greyscale
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen, Badaroglu, and Freye, and incorporate with the teachings of Tran21 by using the teachings of Chen, Badaroglu, and Freye of using a computation-in-memory system to perform operations within a neural network layer, with Tran21’s teaching of split-gate flash memory cells.
One of ordinary skill in the art would be motivated to do so because by integrating Tran21’s framework into the methods of Chen, Badaroglu, and Freye, one with ordinary skill in the art would achieve “By performing the multiplication and addition function, VMM array 33 negates the need for separate multiplication and addition logic circuits and is also power efficient due to its in-situ memory computation,” (see Tran21 in [0094]).
Claim 17 is rejected under 35 U.S.C. 103 over Chen in view of Badaroglu, further in view of Freye, further in view of Kumar, and further in view of Hoang.
Claim 17:
Regarding claim 17, Chen in view of Badaroglu, further in view of Freye, and further in view of Kumar, teach the limitations of claim 16. However, Chen in view of Badaroglu, further in view of Freye, and further in view of Kumar, did not teach “The system of claim 16, wherein the plurality of CIM digital cells respectively comprise a multiply logic cell and a 2-bit adder logic cell.”
In an analogous field, Hoang teaches “The system of claim 16, wherein the plurality of CIM digital cells respectively comprise a multiply logic cell and a 2-bit adder logic cell”
See Hoang in [0028] describe “For the embodiments described below, the in-array multiplication is performed between multi-bit valued inputs, or activations, for a layer of the DNN and multi-bit valued weights of the layer. Each bit of a weight value is stored in a binary valued memory cell of the memory array and each bit of the input is applied as a binary input to a word line of the array for the multiplication of the input with the weight. To perform a multiply and accumulate operation, the results of the multiplications are accumulated by adders connected to sense amplifiers along the bit lines of the array. The adders can be configured to multiple levels of precision, so that the same structure can accommodate weights and activations of 8-bit, 4-bit, and 2-bit precision.” Here, Hoang describes using an adder cell that can calculate with various levels of precision, including precision of 2 - bit (i.e. 2-bit adder logic cell).
Further, see Hoang in [0078] mention “The following presents embodiments of a digital multi-precision Compute-in-Memory Deep Neural Network (CIM-DNN) engine for flexible and energy-efficient inferencing. The described memory array architectures support multi-precisions for both activations (inputs to a network layer) and weights…” Here, Hoang mentions information related to a digital CIM system.
Further, see Hoang in [0076] describe “A common technique for executing the matrix multiplications is by use of a multiplier-accumulator (MAC, or MAC unit). However, this has a number of issues. Referring back to FIG. 9B, the inference phase loads the neural network weights at step 922 before the matrix multiplications are performed by the propagation at step 923. However, as the amount of data involved can be extremely large, use of a multiplier-accumulator for inferencing has several issues related to the loading of weights. One of these is high energy dissipation due to having to use large MAC arrays with the required bit-width. Another is high energy dissipation due to the limited size of MAC arrays, resulting in high data movement between logic and memory and an energy dissipation that can be much higher than used in the logic computations themselves.” Here, Hoang describes a MAC unit is another related term for a multiply logic cell, since this reflects performing multiplication operations inside the memory array of the digital CIM.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Chen, Badaroglu, Freye, and Kumar, and incorporate with the teachings of Hoang by using the teachings of Chen, Badaroglu, Freye, and Kumar of using a computation-in-memory system to perform operations within a neural network layer, with Hoang’s teaching of using a CIM digital cells that include a multiply logic cell and a 2-bit adder logic cell.
One of ordinary skill in the art would be motivated to do so because by integrating Hoang’s framework into the methods of Chen, Badaroglu, Freye, and Kumar, one with ordinary skill in the art would achieve a method where “each PP can be computed in-memory by using a single binary value (or single level cell, SLC) NAND Flash memory cell or SCM memory cell without the need of reading out data, improving performance and energy efficiency of the inference engine,” (see Hoang in [0081]).
Claims 19 and 20 are rejected under 35 U.S.C. 103 over Tran21 in view of Badaroglu.
Claim 19:
Regarding claim 19, Tran21 teaches “A method comprising: performing vector-by-matrix multiplication operations in a first layer of a neural network using a digital computation-in-memory engine;”
See Tran 21 in [0071] describe “...Digital non-volatile memories are well known. ...discloses an array of split gate non-volatile memory cells, which are a type of flash memory cells. Such a memory cell 210 is shown in FIG. 2.,” Also, see Tran21 in [0098] also mention " FIG. 10 is a block diagram depicting the usage of numerous layers of VMM systems 32, here labeled as VMM systems 32 a, 32 b, 32 c, 32 d, and 32 e. As shown in FIG. 10, the input, denoted Inputx, is converted from digital to analog by a digital-to-analog converter 31, and provided to input VMM system 32 a. The converted analog inputs could be voltage or current. The input D/A conversion for the first layer could be done by using a function or a LUT (look up table) that maps the inputs Inputx to appropriate analog levels for the matrix multiplier of input VMM system 32 a. The input conversion could also be done by an analog to analog (A/A) converter to convert an external analog input to a mapped analog input to the input VMM system 32 a. The input conversion could also be done by a digital-to-digital pul[s]e (D/P) converter to convert an external digital input to a mapped digital pulse or pulses to the input VMM system 32." Here, Tran21 mentions that vector by matrix multiplication or VMM can be performed in digital CIM systems for digital inputs.
Further, see Tran21 in [0093] describe "FIG. 9 is a block diagram of a system that can be used for that purpose. Vector-by-matrix multiplication (VMM) system 32 includes non-volatile memory cells and is utilized as the synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer." Here, Tran21 shows that VMM or vector-by- matrix multiplication calculations are run among layers of a neural network.
Further, see Tran21 in paragraph [0008] mention “Precision and accuracy are extremely important in operations involving VMM arrays, as each individual memory cell can store one of N different levels, where N can be greater than 2, as opposed to a traditional memory cell where N is always 2. This makes testing an extremely important operation. For example, verification of a programming operation is required to ensure that each individual cell or a column of cells is accurately programmed to the desired value. As another example, it is critical to identify bad cells or groups of cells so that they can be removed from the set of cells used to store data during operation of the VMM array.” Here, Tran21 describes VMM operations on memory cells, which are part of memory systems.
Further, Tran21 teaches “performing vector-by-matrix multiplication operations in a second layer different than the first layer of the neural network using an analog computation-in-memory engine”
See Tran21 in paragraphs [0006-0007] mentions " The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs…Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells used in this manner can be referred to as a vector by matrix multiplication (VMM) array... Each non-volatile memory cells used in the analog neural memory system must be erased and programmed to hold a very specific and precise amount of charge, i.e., the number of electrons, in the floating gate. " Here, Tran21 describes performing a vector by matrix multiplication operation in analog memory cells.
However, Tran21 did not teach “performing vector-by-matrix multiplication operations in a second layer different than the first layer of the neural network using an analog computation-in-memory engine.”
In an analogous art, Badaroglu teaches “performing vector-by-matrix multiplication operations in a second layer different than the first layer of the neural network using an analog computation-in-memory engine”
See Badaroglu in [0028] describe “Some aspects provide a hybrid neural network architecture using both compute-in-memory (CIM) and neural processing unit (NPU) processing elements (PEs), where the CIM PEs and the NPU PEs can share resources (e.g., memory), can concurrently operate, and can transfer data from one type of PE to another type of PE within the same neural network layer or in different neural network layers (e.g., adjacent layers). For example, the CIM and NPU PEs may be coupled to the same tightly coupled memory (TCM) bus for transferring weights, activation inputs, and/or outputs. A hybrid architecture as presented herein may offer the best (or at least better) energy consumption and speed trade-offs than conventional neural network architectures,” Here, Badaroglu also mentions for each CIM system, the performance operations work in a second layer or a third layer, each different from a first neural network layer within the same neural network model.
Further, see Badaroglu in [0059] describe "This processing may be repeated for each layer of the image data, and the outputs (e.g., output activations) may be stored in the output buffers and then used by the mobile device for an ML/AI task, such as facial recognition." In addition, Badaroglu mentions that the ‘processing may be repeated for each layer’ relates to performing calculations in subsequent layers like first, second, third or additional layers of a neural network model.
Also, see Badaroglu in [0060] describe “As used herein, the term “CIM” may refer to either or both analog CIM and digital CIM, unless it is clear from context that only analog CIM or only digital CIM is meant.” Badaroglu mentions using either an analog CIM or a digital CIM system.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Tran21 and incorporate into the teachings of Badaroglu because both references teach using a type of computation-in-memory system to perform operations within a neural network layer.
One of ordinary skill in the art would be motivated to do so because “CIM units are generally better than NPUs in terms of energy efficiency, whereas NPUs are generally better than CIM units for depth-wise convolution. Due to bit-serial operation, digital compute-in-memory (DCIM) units may have lower TOPS, but comparable or better performance for a given area (e.g., in terms of TOPS/mm2) than NPUs,” (see Badaroglu in [0082]).
Claim 20:
Regarding claim 20, Tran21 in view of Badaroglu, teaches the limitations of claim 19.
Further, Tran21 teaches “20. The method of claim 19, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns,”
See Tran21 in [0006] mention “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells,” and see Tran21 in [0002] describe “Testing circuitry and methods are disclosed for use with analog neural memory in deep learning artificial neural networks. The analog neural memory comprises one or more arrays of non-volatile flash memory cells.” Here, Tran21 describes an analog CIM that performs calculations on neural networks. Tran21 specifies that the memory cells do the work, meaning the system computes data right inside the memory device instead of moving it to a separate processor.
Further, see Tran21 in paragraph [0011] describe “One embodiment comprises a method of verifying values programmed into a plurality of non-volatile memory cells in an array of analog neural non-volatile memory cells, wherein the array is arranged in rows and columns, wherein each row is coupled to a word line and each column is coupled to a bit line, and wherein each word line is selectively coupled to a row decoder and each bit line is selectively coupled to a column decoder, the method comprising: asserting, by the row decoder, all word lines in the array; asserting, by the column decoder, a bit line in the array; sensing, by a sense amplifier, a current received from the bit line; and comparing the current to a reference current to determine if the non-volatile memory cells coupled to the bit line contain the desired values.” Here, Tran21 describes that the analog CIM has non-volatile memory cell arrays arranged in rows and columns (i.e. non-volatile memory cells arranged into rows and columns).
Claims 21, 22 and 23 are rejected under 35 U.S.C. 103 over Tran21 in view of Badaroglu, further in view of Freye.
Claim 21:
Regarding claim 21, Tran21 teaches “A method comprising: performing vector-by-matrix multiplication operations in a first layer of a neural network using a digital computation-in-memory engine;”
See Tran 21 in [0071] describe “...Digital non-volatile memories are well known. ...discloses an array of split gate non-volatile memory cells, which are a type of flash memory cells. Such a memory cell 210 is shown in FIG. 2.,” Also, see Tran21 in [0098] also mention " FIG. 10 is a block diagram depicting the usage of numerous layers of VMM systems 32, here labeled as VMM systems 32 a, 32 b, 32 c, 32 d, and 32 e. As shown in FIG. 10, the input, denoted Inputx, is converted from digital to analog by a digital-to-analog converter 31, and provided to input VMM system 32 a. The converted analog inputs could be voltage or current. The input D/A conversion for the first layer could be done by using a function or a LUT (look up table) that maps the inputs Inputx to appropriate analog levels for the matrix multiplier of input VMM system 32 a. The input conversion could also be done by an analog to analog (A/A) converter to convert an external analog input to a mapped analog input to the input VMM system 32 a. The input conversion could also be done by a digital-to-digital pul[s]e (D/P) converter to convert an external digital input to a mapped digital pulse or pulses to the input VMM system 32." Here, Tran21 mentions that vector by matrix multiplication or VMM can be performed in digital CIM systems for digital inputs.
Further, see Tran21 in [0093] describe "FIG. 9 is a block diagram of a system that can be used for that purpose. Vector-by-matrix multiplication (VMM) system 32 includes non-volatile memory cells and is utilized as the synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer." Here, Tran21 shows that VMM or vector-by-matrix multiplication calculations are run among layers of a neural network.
Further, see Tran21 in paragraph [0008] mention “Precision and accuracy are extremely important in operations involving VMM arrays, as each individual memory cell can store one of N different levels, where N can be greater than 2, as opposed to a traditional memory cell where N is always 2. This makes testing an extremely important operation. For example, verification of a programming operation is required to ensure that each individual cell or a column of cells is accurately programmed to the desired value. As another example, it is critical to identify bad cells or groups of cells so that they can be removed from the set of cells used to store data during operation of the VMM array.” Here, Tran21 describes VMM operations on memory cells, which are part of memory systems.
Further, Tran21 teaches “performing vector-by-matrix multiplication operations in a second layer of the neural network using an analog computation-in-memory engine”
See Tran21 in paragraphs [0006-0007] mentions “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs… Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells used in this manner can be referred to as a vector by matrix multiplication (VMM) array... Each non-volatile memory cells used in the analog neural memory system must be erased and programmed to hold a very specific and precise amount of charge, i.e., the number of electrons, in the floating gate.” Here, Tran21 describes performing a vector by matrix multiplication operation in analog memory cells.
Later, see Tran21 in [0099] describe “The output generated by input VMM system 32 a is provided as an input to the next VMM system (hidden level 1) 32 b, which in turn generates an output that is provided as an input to the next VMM system (hidden level 2) 32 c, and so on. The various layers of VMM system 32 function as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM system 32 a, 32 b, 32 c, 32 d, and 32 e can be a stand-alone, physical system comprising a respective non-volatile memory array, or multiple VMM systems could utilize different portions of the same physical non-volatile memory array, or multiple VMM systems could utilize overlapping portions of the same physical non-volatile memory array.” Here, Tran21 mentions that the vector by matrix multiplication calculations can be performed layer by layer within a neural network, and accounts for an analog CIM system perform calculations on a second layer of the neural network.
Further, Tran21 teaches “performing vector-by-matrix multiplication operations in a third layer of the neural network using a … engine”
See Tran21 in [0099] describe “The output generated by input VMM system 32 a is provided as an input to the next VMM system (hidden level 1) 32 b, which in turn generates an output that is provided as an input to the next VMM system (hidden level 2) 32 c, and so on. The various layers of VMM system 32 function as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM system 32 a, 32 b, 32 c, 32 d, and 32 e can be a stand-alone, physical system comprising a respective non-volatile memory array, or multiple VMM systems could utilize different portions of the same physical non-volatile memory array, or multiple VMM systems could utilize overlapping portions of the same physical non-volatile memory array.” Here, Tran21 mentions that the vector by matrix multiplication calculations can be performed layer by layer within a neural network, and accounts for an analog CIM system perform calculations on a subsequent layer, such as a third layer of the neural network.
Further, see Tran21 in [0098] describe “FIG. 10 is a block diagram depicting the usage of numerous layers of VMM systems 32, here labeled as VMM systems 32 a, 32 b, 32 c, 32 d, and 32 e. As shown in FIG. 10, the input, denoted Inputx, is converted from digital to analog by a digital-to-analog converter 31, and provided to input VMM system 32 a. The converted analog inputs could be voltage or current. The input D/A conversion for the first layer could be done by using a function or a LUT (look up table) that maps the inputs Inputx to appropriate analog levels for the matrix multiplier of input VMM system 32 a.” Here, Tran21 describes various VMM systems, which relate to an engine that performs vector by matrix multiplication calculations. See Tran21 in figure 10 for details.
PNG
media_image4.png
956
777
media_image4.png
Greyscale
However, Tran21 did not teach “performing vector-by-matrix multiplication operations in a third layer of the neural network using a dynamic weight engine; wherein the first layer, the second layer, and the third layer are different layers in the neural network.”
In an analogous art, Badaroglu teaches “… wherein the first layer, the second layer, and the third layer are different layers in the neural network”
See Badaroglu in [0028] describe “Some aspects provide a hybrid neural network architecture using both compute-in-memory (CIM) and neural processing unit (NPU) processing elements (PEs), where the CIM PEs and the NPU PEs can share resources (e.g., memory), can concurrently operate, and can transfer data from one type of PE to another type of PE within the same neural network layer or in different neural network layers (e.g., adjacent layers). For example, the CIM and NPU PEs may be coupled to the same tightly coupled memory (TCM) bus for transferring weights, activation inputs, and/or outputs. A hybrid architecture as presented herein may offer the best (or at least better) energy consumption and speed trade-offs than conventional neural network architectures,” Here, Badaroglu also mentions for each CIM system, the performance operations work in a second layer or a third layer, each different from a first neural network layer within the same neural network model.
Further, see Badaroglu in [0059] describe "This processing may be repeated for each layer of the image data, and the outputs (e.g., output activations) may be stored in the output buffers and then used by the mobile device for an ML/AI task, such as facial recognition." In addition, Badaroglu mentions that the ‘processing may be repeated for each layer’ relates to performing calculations in subsequent layers like first, second, third or additional layers of a neural network model.
Also, see Badaroglu in [0060] describe “As used herein, the term “CIM” may refer to either or both analog CIM and digital CIM, unless it is clear from context that only analog CIM or only digital CIM is meant.” Badaroglu mentions using either an analog CIM or a digital CIM system.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Tran21 and incorporate into the teachings of Badaroglu because both references teach using a type of computation-in-memory system to perform operations within various layers in a neural network.
One of ordinary skill in the art would be motivated to do so because “CIM units are generally better than NPUs in terms of energy efficiency, whereas NPUs are generally better than CIM units for depth-wise convolution. Due to bit-serial operation, digital compute-in-memory (DCIM) units may have lower TOPS, but comparable or better performance for a given area (e.g., in terms of TOPS/mm2) than NPUs,” (see Badaroglu in [0082]).
However, Tran21 in view of Badaroglu did not teach “performing vector-by-matrix multiplication operations in a third layer of the neural network using a dynamic weight engine,”
In an analogous art, Freye teaches “performing … operations in … the neural network using a dynamic weight engine”
See Freye on page 120 in Section II. Cascaded TDCIM Architecture describe “The operation of the typical convolution layer is shown in (2), where x is the input activation and w is the weight vector. f is the activation function and usually the binarize function for binary neural networks (BNNs) and the rectified linear unit (RELU) function for CNNs
Z=f(w⋅x).(2)
In TD computing, cascaded variable delay elements can implement an accumulation. Each element realizes the delay to encode one multiplication result, thus realizing the MAC function. Unlike in the traditional digital circuits, the convolution result is therefore presented as an accumulated delay.
For the TDCIM architecture, the weights of one kernel are stored in the memory of one computing chain. For a kernel size of N with M computing chains in parallel, the total area, Atot , is given by M⋅N⋅Acell+ATDC with the cell area, Acell , and the area for time-to-digital converter (TDC), ATDC . After computing a convolution, the activation signals are changed, whereas the weights can remain in memory. The weights are only updated after the complete output feature map is computed, thus reducing the data movement from the main memory. The input activation can be shared by all the computing chains and further reduce the data movement.” Here, Freye mentions a time domain compute-in memory or TDCIM that relates to a third separate engine that perform operations by helping store and update weights within a CNN neural network, which relates to being a dynamic weight engine. The examiner construes dynamic weight to mean any value or quantity that can be updated, see reference from specification in [0153] stating “dynamic weight engine 3603 is a device whose stored weights can be modified by changing a bias voltage or bias current without performing a separate erase or program operation. Since the weights can be modified without performing a separate erase or program operation, these weights are considered dynamic weights.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Tran21 and Badaroglu, and incorporate with the teachings of Freye by using the teachings of Tran21 and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Freye’s teaching of using a dynamic weight engine.
One of ordinary skill in the art would be motivated to do so because by integrating Freye’s framework into the methods of Tran21 and Badaroglu, one with ordinary skill in the art would achieve “In this domain, analog computing is often considered to decrease power consumption further. ... A different compute scheme is time-domain CIM (TDCIM). In time-domain (TD) computing, the values are encoded as discrete arrival times of signal edges. While signaling is sample discrete, the arrival time is fundamentally continuous. Similar to charge and current, time is inherently additive, allowing for efficient accumulation operations,” (see Freye in page 119, second paragraph of Introduction).
Claim 22:
Regarding claim 22, Tran21 in view of Badaroglu, further in view of Freye, teaches the limitations of claim 21.
Further, Tran21 teaches “The method of claim 21, wherein the analog computation-in-memory engine comprises a plurality of non-volatile memory cells arranged into rows and columns.”
See Tran21 in [0006] mention “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells,” and see Tran21 in [0002] describe “Testing circuitry and methods are disclosed for use with analog neural memory in deep learning artificial neural networks. The analog neural memory comprises one or more arrays of non-volatile flash memory cells.” Here, Tran21 describes an analog CIM that performs calculations on neural networks. Tran21 specifies that the memory cells do the work, meaning the system computes data right inside the memory device instead of moving it to a separate processor.
Further, see Tran21 in paragraph [0011] describe “One embodiment comprises a method of verifying values programmed into a plurality of non-volatile memory cells in an array of analog neural non-volatile memory cells, wherein the array is arranged in rows and columns, wherein each row is coupled to a word line and each column is coupled to a bit line, and wherein each word line is selectively coupled to a row decoder and each bit line is selectively coupled to a column decoder, the method comprising: asserting, by the row decoder, all word lines in the array; asserting, by the column decoder, a bit line in the array; sensing, by a sense amplifier, a current received from the bit line; and comparing the current to a reference current to determine if the non-volatile memory cells coupled to the bit line contain the desired values.” Here, Tran21 describes that the analog CIM has non-volatile memory cell arrays arranged in rows and columns (i.e. non-volatile memory cells arranged into rows and columns).
Claim 23:
Regarding claim 23, Tran21 teaches “performing vector-by-matrix multiplication operations in a first layer of a neural network using one of an analog computation-in-memory engine, a digital computation-in-memory engine, …;”
See Tran21 in [0071] describe "...Digital non-volatile memories are well known. ...discloses an array of split gate non-volatile memory cells, which are a type of flash memory cells. Such a memory cell 210 is shown in FIG. 2.,” Also, see Tran21 in [0098] also mention " FIG. 10 is a block diagram depicting the usage of numerous layers of VMM systems 32, here labeled as VMM systems 32 a, 32 b, 32 c, 32 d, and 32 e. As shown in FIG. 10, the input, denoted Inputx, is converted from digital to analog by a digital-to-analog converter 31, and provided to input VMM system 32 a. The converted analog inputs could be voltage or current. The input D/A conversion for the first layer could be done by using a function or a LUT (look up table) that maps the inputs Inputx to appropriate analog levels for the matrix multiplier of input VMM system 32 a. The input conversion could also be done by an analog to analog (A/A) converter to convert an external analog input to a mapped analog input to the input VMM system 32 a. The input conversion could also be done by a digital-to-digital pul[s]e (D/P) converter to convert an external digital input to a mapped digital pulse or pulses to the input VMM system 32 ." Here, Tran21 mentions that vector by matrix multiplication or VMM can be performed in digital CIM systems for digital inputs.
Also, see Tran21 in paragraphs [0006-0007] mentions “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs … Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells used in this manner can be referred to as a vector by matrix multiplication (VMM) array... Each non-volatile memory cells used in the analog neural memory system must be erased and programmed to hold a very specific and precise amount of charge, i.e., the number of electrons, in the floating gate.” Here, Tran21 describes performing a vector by matrix multiplication operation in analog memory cells.
Later, see Tran21 in [0099] describe “The output generated by input VMM system 32 a is provided as an input to the next VMM system (hidden level 1) 32 b, which in turn generates an output that is provided as an input to the next VMM system (hidden level 2) 32 c, and so on. The various layers of VMM system 32 function as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM system 32 a, 32 b, 32 c, 32 d, and 32 e can be a stand-alone, physical system comprising a respective non-volatile memory array, or multiple VMM systems could utilize different portions of the same physical non-volatile memory array, or multiple VMM systems could utilize overlapping portions of the same physical non-volatile memory array.” Here, Tran21 mentions that the vector by matrix multiplication calculations can be performed layer by layer within a neural network, and accounts for an analog CIM system perform calculations on a layer of the neural network.
Further, Tran21 teaches “and performing vector-by-matrix multiplication operations in a second layer of a neural network using …of the analog computation-in-memory engine, the digital computation-in-memory engine, …”
See Tran 21 in [0071] describe "..Digital non-volatile memories are well known. ...discloses an array of split gate non-volatile memory cells, which are a type of flash memory cells. Such a memory cell 210 is shown in FIG. 2.,” Also, see Tran21 in [0098] also mention " FIG. 10 is a block diagram depicting the usage of numerous layers of VMM systems 32, here labeled as VMM systems 32 a, 32 b, 32 c, 32 d, and 32 e. As shown in FIG. 10, the input, denoted Inputx, is converted from digital to analog by a digital-to-analog converter 31, and provided to input VMM system 32 a. The converted analog inputs could be voltage or current. The input D/A conversion for the first layer could be done by using a function or a LUT (look up table) that maps the inputs Inputx to appropriate analog levels for the matrix multiplier of input VMM system 32 a. The input conversion could also be done by an analog to analog (A/A) converter to convert an external analog input to a mapped analog input to the input VMM system 32 a. The input conversion could also be done by a digital-to-digital pul[s]e (D/P) converter to convert an external digital input to a mapped digital pulse or pulses to the input VMM system 32 ." Here, Tran21 mentions that vector by matrix multiplication or VMM can be performed in digital CIM systems for digital inputs.
Also, see Tran21 in paragraphs [0006-0007] mentions “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs…Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells used in this manner can be referred to as a vector by matrix multiplication (VMM) array... Each non-volatile memory cells used in the analog neural memory system must be erased and programmed to hold a very specific and precise amount of charge, i.e., the number of electrons, in the floating gate.” Here, Tran21 describes performing a vector by matrix multiplication operation in analog memory cells.
Later, see Tran21 in [0099] describe “The output generated by input VMM system 32 a is provided as an input to the next VMM system (hidden level 1) 32 b, which in turn generates an output that is provided as an input to the next VMM system (hidden level 2) 32 c, and so on. The various layers of VMM system 32 function as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM system 32 a, 32 b, 32 c, 32 d, and 32 e can be a stand-alone, physical system comprising a respective non-volatile memory array, or multiple VMM systems could utilize different portions of the same physical non-volatile memory array, or multiple VMM systems could utilize overlapping portions of the same physical non-volatile memory array.” Here, Tran21 mentions that the vector by matrix multiplication calculations can be performed layer by layer within a neural network, and accounts for an analog CIM system perform calculations on a second layer of the neural network.
However, Tran21 did not teach “performing … operations in a first layer of a neural network using … a dynamic weight engine;” or “and performing … operations in a second layer of a neural network using another …engine, and the dynamic weight engine”.
In an analogous art, Badaroglu teaches “and performing … operations in a second layer of a neural network using another … engine, …”
See Badaroglu in [0123] describe “Clause 12: The neural network circuit of Clause 9 or 10, wherein the at least one of the plurality of CIM PEs is in a first neural network layer and wherein the at least one of the plurality of NPU PEs is in a second neural network layer, different from the first neural network layer.” Badaroglu shows CIM PEs or CIM systems perform operations in a first layer of a neural network.
Further, see Badaroglu in [0128] describe “Clause 17: The neural network circuit of Clause 14 or 15, wherein the at least one of the plurality of NPU PEs is in a first neural network layer and wherein the at least one of the plurality of CIM PEs is in a second neural network layer, different from the first neural network layer.” Here, Badaroglu explicitly mentions the CIM PEs or compute-in memory processing elements (which examiner interprets to be synonymous with a CIM system or engine), perform operations using one CIM system on one layer of the neural network, then using a different CIM system on another layer (could be a second layer ) of the neural network. The examiner construes the term ‘another’ to be a completely separate or different type of CIM system or engine listed among the limitation ‘analog computation-in-memory engine, the digital computation-in-memory engine, and the dynamic weight engine’ that perform operations in a neural network layer.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Tran21 and incorporate into the teachings of Badaroglu since the references teach using a type of computation-in-memory system to perform operations within various layers in a neural network using different CIMs.
One of ordinary skill in the art would be motivated to do so because “CIM units are generally better than NPUs in terms of energy efficiency, whereas NPUs are generally better than CIM units for depth-wise convolution. Due to bit-serial operation, digital compute-in-memory (DCIM) units may have lower TOPS, but comparable or better performance for a given area (e.g., in terms of TOPS/mm2) than NPUs,” (see Badaroglu in [0082]).
However, Tran21 in view of Badaroglu, did not teach “performing … operations in a first layer of a neural network using … a dynamic weight engine;” or “and performing … operations in a second layer of a neural network using …, the dynamic weight engine”
In an analogous method, Freye teaches “performing … operations in a first layer of a neural network using … a dynamic weight engine;” or “and performing … operations in a second layer of a neural network using …, the dynamic weight engine”
See Freye on page 120 in Section II. Cascaded TDCIM Architecture describe “the operation of the typical convolution layer is shown in (2), where x is the input activation and w is the weight vector. f is the activation function and usually the binarize function for binary neural networks (BNNs) and the rectified linear unit (RELU) function for CNNs
Z=f(w⋅x).(2)
In TD computing, cascaded variable delay elements can implement an accumulation. Each element realizes the delay to encode one multiplication result, thus realizing the MAC function. Unlike in the traditional digital circuits, the convolution result is therefore presented as an accumulated delay. For the TDCIM architecture, the weights of one kernel are stored in the memory of one computing chain. For a kernel size of N with M computing chains in parallel, the total area, Atot , is given by M⋅N⋅Acell+ATDC with the cell area, Acell , and the area for time-to-digital converter (TDC), ATDC . After computing a convolution, the activation signals are changed, whereas the weights can remain in memory. The weights are only updated after the complete output feature map is computed, thus reducing the data movement from the main memory. The input activation can be shared by all the computing chains and further reduce the data movement.” Here, Freye mentions a time domain compute-in memory or TDCIM that relates to a third separate engine that perform operations by helping store and update weights within a CNN neural network, which relates to being a dynamic weight engine. The examiner construes dynamic weight to mean any value or quantity that can be updated, see reference from specification in [0153] stating “dynamic weight engine 3603 is a device whose stored weights can be modified by changing a bias voltage or bias current without performing a separate erase or program operation. Since the weights can be modified without performing a separate erase or program operation, these weights are considered dynamic weights.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Tran21 and Badaroglu, and incorporate with the teachings of Freye by using the teachings of Tran21 and Badaroglu of using a computation-in-memory system to perform operations within a neural network layer, with Freye’s teaching of using a dynamic weight engine.
One of ordinary skill in the art would be motivated to do so because by integrating Freye’s framework into the methods of Tran21 and Badaroglu, one with ordinary skill in the art would achieve “In this domain, analog computing is often considered to decrease power consumption further. ... A different compute scheme is time-domain CIM (TDCIM). In time-domain (TD) computing, the values are encoded as discrete arrival times of signal edges. While signaling is sample discrete, the arrival time is fundamentally continuous. Similar to charge and current, time is inherently additive, allowing for efficient accumulation operations,” (see Freye in page 119, second paragraph of Introduction).
Claim 24 is rejected under 35 U.S.C. 103 over Tran21 in view of Freye.
Claim 24:
Regarding claim 24, Tran21 teaches “A method comprising: storing weights received from an analog computation-in-memory engine in one or more of a digital computation-in-memory engine, a digital computation engine, and a dynamic weight engine.”
See Tran 21 in [0071] describe "..Digital non-volatile memories are well known. ...discloses an array of split gate non-volatile memory cells, which are a type of flash memory cells. Such a memory cell 210 is shown in FIG. 2.,” Also, see Tran21 in [0098] also mention " FIG. 10 is a block diagram depicting the usage of numerous layers of VMM systems 32, here labeled as VMM systems 32 a, 32 b, 32 c, 32 d, and 32 e. As shown in FIG. 10, the input, denoted Inputx, is converted from digital to analog by a digital-to-analog converter 31, and provided to input VMM system 32 a. The converted analog inputs could be voltage or current. The input D/A conversion for the first layer could be done by using a function or a LUT (look up table) that maps the inputs Inputx to appropriate analog levels for the matrix multiplier of input VMM system 32 a. The input conversion could also be done by an analog to analog (A/A) converter to convert an external analog input to a mapped analog input to the input VMM system 32 a. The input conversion could also be done by a digital-to-digital pul[s]e (D/P) converter to convert an external digital input to a mapped digital pulse or pulses to the input VMM system 32 ." Here, Tran21 mentions that vector by matrix multiplication or VMM can be performed in digital CIM systems for digital inputs.
Also, see Tran21 in paragraphs [0006-0007] mentions “The non-volatile memory arrays operate as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and to generate therefrom a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs… Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells used in this manner can be referred to as a vector by matrix multiplication (VMM) array... Each non-volatile memory cells used in the analog neural memory system must be erased and programmed to hold a very specific and precise amount of charge, i.e., the number of electrons, in the floating gate.” Here, Tran21 describes performing a vector by matrix multiplication operation in analog memory cells. Tran21 also mentions each memory cell from a CIM system can store weights.
However, Tran21 did not teach “A method comprising: storing weights received from an analog computation-in-memory engine in one or more of a digital computation-in-memory engine, a digital computation engine, and a dynamic weight engine.”
In an analogous art, Freye teaches “A method comprising: storing weights received from an analog computation-in-memory engine in one or more of a digital computation-in-memory engine, a digital computation engine, and a dynamic weight engine.”
See Freye on page 120 in Section II. Cascaded TDCIM Architecture describe “the operation of the typical convolution layer is shown in (2), where x is the input activation and w is the weight vector. f is the activation function and usually the binarize function for binary neural networks (BNNs) … For the TDCIM architecture, the weights of one kernel are stored in the memory of one computing chain. For a kernel size of N with M computing chains in parallel, the total area, Atot , is given by M⋅N⋅Acell+ATDC with the cell area, Acell , and the area for time-to-digital converter (TDC), ATDC . After computing a convolution, the activation signals are changed, whereas the weights can remain in memory. The weights are only updated after the complete output feature map is computed, thus reducing the data movement from the main memory. The input activation can be shared by all the computing chains and further reduce the data movement.” Here, Freye mentions a time domain compute-in memory or TDCIM system that relates to a third separate engine that perform operations by storing and updating weights (i.e. a method in storing weights) within a CNN neural network, which relates to being a dynamic weight engine. The examiner construes dynamic weight to mean any value or quantity that can be updated, see reference from specification in [0153] stating “dynamic weight engine 3603 is a device whose stored weights can be modified by changing a bias voltage or bias current without performing a separate erase or program operation. Since the weights can be modified without performing a separate erase or program operation, these weights are considered dynamic weights.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Tran21, and incorporate with the teachings of Freye by using the teachings of Tran21 of using a computation-in-memory system to perform operations within a neural network layer, with Freye’s teaching of using a dynamic weight engine.
One of ordinary skill in the art would be motivated to do so because by integrating Freye’s framework into the methods of Tran21, one with ordinary skill in the art would achieve “In this domain, analog computing is often considered to decrease power consumption further. ... A different compute scheme is time-domain CIM (TDCIM). In time-domain (TD) computing, the values are encoded as discrete arrival times of signal edges. While signaling is sample discrete, the arrival time is fundamentally continuous. Similar to charge and current, time is inherently additive, allowing for efficient accumulation operations,” (see Freye in page 119, second paragraph of Introduction).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WenWei Zeng/Examiner, Art Unit 2146
/SHAHID K KHAN/Primary Examiner, Art Unit 2146