Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-19 are pending for examination. Claims 1, 4, and 9 are independent.
Examiner notes
Examiner attempted to reach out to Vincent De Luca and propose an Examiners amendment for this application. No response was received at the time.
Response to Arguments
Applicant's arguments filed 03/23/2026 have been fully considered but they are not fully persuasive.
Applicant arguments regarding 35 U.S.C. § 101:
Claim 1 is directed to processing data according to a neural network process.
Contrary to the assertion at page 2 of the Office action ("[t]his involves a human
performing a set of elementary neural network operations") a human mind cannot
perform a neural network process. Claim 1 further requires processing the data
according to the neural network process, using the fixed-function circuitry of the
hardware accelerator to perform the neural network with the associated pooling
function. Thus, the claim language itself refutes the Examiner's contention that the
claim is seeking to patent a mental process. The same analysis applies to method
claims 4 and 9. Claims 13 - 15 are directed to a statutory article of manufacture, that
stores computer readable code which causes a computer to perform a useful function
when run. These claims are not directed to a "mental process." Claims 16 - 19 are
drawn to a data processing system including the structural combination of a hardware
accelerator and a controller; claims 16 - 19 are clearly not directed to a "mental
process." […]
The present claims fall within this latter category of patent-eligible improvements to computer technology. The claimed invention provides a novel approach to implementing binary argmax/argmin functions (and related functions) in hardware accelerators. In particular, the proposed approach improves the functioning of a hardware accelerator with fixed-function circuitry by at least improving or expanding the capabilities of the fixed function circuitry to perform a binary argmax, binary argmin or a related function.
The claimed invention recognizes that a hardware accelerator has a limited set of
functions. As explained in the Specification (para. [0050]): "Unlike the operations
performed by a CPU, the (neural network) operations performed by an NNA are not
designed to be a flexible or complete set of general-purpose operations. Instead, each
neural network operation is specialised to perform a particular computationally intensive
neural network calculation quickly and efficiently."
The claims make use of this specialized hardware in a novel way to implement a binary argmax/argmin function (or related function), providing concrete improvements in
neural network processing efficiency and hardware resource utilization.
Examiners response: Examiner respectfully disagrees, under broadest reasonable interpretation there are neural network processes (e.g. elementary operation or making predictions) that can be performed by the human mind. Claim limitations describing “using the fixed-function circuitry of the hardware accelerator to perform the neural network” and using a neural network are understood as adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f).
Examiner respectfully disagrees, the current claim language does not reflect specialized hardware and instead broadly describes a hardware accelerator and fixed-function circuitry. These elements are understood to be generic computer elements (See MPEP 2106.05(f)), Further detailing the circuity and performed actions of the circuity could overcome the 35 U.S.C. § 101 rejection.
Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. MPEP 2106.05(a) also indicates not only does the specification have to disclose the improvement for a claim not to be an abstract idea, the claim elements themselves have to provide limitations that reflect the improvement.
Applicants claimed improvement is describing an improvement to an abstract idea and not to an improvement to a computer or technical field. MPEP 2106.05(a) says an improvement in the abstract idea itself is not an improvement in technology. The claim limitations overall are a combination of mental steps under step 2A Prong 1, and additional elements under steps 2A Prong 2 & 2B as detailed in the 101 rejection below.
Applicant arguments regarding 35 U.S.C. § 103:
Nakahara
Claim 1 specifically requires "a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations." Nakahara's FPGA-based implementation is the antithesis of fixed-function circuitry because it is programmable logic that can be reconfigured to implement different functions (Yu similarly fails to teach fixed-function circuitry, as Yu also discloses an FPGA-based implementation, which is reconfigurable logic rather than fixed-function circuitry).
Applicant further submits that the Office action has failed to establish that any of the cited references teaches "mapping the pooling function to a set of elementary neural network operations" as recited by claim 1. […]
Yu
Applicant also submits that the characterization of Yu's pooling indices as a one-hot vector is inaccurate. Yu explicitly discloses that indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2x2 pooling window. A 2-bit encoded index is fundamentally different from a one-hot vector. A one-hot vector for a 2x2 pooling window would require 4 bits, with exactly one bit set to 1 and all others set to 0.
Claim 1 requires that the pooling function outputs a maximum or minimum value of data input to the pooling function and a one-hot vector identifying the index of the maximum or minimum value as a single integrated operation. Neither Nakahara nor Yu teaches a pooling function that outputs both the pooled value and a one-hot vector as part of the same pooling operation. Nakahara discloses max pooling operations but does not teach outputting any index information, let alone a one-hot vector, as part of the pooling function output. Yu's 2-bit encoded indices are merely a compact storage format for position information, not a one-hot vector as required by claim 1.
Satti
Satti is cited only for teaching an element-wise minimum operation and does not cure the deficiencies of Nakahara and Yu regarding the fixed-function circuitry, the mapping of the pooling function to elementary neural network operations, or the one-hot vector output of the pooling function.
Accordingly, claim 1 is considered novel and inventive over the cited prior art, in full compliance with 35 U.S.C. § 103. Claims 2, 3, 13, 16 and 19 are considered novel and inventive at least by virtue of their dependence on a novel and inventive independent claim 1.
Examiner response: Examiner respectfully disagrees the FPGA is mapped to a hardware accelerator, with further components within the FPGA that can be considered fixed-function circuitry (e.g. pooling circuit), which is broadly stated. Under broadest reasonable interpretation, examiner respectfully disagrees performing pooling operations involve performing elementary neural network operations. Yu describes pooing indices for a max pooing input.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Subject Matter Eligibility Analysis Step 1:
Claim 1 recites “A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations….” which is a process, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 1 recites the steps:
“A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations….”: This involves a human performing a set of elementary neural network operations; therefore, this is a mental process.
“mapping the pooling function to a set of elementary neural network operations, wherein the set of elementary neural network operations comprises only elementary neural network operations from the set of available elementary neural network operations”: This involves a human mapping a pooling function to a set of neural network operations. Thus, this is a mental process.
“wherein each of the set of elementary neural network operations is selected from a list consisting of: an element-wise subtraction operation, an element-wise addition operation, an element-wise multiplication operation, an element-wise maximum operation, an element-wise minimum operation, a max pooling operation or min pooling operation, a magnitude operation, one or more lookups operations using one or more look-up tables, a convolution operation, and a deconvolution operation”: This involves a human performing an operation from a list in step (I). Therefore, this is a mental process.
Claim 1 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 1 discloses the additional elements:
“receiving a definition of a neural network process to be performed on the data, the neural network process comprising a neural network with an associated pooling function, wherein the pooling function outputs a maximum or minimum value of data input to the pooling function and a one-hot vector identifying the index of the maximum or minimum value of the data input to the pooling function”: This element does not integrate the element into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated pooling function, wherein the pooling function is performed using the set of elementary neural network operations”: This element does not integrate the element into a practical application because the element recites generic computing components (neural network and hardware accelerator) to perform the abstract ideas (MPEP 2106.05(f)).
“the data comprises image data and/or audio data”: This element does not integrate into a practical application because it further defines the input data in the additional element (I) that was received. Hence, the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
Thus, claim 1 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 1 do not provide more than the abstract ideas themselves taken alone and in combination because:
“receiving a definition of a neural network process to be performed on the data, the neural network process comprising a neural network with an associated pooling function, wherein the pooling function outputs a maximum or minimum value of data input to the pooling function and a one-hot vector identifying the index of the maximum or minimum value of the data input to the pooling function”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information])
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated pooling function, wherein the pooling function is performed using the set of elementary neural network operations”: This element recites generic computing components (neural network and hardware accelerator) to perform the abstract ideas (MPEP 2106.05(f)).
“the data comprises image data and/or audio data”: This element does not integrate into a practical application because it further defines the input data in the additional element (I) that was received. Hence, the element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 1 is subject-matter ineligible.
Regarding claim 2:
Subject Matter Eligibility Analysis Step 1:
Claim 2 is a process as in claim 1
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental concepts in claim 1, claim 2 recites:
“the data input to the pooling function is an input tensor, and the set of elementary neural network operations implements, to perform the pooling function”: This involves a human performing the pooling function using the elementary neural network operations from claim 1. Therefore, this is a mental process.
“a maximum or minimum pooling operation, applied to the input tensor, that identifies the maximum or minimum value contained in the input tensor”: This involves a human implementing a maximum/minimum pooling operation to an input tensor by identifying the maximum/minimum value in the input tensor. Hence, this is a mental process.
Thus claim 2 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 1, claim 2 recites:
“the data input to the pooling function is an input tensor, and the set of elementary neural network operations implements, to perform the pooling function”: This element does not integrate the element into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
“a binary argmax or binary argmin function that outputs a one-hot vector representing an argmax or argmin of the input tensor”: This element does not integrate the element into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g))
Claim 2 therefore is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 2 do not provide more than the abstract ideas themselves taken alone and in combination because:
“the data input to the pooling function is an input tensor, and the set of elementary neural network operations implements, to perform the pooling function”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]).
“a binary argmax or binary argmin function that outputs a one-hot vector representing an argmax or argmin of the input tensor”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 2 is subject-matter ineligible.
Regarding claim 3:
Subject Matter Eligibility Analysis Step 1:
Claim 3 is a process as in claim 2.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental concepts in claim 2, claim 3 recites:
“an argmax or argmin operation, applied to the input tensor, that identifies an index of a maximum value or minimum value respectively of the input tensor”: This involves a human applying an argmax/argmin operation to an input tensor by identifying the index of the maximum/minimum value of the input tensor, therefore this is a mental process.
“a subtraction operation, applied to each element of the first intermediate vector, that subtracts the identified index of the input tensor from the said element of the first intermediate vector to produce a second intermediate vector”: This involves a human applying a subtraction operation to each element of an immediate vector by subtracting the index from step (I) with each element of the immediate vector in order to produce a second immediate vector, thus this is a mental process.
“a zero-identification operation, applied to the second intermediate vector, that replaces any zero values in the second intermediate vector with a first binary value and any non-zero values in the second intermediate vector with a second, different binary value, to thereby produce the one-hot vector”: This involves a human replacing any zero values from the vector in step (II) with a binary value and any non-zero value with another binary value to perform a zero-identification operation. Thus, this is a mental process.
Claim 3, hence recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 2, claim 3 recites:
“a first intermediate vector obtaining operation that obtains a first intermediate vector having a same number of entries as the input tensor, each entry of the first intermediate vector containing an index value of a different entry of the input tensor”: This element does not integrate into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g))
Claim 3 therefore is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional element in claim 3 does not provide more than the abstract ideas themselves taken alone and in combination because:
“a first intermediate vector obtaining operation that obtains a first intermediate vector having a same number of entries as the input tensor, each entry of the first intermediate vector containing an index value of a different entry of the input tensor”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 3 is subject-matter ineligible.
Regarding claim 4:
Subject Matter Eligibility Analysis Step 1:
Claim 4 recites “A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations….” which is a process, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 4 recites the steps:
“A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations….”: This involves a human performing a set of elementary neural network operations; therefore, this is a mental process.
“receiving a definition of a neural network process to be performed on the data, the neural network process comprising a neural network with an associated unpooling or backward pooling function, wherein the unpooling or backward pooling function is configured to map an input value to an original position in a tensor using a one-hot vector that represents an argmax or argmin of a previous pooling function”: This involves a human mapping an input value to a position in a tensor using a one-hot vector from a pooling function, thus this is a mental process.
“mapping the unpooling or backward pooling function to a set of elementary neural network operations, wherein the set of elementary neural network operations comprises only elementary neural network operations from the set of available elementary neural network operations”: This involves a human mapping a pooling function to a set of neural network operations. Thus, this is a mental process.
“wherein each of the set of elementary neural network operations is selected from a list consisting of: an element-wise subtraction operation, an element-wise addition operation, an element-wise multiplication operation, an element-wise maximum operation, an element-wise minimum operation, a max pooling operation or min pooling operation, a magnitude operation, one or more lookups operations using one or more look-up tables, a convolution operation, and a deconvolution operation”: This involves a human performing an operation from a list in step (I). Therefore, this is a mental process.
Claim 4 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 4 discloses the additional elements:
“receiving a definition of a neural network process to be performed on the data, the neural network process comprising a neural network with an associated unpooling or backward pooling function, wherein the unpooling or backward pooling function is configured to map an input value to an original position in a tensor using a one-hot vector that represents an argmax or argmin of a previous pooling function”: This element does not integrate the element into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated unpooling or backward pooling function, wherein the pooling function is performed using the set of elementary neural network operations”: This element does not integrate the element into a practical application because the element recites generic computing components (neural network and hardware accelerator) to perform the abstract ideas (MPEP 2106.05(f)).
“the data comprises image data and/or audio data”: This element does not integrate into a practical application because it further defines the input data in the additional element (I) that was received. Hence, the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
Thus, claim 4 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 4 do not provide more than the abstract ideas themselves taken alone and in combination because:
“receiving a definition of a neural network process to be performed on the data, the neural network process comprising a neural network with an associated unpooling or backward pooling function, wherein the unpooling or backward pooling function is configured to map an input value to an original position in a tensor using a one-hot vector that represents an argmax or argmin of a previous pooling function”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information])
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated unpooling or backward pooling function, wherein the pooling function is performed using the set of elementary neural network operations”: This element recites generic computing components (neural network and hardware accelerator) to perform the abstract ideas (MPEP 2106.05(f)).
“the data comprises image data and/or audio data”: This element does not integrate into a practical application because it further defines the input data in the additional element (I) that was received. Hence, the element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 4 is subject-matter ineligible.
Regarding claim 5:
Subject Matter Eligibility Analysis Step 1:
Claim 5 is a process as in claim 4.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental concepts in claim 4, claim 5 discloses:
“a multiplication function configured to multiply each entry in the one-hot vector by the input value to produce a product one-hot vector”: This involves a human producing a one-hot vector by multiplying each entry with the input value from claim 4. Therefore, this is a mental process.
Hence, claim 5 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 4, claim 5 mentions:
“a binary argmax/argmin acquisition function, configured to obtain the one-hot vector representing an argmax or argmin of a previous pooling function” : This element does not integrate the element into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g))
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor”: This element does not integrate the element into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g))
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor”: This element does not integrate the element into a practical application because the element recites a generic computing component (binary constant filter) to perform the abstract ideas (MPEP 2106.05(f)).
Claim 5 therefore is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 5 do not provide more than the abstract ideas themselves taken alone and in combination because:
“a binary argmax/argmin acquisition function, configured to obtain the one-hot vector representing an argmax or argmin of a previous pooling function” : This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information])
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information])
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor”: This element does not integrate the element into a practical application because the element recites a generic computing component (binary constant filter) to perform the abstract ideas (MPEP 2106.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 5 is subject-matter ineligible.
Regarding claim 6:
Subject Matter Eligibility Analysis Step 1:
Claim 6 is a process as in claim 5.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental concepts in claim 5, claim 6 mentions:
“the multiplication function is performed using an element-wise multiplication operation.”: This involves a human using element-wise multiplication as the multiplication function. Thus, this is a mental process.
Therefore, claim 6 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 6 recites the same additional elements as claim 5, thus claim 6 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
Since claim 6 recites the same additional elements as claim 5 and there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 6 is subject-matter ineligible.
Regarding claim 7:
Subject Matter Eligibility Analysis Step 1:
Claim 7 is a process as in claim 5.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental concepts in claim 5, claim 7 mentions:
“deconvolution function is performed using a deconvolution operation.”: This involves a human using a deconvolution operation (e.g. blurring, adding noise, etc.) as the deconvolution function. Thus, this is a mental process.
Therefore, claim 7 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 7 recites the same additional elements as claim 5, thus claim 7 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
Since claim 7 recites the same additional elements as claim 5 and there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 7 is subject-matter ineligible.
Regarding claim 8:
Subject Matter Eligibility Analysis Step 1:
Claim 8 is a process as in claim 5.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 8 recites the same mental concepts as claim 5. Hence, claim 8 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the mental concepts in claim 5, claim 8 recites:
“deconvolution function is a grouped deconvolution function.”: This element does not integrate the abstract ideas into a practical application because it further defines the deconvolution function which generates the output in claim 5, thus this element is also an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
Claim 8 therefore is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional element does not provide significantly more than the abstract ideas themselves, taken alone and in combination because:
“deconvolution function is a grouped deconvolution function.”: This element further defines the deconvolution function which generates the output in claim 5, thus this element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information])
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 8 is subject-matter ineligible.
Regarding claim 9:
Subject Matter Eligibility Analysis Step 1:
Claim 9 recites “A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations….” which is a process, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 9 recites the steps:
“A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations….”: This involves a human performing a set of elementary neural network operations; therefore, this is a mental process.
“mapping the binary argmax or binary argmin function to a set of elementary neural network operations, wherein the set of elementary neural network operations comprises only elementary neural network operations from the set of available elementary neural network operations”: This involves a human mapping a binary argmax or binary argmin function to a set of neural network operations. Thus, this is a mental process.
“wherein each of the set of elementary neural network operations is selected from a list consisting of: an element-wise subtraction operation, an element-wise addition operation, an element-wise multiplication operation, an element-wise maximum operation, an element-wise minimum operation, a max pooling operation or min pooling operation, a magnitude operation, one or more lookups operations using one or more look-up tables, a convolution operation, and a deconvolution operation”: This involves a human performing an operation from a list in step (I). Therefore, this is a mental process.
Claim 9 therefore recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 9 discloses the additional elements:
“receiving a definition of a neural network process to be performed, the neural network process comprising a neural network with an associated binary argmax or binary argmin function”: This element does not integrate the element into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated binary argmax or binary argmin function, wherein the binary argmax or binary argmin function is performed using the set of elementary neural network operations”: This element does not integrate the element into a practical application because the element recites generic computing components (neural network and hardware accelerator) to perform the abstract ideas (MPEP 2106.05(f)).
“the data comprises image data and/or audio data”: This element does not integrate into a practical application because it further defines the input data in the additional element (I) that was received. Hence, the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)).
Thus, claim 9 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 9 do not provide more than the abstract ideas themselves taken alone and in combination because:
“receiving a definition of a neural network process to be performed, the neural network process comprising a neural network with an associated binary argmax or binary argmin function”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information])
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated binary argmax or binary argmin function, wherein the binary argmax or binary argmin function is performed using the set of elementary neural network operations”: This element recites generic computing components (neural network and hardware accelerator) to perform the abstract ideas (MPEP 2106.05(f)).
“the data comprises image data and/or audio data”: This element does not integrate into a practical application because it further defines the input data in the additional element (I) that was received. Hence, the element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 9 is subject-matter ineligible.
Regarding claim 10:
Subject Matter Eligibility Analysis Step 1:
Claim 10 is a process as in claim 9.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental concepts in claim 9, claim 10 recites:
“an argmax or argmin operation, applied to a first input tensor, that identifies an index of a maximum value or minimum value respectively of the first input tensor”: This involves a human applying an argmax/argmin operation to an input tensor by identifying the index of the maximum/minimum value of the input tensor, therefore this is a mental process.
“a subtraction operation, applied to each element of the first intermediate vector, that subtracts the identified index of the first input tensor from the said element of the first intermediate vector to produce a second intermediate vector”: This involves a human applying a subtraction operation to each element of an immediate vector by subtracting the index from step (I) with each element of the immediate vector in order to produce a second immediate vector, thus this is a mental process.
“a zero-identification operation, applied to the second intermediate vector, that replaces any zero values in the second intermediate vector with a first binary value and any non-zero values in the second intermediate vector with a second, different binary value, to thereby produce the one-hot vector”: This involves a human replacing any zero values from the vector in step (II) with a binary value and any non-zero value with another binary value to perform a zero-identification operation. Thus, this is a mental process.
Claim 10 hence recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 9, claim 10 recites:
“a vector obtaining operation that obtains a first intermediate vector having a same number of entries as the first input tensor, each entry of the first intermediate vector containing an index value of a different entry of the input tensor”: This element does not integrate into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g))
Claim 10 therefore is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional element in claim 10 does not provide more than the abstract ideas themselves taken alone and in combination because:
“a vector obtaining operation that obtains a first intermediate vector having a same number of entries as the first input tensor, each entry of the first intermediate vector containing an index value of a different entry of the input tensor”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1369 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 10 is subject-matter ineligible.
Regarding claim 11:
Subject Matter Eligibility Analysis Step 1:
Claim 11 is a process as in claim 10.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 11 recites the same mental concepts as claim 10, thus claim 11 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 10, claim 11 mentions:
“the first intermediate vector is stored in a look-up table”: This element does not integrate the abstract ideas into a practical application because the element mentions an insignificant extra solution activity of data storage (MPEP 2106.05 (g)).
Thus, claim 11 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 11 does not provide significantly more than the abstract ideas themselves taken alone and in combination because:
“the first intermediate vector is stored in a look-up table”: This element mentions a well-understood, routine and conventional activity of “storing and retrieving information in memory” (MPEP 2106.05(d)(IV) “Storing and retrieving information in memory”, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93)
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 11 is subject-matter ineligible.
Regarding claim 12:
Subject Matter Eligibility Analysis Step 1:
Claim 12 is a process as in claim 10.
Subject Matter Eligibility Analysis Step 2A Prong 1:
In addition to the mental concepts in claim 10, claim 12 mentions:
“a binary maximum/minimum operation, applied to the first input tensor, to produce a binary tensor of the same spatial size as the first input tensor, wherein each element in the binary tensor: corresponds to a different element of the first input tensor, and contains a binary value indicating whether or not the corresponding element of the first input tensor has a value equal to the maximum/minimum value contained in the first input tensor”: This involves a human performing a binary maximum/minimum operation to the first input tensor to produce a binary tensor where each element corresponds to a binary value and is different from the elements in the first tensor. Thus, this is a mental process.
“an integer index operation, applied to the binary tensor, that identifies one or more indexes of the binary tensor, the identified one or more indexes being indexes of the one or more elements of the binary tensor that have a binary value that indicates the corresponding element of the first input tensor has a value equal to the maximum/minimum value contained in the first input tensor”: This involves a human performing an integer index operation on the tensor from step (I) by identifying one or more indexes that corresponds to the maximum/minimum value contained in the first input tensor. Therefore, this is a mental process.
“a tie elimination operation, applied to the identified indexes, that selects a single one of the one or more identified indexes to provide the output of the argmax or argmin function”: This involves a human selecting an index from step (II) as the output, thus this is a mental process.
Therefore, claim 12 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 12 recites the same additional elements as claim 10, thus claim 12 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
Since claim 12 recites the same additional elements as claim 10 and there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 12 is subject-matter ineligible.
Regarding claim 13:
Subject Matter Eligibility Analysis Step 1:
Claim 13 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Since claim 13 recites the same mental process as claim 1, claim 13 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional element in claim 1, claim 13 mentions:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 1 to be performed when the code is run”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
Hence, claim 13 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 13 does not provide more than the abstract ideas themselves taken alone and in combination because:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 1 to be performed when the code is run”: This element recites a generic computing component (MPEP 2106.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 13 is subject-matter ineligible.
Regarding claim 14:
Subject Matter Eligibility Analysis Step 4:
Claim 14 is a process as in claim 4.
Subject Matter Eligibility Analysis Step 2A Prong 4:
Since claim 14 recites the same mental process as claim 4, claim 14 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional element in claim 4, claim 14 mentions:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 4 to be performed when the code is run”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2406.05(f)).
Hence, claim 14 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 14 does not provide more than the abstract ideas themselves taken alone and in combination because:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 4 to be performed when the code is run”: This element recites a generic computing component (MPEP 2406.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 14 is subject-matter ineligible.
Regarding claim 15:
Subject Matter Eligibility Analysis Step 9:
Claim 15 is a process as in claim 9.
Subject Matter Eligibility Analysis Step 2A Prong 9:
Since claim 15 recites the same mental process as claim 9, claim 15 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional element in claim 9, claim 15 mentions:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 9 to be performed when the code is run”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2906.05(f)).
Hence, claim 15 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 15 does not provide more than the abstract ideas themselves taken alone and in combination because:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 9 to be performed when the code is run”: This element recites a generic computing component (MPEP 2906.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 15 is subject-matter ineligible.
Regarding claim 16:
Subject Matter Eligibility Analysis Step 1:
Claim 16 is a process as in claim 1.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Since claim 16 recites the same mental process as claim 1, claim 16 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 1, claim 16 mentions:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operation”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
“a controller configured to perform the method as set forth in claim 1”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
Hence, claim 16 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 16 does not provide more than the abstract ideas themselves taken alone and in combination because:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operation”: This element recites a generic computing component (MPEP 2106.05(f)).
“a controller configured to perform the method as set forth in claim 1”: This recites a generic computing component (MPEP 2106.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 16 is subject-matter ineligible.
Regarding claim 17:
Subject Matter Eligibility Analysis Step 4:
Claim 17 is a process as in claim 4.
Subject Matter Eligibility Analysis Step 2A Prong 4:
Since claim 17 recites the same mental process as claim 4, claim 17 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 4, claim 17 mentions:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operation”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
“a controller configured to perform the method as set forth in claim 4”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
Hence, claim 17 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 17 does not provide more than the abstract ideas themselves taken alone and in combination because:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operation”: This element recites a generic computing component (MPEP 2106.05(f)).
“a controller configured to perform the method as set forth in claim 4”: This recites a generic computing component (MPEP 2106.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 17 is subject-matter ineligible.
Regarding claim 18:
Subject Matter Eligibility Analysis Step 9:
Claim 18 is a process as in claim 9.
Subject Matter Eligibility Analysis Step 2A Prong 9:
Since claim 18 recites the same mental process as claim 9, claim 18 recites abstract ideas.
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 9, claim 18 mentions:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operation”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
“a controller configured to perform the method as set forth in claim 9”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
Hence, claim 18 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 18 does not provide more than the abstract ideas themselves taken alone and in combination because:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operation”: This element recites a generic computing component (MPEP 2106.05(f)).
“a controller configured to perform the method as set forth in claim 9”: This recites a generic computing component (MPEP 2106.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 18 is subject-matter ineligible.
Regarding claim 19:
Subject Matter Eligibility Analysis Step 1:
Claim 19 is a process as in claim 14.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 19 contains the same mental concepts as claim 14, thus claim 19 recites abstract ideas
Subject Matter Eligibility Analysis Step 2A Prong 2:
In addition to the additional elements in claim 14, claim 19 discloses:
“an activation unit, comprising an LUT”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
“a local response normalisation unit, configured to perform a local response normalisation”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
“an element-wise operations unit, configured to apply a selected operation to every pair of respective elements of two tensor of identical size”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
“one or more convolution engines, configured to perform convolution operations”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
“a pooling unit, configured to perform pooling operations, including max pooling and/or min pooling”: This element does not integrate the abstract ideas into a practical application because the element recites a generic computing component (MPEP 2106.05(f)).
Hence, claim 19 is directed to the abstract ideas.
Subject Matter Eligibility Analysis Step 2B:
The additional elements in claim 19 does not provide more than the abstract ideas themselves taken alone and in combination because:
“an activation unit, comprising an LUT”: This element recites a generic computing component (MPEP 2106.05(f)).
“a local response normalisation unit, configured to perform a local response normalisation”: This element recites a generic computing component (MPEP 2106.05(f)).
“an element-wise operations unit, configured to apply a selected operation to every pair of respective elements of two tensor of identical size”: This element recites a generic computing component (MPEP 2106.05(f)).
“one or more convolution engines, configured to perform convolution operations”: This element recites a generic computing component (MPEP 2106.05(f)).
“a pooling unit, configured to perform pooling operations, including max pooling and/or min pooling”: This element recites a generic computing component (MPEP 2106.05(f)).
Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 19 is subject-matter ineligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 2, 9, 13, 15, 16, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Nakahara et al.’s “A Fully Connected Layer Elimination for a Binarized Convolutional Neural Network on an FPGA” in view of Yu et al.’s “Optimizing FPGA-based Convolutional Encoder-Decoder Architecture for Semantic Segmentation” and in further view of Satti et al.’s “Min-Max Average Pooling Based Filter for Impulse Noise Removal”.
Regarding claim 1, Nakahara et al. teaches “A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations,” (Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA… As for the number of operations for the implemented CNN, sine it performed 3 × 3 MAC operations for 256 feature maps at 143 MHz, it was 329.47 GOPS (Giga Operations Per Seconds).”; “FPGA” corresponds to “a hardware accelerator”; “MAC operations” correspond to “elementary neural network operations”)“the method comprising:”
“mapping the pooling function to a set of elementary neural network operations, wherein the set of elementary neural network operations comprises only elementary neural network operations from the set of available elementary neural network operations”(Nakahara et al. Fig. 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “conv” stands for “convolutional”; since the feature maps are mapped to “conv+max pooling” involves pooling layers that use convolutional and max pooling, this corresponds to “mapping the pooling function to a set of elementary neural network operations”); and
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated pooling function, wherein the pooling function is performed using the set of elementary neural network operations” (Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA”; Fig. 5
PNG
media_image2.png
584
663
media_image2.png
Greyscale
displays “max-pooling” is performed on the FPGA which includes an “ARM Cortex A-9 Processor”, thus this corresponds to “ processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated pooling function”; since “max-pooling” performs a max pooling operation, this corresponds to “the pooling function is performed using the set of elementary neural network operations”);
“wherein the data comprises image data and/or audio data” (Nakamura et al. Section III “We designed both VGG11s by using Chainer [1]. Then, we trained the VGG11s using the CIFAR10 training images on the NVidia GTX Titan X.”) [note: since the limitation recites “and/or”, it can be read as “the data comprises image data”]; and
“wherein each of the set of elementary neural network operations is selected from a list consisting of:”
“an element-wise addition operation” (Nakahara et al. Section III
PNG
media_image3.png
230
686
media_image3.png
Greyscale
; Equation (2) mathematically defines an average pooling layer which sums each element labeled as “xK/2” thus this corresponds to “an element-wise addition operation”)
“an element-wise multiplication operation”(Nakahara et al. Table 1 defines a binarized multiplication at a pooling layer:
PNG
media_image4.png
248
344
media_image4.png
Greyscale
in which “x” corresponds to an “element” in the input data)
“a max pooling operation or min pooling operation” (Nakahara et al. Fig. 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “Max Pooling” corresponds to “a max pooling operation”)[note: since the limitation recites “or”, it can be read as “a max pooling operation”]
“a convolution operation” (Nakahara et al. Fig 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “CONV” corresponds to “a convolution operation”);
Nakahara et al. fails to teach:
“receiving a definition of a neural network process to be performed on the data, the neural network process comprising a neural network with an associated pooling function, wherein the pooling function outputs a maximum or minimum value of data input to the pooling function and a one-hot vector identifying the index of the maximum or minimum value of the data input to the pooling function”; and
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise subtraction operation…an element-wise maximum operation…an element-wise minimum operation…a magnitude operation…one or more lookups operations using one or more look-up tables…a deconvolution operation”
However, Yu et al. teaches:
“receiving a definition of a neural network process to be performed on the data, the neural network process comprising a neural network with an associated pooling function, wherein the pooling function outputs a maximum or minimum value of data input to the pooling function and a one-hot vector identifying the index of the maximum or minimum value of the data input to the pooling function” (Yu et al. Section II “Semantic segmentation CNNs, such as UNet [6], DeconvNet [7], SegNet [8], etc., generally have a convolutional encoder-decoder architecture… We use the SegNet-basic as an example to introduce our work. The architecture of SegNetbasic is shown in Figure 1. It has four encoders and four decoders. The un-pooling layer in a decoder uses the pooling indices generated by the max pooling layer in its corresponding encoder to complete the up-sampling… indices store the absolute position of the maximum value in the feature map before pooling, which requires multiple bits.”; since the “max pooling layer” outputs “pooling indices” which indicate the “maximum value”, “SegNetbasic” and “max pooling layer” corresponds to “a neural network with an associated pooling function, wherein the pooling function outputs a maximum or minimum value of data input to the pooling function” and “indices” correspond to “one-hot vector identifying the index of the maximum or minimum value of the data input to the pooling function”);
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise subtraction operation” (Yu et al. Section II Equation 4 displays bit truncation that is done during pooling:
PNG
media_image5.png
127
475
media_image5.png
Greyscale
; in which x’j represents each element in the input data and
PNG
media_image6.png
36
274
media_image6.png
Greyscale
is a subtraction operation that is performed at each x’j );
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise maximum operation”(Yu et al. Section II Equation (1):
PNG
media_image7.png
120
582
media_image7.png
Greyscale
; in which “xj” corresponds to “an element” in the input data );
“wherein each of the set of elementary neural network operations is selected from a list consisting of… a magnitude operation”(Yu et al. Section II Equation (1):
PNG
media_image7.png
120
582
media_image7.png
Greyscale
; since the magnitude of the maximum value (labeled as “|Max|”) is taken, this corresponds to “a magnitude operation”);
“wherein each of the set of elementary neural network operations is selected from a list consisting of… one or more lookups operations using one or more look-up tables”(Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption. In the Caffe [9] framework widely used on CPU and GPU, indices store the absolute position of the maximum value in the feature map before pooling, which requires multiple bits. Memory consumption is a stiff challenge for embedded platforms with limited storage resources. Therefore, we use a similar approach [10] as shown in Figure 2. Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2ൈ2 pooling window. For SegNetbasic, as shown in Table I, Caffe method of storing pooling indices for the entire network requires 6.72 MB, while our method only requires 0.87 MB, which reduces 87% memory footprint.”; Yu et al. Table I displays:
PNG
media_image8.png
383
595
media_image8.png
Greyscale
; since Table I stores pooling indices which are “used in the last decoder”, Table I corresponds to “a lookup table”); and
“wherein each of the set of elementary neural network operations is selected from a list consisting of… a deconvolution operation”(Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption….”; Yu et al. Abstract “However, CNNs for semantic segmentation usually contain some symmetrical encoders and decoders, corresponding to the down-sampling process (e.g., pooling, convolution) and the up-sampling process (e.g., unpooling, deconvolution)”; since “unpooling” and “deconvolution” are synonymous, “unpooling operation” corresponds to “a deconvolution operation”)
Nakahara et al. and Yu et al. are both analogous to the claimed invention because they both utilize FPGAs to run CNNs. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replace the neural network disclosed by Nakahara et al. with a neural network that contains a pooling function that outputs a one-hot vector as taught by Yu et al. Hence, this would be a substitution of one known element (neural network) to another (neural network with an associated pooling function) to obtain predictable results (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Both Nakahara et al. and Yu et al. fail to teach:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise minimum operation…”
However, Satti et al. teaches:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise minimum operation…””(Satti et al. Algorithm 1 defines the pooling layer as:
PNG
media_image9.png
383
596
media_image9.png
Greyscale
; line 9 takes the minimum of the elements in the window “W-c” which includes each element “Pi,j” for an image input data, thus this corresponds to “an element wise minimum operation”);
Nakahara et al., Yu et al., and Satti et al. are all analogous to the claimed invention because they are all in the same field of utilizing max pooling layers in CNN. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the operations listed by Nakahara et al. and Yu et al. to further include an element-wise minimum operation as taught by Satti et al. Doing so would remove noise from input images more effectively (Satti et al. Section IV).
Regarding claim 2, the rejections of claim 1 are incorporated. Yu et al. further teaches:
“a maximum or minimum pooling operation, applied to the input tensor, that identifies the maximum or minimum value contained in the input tensor” (Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
illustrates the pooling pipeline to generate pooling indices and a “max pooling output”; Yu et al. further mathematically defines “max pooling” in Section II Equation (1):
PNG
media_image11.png
114
587
media_image11.png
Greyscale
; since the max pooling layer uses a max function as shown in Equation (1), this corresponds to “a max…pooling operation” and “identifies the maximum value contained in the input tensor”; “max pooling input” corresponds to “input tensor”; ); and
“a binary argmax or binary argmin function that outputs a one-hot vector representing an argmax or argmin of the input tensor” (Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; Figure 2 displays that in addition to finding the maximum values of the “max pooling input”, the max pooling layer also outputs “pooling indices” which represent the indices of each entry in the max pooling output, thus this corresponds to “a one-hot vector representing an argmax…of the input tensor”; Yu et al. also defines their “method” using Equation (4):
PNG
media_image12.png
129
499
media_image12.png
Greyscale
which converts the indices into binary values as shown in “our method” in Figure 2, thus this corresponds to “a binary argmax…function”);
Nakahara et al. and Yu et al. are both analogous to the claimed invention because they both utilize FPGAs to run CNNs. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replace the neural network disclosed by Nakahara et al. with a neural network that contains a pooling function that outputs a one-hot vector as taught by Yu et al. Hence, this would be a substitution of one known element (neural network) to another (neural network with an associated pooling function) to obtain predictable results (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Regarding claim 9, Nakahara et al. discloses “A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations,” (Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA… As for the number of operations for the implemented CNN, sine it performed 3 × 3 MAC operations for 256 feature maps at 143 MHz, it was 329.47 GOPS (Giga Operations Per Seconds).”; “FPGA” corresponds to “a hardware accelerator”; “MAC operations” correspond to “elementary neural network operations”)“the method comprising:”
“receiving a definition of a neural network process to be performed, the neural network process comprising a neural network with an associated binary argmax or binary argmin function” (Nakahara et al. Section IV “Fig. 5 shows the architecture for the proposed BN free binarized CNN”; Fig. 5 displays that the CNN contains a “Binarized Max-Pooling (OR gate)” and Section II defines it as
PNG
media_image13.png
92
363
media_image13.png
Greyscale
which corresponds to “binary argmax…function”);
“mapping the binary argmax or binary argmin function to a set of elementary neural network operations, wherein the set of elementary neural network operations comprises only elementary neural network operations from the set of available elementary neural network operations” (Nakahara et al. Fig. 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “conv” stands for “convolutional”; since the feature maps are mapped to “conv+max pooling” involves pooling layers that use convolutional and max pooling and the max pooling layers use a binary max function
PNG
media_image13.png
92
363
media_image13.png
Greyscale
in Section II, this corresponds to “mapping the binary argmax…function to a set of elementary neural network operations” in which a “convolution” corresponds to “a set of elementary neural network operations” ) and
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated binary argmax or binary argmin function, wherein the binary argmax or binary argmin function is performed using the set of elementary neural network operations”(Nakahara et al. Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA”; Fig. 5
PNG
media_image2.png
584
663
media_image2.png
Greyscale
displays “ binarized max-pooling” which includes a binary argmax function is performed on the FPGA which includes an “ARM Cortex A-9 Processor”, thus this corresponds to “ processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated binary argmax… function”; since “max-pooling” performs a max pooling operation, this corresponds to “the binary argmax…function is performed using the set of elementary neural network operations” );
“wherein the data comprises image data and/or audio data”(Nakamura et al. Section III “We designed both VGG11s by using Chainer [1]. Then, we trained the VGG11s using the CIFAR10 training images on the NVidia GTX Titan X.”); and
“wherein each of the set of elementary neural network operations is selected from a list consisting of:”
“an element-wise addition operation” (Nakahara et al. Section III
PNG
media_image3.png
230
686
media_image3.png
Greyscale
; Equation (2) mathematically defines an average pooling layer which sums each element labeled as “xK/2” thus this corresponds to “an element-wise addition operation”)
“an element-wise multiplication operation”(Nakahara et al. Table 1 defines a binarized multiplication at a pooling layer:
PNG
media_image4.png
248
344
media_image4.png
Greyscale
in which “x” corresponds to an “element” in the input data)
“a max pooling operation or min pooling operation” (Nakahara et al. Fig. 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “Max Pooling” corresponds to “a max pooling operation”)[note: since the limitation recites “or”, it can be read as “a max pooling operation”]
“a convolution operation” (Nakahara et al. Fig 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “CONV” corresponds to “a convolution operation”);
Nakahara et al. fails to teach:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise subtraction operation…an element-wise maximum operation…an element-wise minimum operation…a magnitude operation…one or more lookups operations using one or more look-up tables…a deconvolution operation”
However, Yu et al. teaches:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise subtraction operation” (Yu et al. Section II Equation 4 displays bit truncation that is done during pooling:
PNG
media_image5.png
127
475
media_image5.png
Greyscale
; in which x’j represents each element in the input data and
PNG
media_image6.png
36
274
media_image6.png
Greyscale
is a subtraction operation that is performed at each x’j );
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise maximum operation”(Yu et al. Section II Equation (1):
PNG
media_image7.png
120
582
media_image7.png
Greyscale
; in which “xj” corresponds to “an element” in the input data );
“wherein each of the set of elementary neural network operations is selected from a list consisting of… a magnitude operation”(Yu et al. Section II Equation (1):
PNG
media_image7.png
120
582
media_image7.png
Greyscale
; since the magnitude of the maximum value (labeled as “|Max|”) is taken, this corresponds to “a magnitude operation”);
“wherein each of the set of elementary neural network operations is selected from a list consisting of… one or more lookups operations using one or more look-up tables”(Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption. In the Caffe [9] framework widely used on CPU and GPU, indices store the absolute position of the maximum value in the feature map before pooling, which requires multiple bits. Memory consumption is a stiff challenge for embedded platforms with limited storage resources. Therefore, we use a similar approach [10] as shown in Figure 2. Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2ൈ2 pooling window. For SegNetbasic, as shown in Table I, Caffe method of storing pooling indices for the entire network requires 6.72 MB, while our method only requires 0.87 MB, which reduces 87% memory footprint.”; Yu et al. Table I displays:
PNG
media_image8.png
383
595
media_image8.png
Greyscale
; since Table I stores pooling indices which are “used in the last decoder”, Table I corresponds to “a lookup table”); and
“wherein each of the set of elementary neural network operations is selected from a list consisting of… a deconvolution operation”(Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption….”; Yu et al. Abstract “However, CNNs for semantic segmentation usually contain some symmetrical encoders and decoders, corresponding to the down-sampling process (e.g., pooling, convolution) and the up-sampling process (e.g., unpooling, deconvolution)”; since “unpooling” and “deconvolution” are synonymous, “unpooling operation” corresponds to “a deconvolution operation”)
Nakahara et al. and Yu et al. are both analogous to the claimed invention because they both utilize FPGAs to run CNNs. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the set of elementary neural operations mentioned by Nakahara et al. to further include element wise subtraction, maximum operations, and a deconvolution operation as mentioned by Yu et al. Doing so would optimize the storage of pooling indices (Yu et al. Section I)
Both Nakahara et al. and Yu et al. fail to teach:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise minimum operation…”
However, Satti et al. teaches:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise minimum operation…””(Satti et al. Algorithm 1 defines the pooling layer as:
PNG
media_image9.png
383
596
media_image9.png
Greyscale
; line 9 takes the minimum of the elements in the window “W-c” which includes each element “Pi,j” for an image input data, thus this corresponds to “an element wise minimum operation”);
Nakahara et al., Yu et al., and Satti et al. are all analogous to the claimed invention because they are all in the same field of utilizing max pooling layers in CNN. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the operations listed by Nakahara et al. and Yu et al. to further include an element-wise minimum operation as taught by Satti et al. Doing so would remove noise from input images more effectively (Satti et al. Section IV).
Regarding claim 16, the rejections of claim 1 are incorporated. Nakahara et al. further recites:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations” (Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA… As for the number of operations for the implemented CNN, sine it performed 3 × 3 MAC operations for 256 feature maps at 143 MHz, it was 329.47 GOPS (Giga Operations Per Seconds).”; “FPGA” corresponds to “a hardware accelerator”; “MAC operations” correspond to “elementary neural network operations”); and
“a controller configured to perform the method as set forth in claim 1” (Nakahara et al. Fig. 4 displays a “controller”);
Regarding claim 18, the rejections of claim 9 are incorporated. Nakahara et al. further teaches:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations” (Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA… As for the number of operations for the implemented CNN, sine it performed 3 × 3 MAC operations for 256 feature maps at 143 MHz, it was 329.47 GOPS (Giga Operations Per Seconds).”; “FPGA” corresponds to “a hardware accelerator”; “MAC operations” correspond to “elementary neural network operations”); and
“a controller configured to perform the method as set forth in claim 9” (Nakahara et al. Fig. 4 displays a “controller”).
Claims 3 and 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Nakahara et al.’s “A Fully Connected Layer Elimination for a Binarized Convolutional Neural Network on an FPGA” in view of Yu et al.’s “Optimizing FPGA-based Convolutional Encoder-Decoder Architecture for Semantic Segmentation”, in further view of Satti et al.’s “Min-Max Average Pooling Based Filter for Impulse Noise Removal” in further view of Zhang et al. (US20200294184A1) and in further view of user Louis in Stackoverflow’s “If function, if value in vector > 0, then replace with number”.
Regarding claim 3, the rejections of claim 1 are incorporated. Yu et al. further teaches:
“an argmax or argmin operation, applied to the input tensor, that identifies an index of a maximum value or minimum value respectively of the input tensor” (Yu et al. Section II “Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value….”; Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; since the “max pooling input” is used to generate pooling indices, this corresponds to “the input tensor”; “indices…store the relative position of the maximum value” corresponds to “identify an index of a maximum value…of the input tensor”);
“a first intermediate vector obtaining operation that obtains a first intermediate vector having a same number of entries as the input tensor, each entry of the first intermediate vector containing an index value of a different entry of the input tensor” (Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; Figure 2 illustrates that the “pooling indices” is a vector, thus corresponding to “a first intermediate vector”; since the “pooling indices” contain index values for the elements “1”, “3”, “5” and “7”, thus it contains “an index value of a different entry of the input tensor” in which “max pooling input” corresponds to “the input tensor”);
“a subtraction operation, applied to each element of the first intermediate vector, that subtracts the identified index of the input tensor from the said element of the first intermediate vector to produce a second intermediate vector” (Yu et al. Section II “Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value….”; Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; since the “max pooling input” is used to generate pooling indices, this corresponds to “the input tensor”; “indices…store the relative position of the maximum value” corresponds to “the identified index of the input tensor”; since the “pooling indices” is a vector this corresponds to “a first intermediate vector”) ;
Nakahara et al. and Yu et al. are both analogous to the claimed invention because they both utilize FPGAs to run CNNs. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the neural network disclosed by Nakahara et al. with a neural network that contains a pooling function that outputs a one-hot vector as taught by Yu et al. Hence, this would be a substitution of one known element (neural network) to another (neural network with an associated pooling function) to obtain predictable results (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Nakahara et al., in view of Yu et al. and in view of Satti et al. do not teach:
“a first intermediate vector obtaining operation that obtains a first intermediate vector having a same number of entries as the input tensor, each entry of the first intermediate vector containing an index value of a different entry of the input tensor”;
“a subtraction operation, applied to each element of the first intermediate vector, that subtracts the identified index of the input tensor from the said element of the first intermediate vector to produce a second intermediate vector”; and
“a zero-identification operation, applied to the second intermediate vector, that replaces any zero values in the second intermediate vector with a first binary value and any non-zero values in the second intermediate vector with a second, different binary value, to thereby produce the one-hot vector”
However, Zhang et al. teaches:
“a first intermediate vector obtaining operation that obtains a first intermediate vector having a same number of entries as the input tensor, each entry of the first intermediate vector containing an index value of a different entry of the input tensor”( Zhang et al. [0017]-[0018]
PNG
media_image14.png
382
790
media_image14.png
Greyscale
; “tensor T” is an input tensor with size “m x n x k” and “set S” is an intermediate tensor generated from tensor T with size “o x p x q”; since o can be m, p can be n, and q can be k as denoted in
PNG
media_image15.png
30
88
media_image15.png
Greyscale
PNG
media_image16.png
27
184
media_image16.png
Greyscale
, since corresponds to “having a same number of entries” as in input tensor);
“a subtraction operation, applied to each element of the first intermediate vector, that subtracts the identified index of the input tensor from the said element of the first intermediate vector to produce a second intermediate vector” (Zhang et al. [0029]
PNG
media_image17.png
272
777
media_image17.png
Greyscale
; In Equation (5), the observation set S from T denoted as “ObserveS(T-Xt*Yt)” is an intermediate vector from an input tensor and each element in S labeled as “T” in Equation (5) is subtracted with the product of “Xt-- * Yt” which denotes the index of minimum values in the input tensor; therefore “T-Xt*Yt” in Equation (5) corresponds to “a subtraction operation applied to each element” to produce a “second intermediate vector” denoted as Xt in Equation (5)); and
“a zero-identification operation, applied to the second intermediate vector, that replaces any zero values in the second intermediate vector with a first binary value and any non-zero values in the second intermediate vector with a second, different binary value, to thereby produce the one-hot vector” (Zhang et al. [0029]
PNG
media_image17.png
272
777
media_image17.png
Greyscale
; “X-t” corresponds to “second intermediate vector”);
Nakahara et al., Yu et al., Satti et al., and Zhang et al. are all analogous to the claimed invention because they are all in the same field of performing operations on input tensors. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Nakahara et al., Yu et al., and Satti et al. to include a second intermediate vector by using the first intermediate vector disclosed by Nakahara et al., Yu et al., and Satti et al. Doing so would prevent data loss (Zhang et al. [0003]).
Nakahara et al., Yu et al., Satti et al., and Zhang et al. all fail to teach:
“a zero-identification operation, applied to the second intermediate vector, that replaces any zero values in the second intermediate vector with a first binary value and any non-zero values in the second intermediate vector with a second, different binary value, to thereby produce the one-hot vector”
However, user Louis from Stackoverflow teaches:
“a zero-identification operation, applied to the second intermediate vector, that replaces any zero values in the second intermediate vector with a first binary value and any non-zero values in the second intermediate vector with a second, different binary value, to thereby produce the one-hot vector” (Louis
PNG
media_image18.png
411
676
media_image18.png
Greyscale
; the line “myVector[!myVector >0] <- 0” indicates that if an element in “myVector” is less than or equal to 0 then the binary value “0” is assigned, thus this corresponds to “replaces any zero values…with a first binary value”; the line “myVector[myVector >0 ]<1” replaces any element greater than 0 [thus this corresponds to “any non-zero values” with 1 which is a different binary value, thus this corresponds to “replaces…any non-zero values…with a second, different binary value”; “the result” corresponds to a “one-hot vector”)
Nakahara et al., Yu et al., Satti et al. Zhang et al., and Louis are all analogous to the claimed invention because they are all in the same field of modifying vectors and encoding them. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the second intermediate vector disclosed by Zhang et al. with a one-hot vector with different binary values as mentioned by Louis. Thus, this would be a substitution of one known element (second intermediate vector) with another (one-hot vector) to obtain predictable results (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Regarding claim 10, the rejections of claim 9 are incorporated and claim 10 is rejected for the same reasons as claim 3.
Regarding claim 11, the rejections of claim 10 are incorporated. Yu et al. further discloses:
“wherein the first intermediate vector is stored in a look-up table” (Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption. In the Caffe [9] framework widely used on CPU and GPU, indices store the absolute position of the maximum value in the feature map before pooling, which requires multiple bits. Memory consumption is a stiff challenge for embedded platforms with limited storage resources. Therefore, we use a similar approach [10] as shown in Figure 2. Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2ൈ2 pooling window. For SegNetbasic, as shown in Table I, Caffe method of storing pooling indices for the entire network requires 6.72 MB, while our method only requires 0.87 MB, which reduces 87% memory footprint.”; Yu et al. Table I displays:
PNG
media_image8.png
383
595
media_image8.png
Greyscale
; since Table I stores pooling indices which are “used in the last decoder”, Table I corresponds to “a lookup table” and “pooling indices” correspond to “the first intermediate vector”)
Nakahara et al. and Yu et al. are both analogous to the claimed invention because they both utilize FPGAs to run CNNs. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the max pooling output mentioned by Nakahara et al. to further include storing the output in a lookup table as mentioned by Yu et al. Doing so would optimize the storage of pooling indices (Yu et al. Section I)
Regarding claim 13, the rejections of claim 1 are incorporated. Nakahara et al. further mentions:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 1 to be performed when the code is run” (Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA”; “FPGA” and “Zedboard” corresponds to “a non-transitory computer-readable medium…having stored thereon computer code”);
Regarding claim 15, the rejections of claim 9 are incorporated. Nakahara et al. further teaches:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 9 to be performed when the code is run” (Nakahara et al. Section V “We implemented the binarized CNN for the VGG-11 on the Xilinx Inc. Zedboard, which has the Xilinx Zynq FPGA”; “FPGA” and “Zedboard” corresponds to “a non-transitory computer-readable medium…having stored thereon computer code”)
Claims 4, 14, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al.’s “Optimizing FPGA-based Convolutional Encoder-Decoder Architecture for Semantic Segmentation” in view of Nakahara et al.’s “A Fully Connected Layer Elimination for a Binarized Convolutional Neural Network on an FPGA” and in further view of Satti et al.’s “Min-Max Average Pooling Based Filter for Impulse Noise Removal”.
Regarding claim 4, Yu et al. discloses “A method of processing data according to a neural network process using a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations,”(Yu et al. Section II “The unpooling layer in a decoder uses the pooling indices generated by the max pooling layer in its corresponding encoder to complete the up-sampling… The unpooling kernel shown in Figure 3 can be replaced with other kernels for up-sampling such as deconvolution kernel to support different CNNs” ; Section III “The proposed FPGA-based convolutional encoder-decoder architecture accelerator is evaluated on CamVid dataset using 8 bits dynamic fixed-point number.”; “FPGA” corresponds to “hardware accelerator” and “FPGA-based convolutional encoder-decoder architecture” corresponds to “fixed-function circuitry of the hardware accelerator to perform the neural network”; Yu et al. Figure 3 displays that the unpooling kernel is located “on-chip” of the FPGA and since the unpooling kernel can use a deconvolution kernel this corresponds to “perform a set of available elementary neural network operations”) “the method comprising”:
“receiving a definition of a neural network process to be performed, the neural network process comprising a neural network with an associated unpooling or backward pooling function, wherein the unpooling or backward pooling function is configured to map an input value to an original position in a tensor using a one-hot vector that represents an argmax or argmin of a previous pooling function” (Yu et al. Section II “The unpooling kernel shown in Figure 3 can be replaced with other kernels for up-sampling such as deconvolution kernel to support different CNNs… For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption. In the Caffe [9] framework widely used on CPU and GPU, indices store the absolute position of the maximum value in the feature map before pooling, which requires multiple bits. Memory consumption is a stiff challenge for embedded platforms with limited storage resources. Therefore, we use a similar approach [10] as shown in Figure 2. Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2x2 pooling window”; Yu et al. Fig 2
PNG
media_image19.png
343
715
media_image19.png
Greyscale
; since the unpooling kernel “support different CNNs”, this means that “a definition of a neural network process” is received; “The unpooling kernel…to support different CNNs” also correspond to “a neural network comprising an associated unpooling or backward pooling function”; since the pooling indices represent indices of the “max pooling output”, “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder” corresponds to “…using a one-hot vector that represents an argmax or argmin of a previous pooling function” in which “pooling indices” correspond to “a one-hot vector”; since the unpooling layer uses the “unpooling input” to map to an “unpooling output” this corresponds to “configured to map an input value to an original position in a tensor”);
“mapping the unpooling or backward pooling function to a set of elementary neural network operations, wherein the set of elementary neural network operations comprises only elementary neural network operations from the set of available elementary neural network operations” (Yu et al. Section II “The unpooling kernel shown in Figure 3 can be replaced with other kernels for up-sampling such as deconvolution kernel to support different CNNs.”; “The unpooling kernel…can be replaced with…deconvolution kernel” corresponds to “mapping the unpooling or backward pooling function to a set of elementary neural network operations” in which “deconvolution kernel” corresponds to a elementary neural network operation “from the set of available elementary neural network operations”) ; and
“processing the data according to the neural network process, using the fixed-function circuitry of the hardware accelerator to perform the neural network with the associated unpooling or backward pooling function, wherein the unpooling or backward pooling function is performed using the set of elementary neural network operations” (Yu et al. Section II “The unpooling layer in a decoder uses the pooling indices generated by the max pooling layer in its corresponding encoder to complete the up-sampling… The unpooling kernel shown in Figure 3 can be replaced with other kernels for up-sampling such as deconvolution kernel to support different CNNs” ; Section III “The proposed FPGA-based convolutional encoder-decoder architecture accelerator is evaluated on CamVid dataset using 8 bits dynamic fixed-point number.”; “FPGA” corresponds to “hardware accelerator” and “FPGA-based convolutional encoder-decoder architecture” corresponds to “fixed-function circuitry of the hardware accelerator to perform the neural network”; Yu et al. Figure 3 displays that the unpooling kernel is located “on-chip” of the FPGA and since the unpooling kernel can use a deconvolution kernel this corresponds to “perform the neural network with the associated unpooling or backward pooling function, wherein the unpooling or backward pooling function is performed using the set of elementary neural network operations”);
“wherein the data comprises image data and/or audio data” (Yu et al. “The proposed FPGA-based convolutional encoder-decoder architecture accelerator is evaluated on CamVid dataset using 8 bits dynamic fixed-point number. The performance of the implementation for SegNet-basic to process an image is 2891 ms”); and
“wherein each of the set of elementary neural network operations is selected from a list consisting of:”
“an element-wise subtraction operation” (Yu et al. Section II Equation 4 displays bit truncation that is done during pooling:
PNG
media_image5.png
127
475
media_image5.png
Greyscale
; in which x’j represents each element in the input data and
PNG
media_image6.png
36
274
media_image6.png
Greyscale
is a subtraction operation that is performed at each x’j );
“an element-wise maximum operation” (Yu et al. Section II Equation (1):
PNG
media_image7.png
120
582
media_image7.png
Greyscale
; in which “xj” corresponds to “an element” in the input data );
“a magnitude operation” (Yu et al. Section II Equation (1):
PNG
media_image7.png
120
582
media_image7.png
Greyscale
; since the magnitude of the maximum value (labeled as “|Max|”) is taken, this corresponds to “a magnitude operation”)
“one or more lookups operations using one or more look-up tables” (Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption. In the Caffe [9] framework widely used on CPU and GPU, indices store the absolute position of the maximum value in the feature map before pooling, which requires multiple bits. Memory consumption is a stiff challenge for embedded platforms with limited storage resources. Therefore, we use a similar approach [10] as shown in Figure 2. Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2ൈ2 pooling window. For SegNetbasic, as shown in Table I, Caffe method of storing pooling indices for the entire network requires 6.72 MB, while our method only requires 0.87 MB, which reduces 87% memory footprint.”; Yu et al. Table I displays:
PNG
media_image8.png
383
595
media_image8.png
Greyscale
; since Table I stores pooling indices which are “used in the last decoder”, Table I corresponds to “a lookup table”); and
“a deconvolution operation”(Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption….”; Yu et al. Abstract “However, CNNs for semantic segmentation usually contain some symmetrical encoders and decoders, corresponding to the down-sampling process (e.g., pooling, convolution) and the up-sampling process (e.g., unpooling, deconvolution)”; since “unpooling” and “deconvolution” are synonymous, “unpooling operation” corresponds to “a deconvolution operation”)
Yu et al. fails to disclose:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise addition operation… an element-wise multiplication operation…an element-wise minimum operation…a max pooling operation or min pooling operation… a convolution operation….”
However, Nakahara et al. discloses:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise addition operation….”(Nakahara et al. Section III
PNG
media_image3.png
230
686
media_image3.png
Greyscale
; Equation (2) mathematically defines an average pooling layer which sums each element labeled as “xK/2” thus this corresponds to “an element-wise addition operation”);
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise multiplication operation….” (Nakahara et al. Table 1 defines a binarized multiplication at a pooling layer:
PNG
media_image4.png
248
344
media_image4.png
Greyscale
in which “x” corresponds to an “element” in the input data);
“wherein each of the set of elementary neural network operations is selected from a list consisting of… a max pooling operation or min pooling operation” (Nakahara et al. Fig. 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “Max Pooling” corresponds to “a max pooling operation”)[note: since the limitation recites “or” this can be read as “wherein each of the set of elementary neural network operations is selected from a list consisting of…a max pooling operation”]; and
“wherein each of the set of elementary neural network operations is selected from a list consisting of…a convolution operation” (Nakahara et al. Fig 1
PNG
media_image1.png
233
555
media_image1.png
Greyscale
; “CONV” corresponds to “a convolution operation”);
Yu et al. and Nakahara et al. are both analogous to the claimed invention because they are in the same field of implementing CNNs through pooling. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have substituted the pooling function disclosed by Yu et al. with a max pooling function disclosed by Nakahara et al. in order to obtain pooled output. Thus, this would be a simple substitution of one known element (pooling function) with another (max pooling function) to obtain predictable results (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Yu et al. and Nakahara et al. both fail to teach:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise minimum operation…”
However, Satti et al. teaches:
“wherein each of the set of elementary neural network operations is selected from a list consisting of…an element-wise minimum operation…””(Satti et al. Algorithm 1 defines the pooling layer as:
PNG
media_image9.png
383
596
media_image9.png
Greyscale
; line 9 takes the minimum of the elements in the window “W-c” which includes each element “Pi,j” for an image input data, thus this corresponds to “an element wise minimum operation”);
Yu et al., Nakahara et al., and Satti et al. are all analogous to the claimed invention because they are all in the same field of utilizing max pooling layers in CNN. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the operations listed by Nakahara et al. and Yu et al. to further include an element-wise minimum operation as taught by Satti et al. Doing so would remove noise from input images more effectively (Satti et al. Section IV).
Regarding claim 14, the rejections of claim 4 are incorporated. Yu et al. further teaches:
“A non-transitory computer-readable medium or data carrier having stored thereon computer readable code configured to cause the method of claim 4 to be performed when the code is run” (Yu et al. Section III “We use the OpenCL to implement the SegNet-basic algorithm with the help of the Intel FPGA SDK. The host side controls data transmission and devices execution, consisting of Intel i7-8700 CPU and 64GB memories. The device side is the Arria-10 GX1150 FPGA, and the kernel is run on it to achieve accelerating.”; “Intel i7-8700 CPU” and “Intel FPGA SDK” corresponds to “A non-transitory computer readable medium…having stored thereon computer readable code”; “algorithm” corresponds to “computer readable code”);
Regarding claim 17, the rejections of claim 4 are incorporated. Yu et al. further mentions:
“a hardware accelerator comprising fixed-function circuitry configured to perform a set of available elementary neural network operations” ,”( Yu et al. Section II “The unpooling layer in a decoder uses the pooling indices generated by the max pooling layer in its corresponding encoder to complete the up-sampling… The unpooling kernel shown in Figure 3 can be replaced with other kernels for up-sampling such as deconvolution kernel to support different CNNs” ; Section III “The proposed FPGA-based convolutional encoder-decoder architecture accelerator is evaluated on CamVid dataset using 8 bits dynamic fixed-point number.”; “FPGA” corresponds to “hardware accelerator” and “FPGA-based convolutional encoder-decoder architecture” corresponds to “fixed-function circuitry of the hardware accelerator to perform the neural network”; Yu et al. Figure 3 displays that the unpooling kernel is located “on-chip” of the FPGA and since the unpooling kernel can use a deconvolution kernel this corresponds to “perform a set of available elementary neural network operations”); and
“a controller configured to perform the method as set forth in claim 4” (Yu et al. Section III “We use the OpenCL to implement the SegNet-basic algorithm with the help of the Intel FPGA SDK. The host side controls data transmission and devices execution, consisting of Intel i7-8700 CPU and 64GB memories. The device side is the Arria-10 GX1150 FPGA, and the kernel is run on it to achieve accelerating.”; “Intel i7-8700 CPU” corresponds to “a controller”).
Regarding claim 19, the rejections of claim 14 are incorporated. Nakahara et al. teaches
“wherein the hardware accelerator comprises any one of, or any combination of two or more of:”
“an activation unit, comprising an LUT” (Nakahara et al. Section IV “We used the Xilinx Inc. SDSoC 2016.4 to generate the bistream with timing constrain 143.78 MHz. Our implementation used 18,325 LUTs, 19,913 FFs, 32 18Kb BRAMs, and a DSP48E.”);
a local response normalisation unit, configured to perform a local response normalisation;
an element-wise operations unit, configured to apply a selected operation to every pair of respective elements of two tensor of identical size;
one or more convolution engines, configured to perform convolution operations; and
“a pooling unit, configured to perform pooling operations, including max pooling and/or min pooling” (Nakahara et al. Fig. 5
PNG
media_image20.png
479
572
media_image20.png
Greyscale
; “binarized max-pooling (OR gate)” corresponds to “a pooling unit” that performs “max pooling”); [ note since the limitation recites “two or more”, it can be read as “the hardware accelerator comprises… an activation unit, comprising an LUT…and a pooling unit, configured to perform pooling operations, including max pooling and/or min pooling”]
Yu et al. and Nakahara et al. are both analogous to the claimed invention because they are both in the same field of executing CNNs on FPGAs. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the hardware accelerator disclosed by Yu et al. with a hardware accelerator with a LUT and max pooling unit as recited by Nakahara et al. Thus, this would be a simple substitution of one known element (hardware accelerator) with another (hardware accelerator with a LUT and max pooling unit) to obtain predictable results (CNN execution) (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Claims 5-8 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al.’s “Optimizing FPGA-based Convolutional Encoder-Decoder Architecture for Semantic Segmentation” in view of Nakahara et al.’s “A Fully Connected Layer Elimination for a Binarized Convolutional Neural Network on an FPGA”, in further view of Satti et al.’s “Min-Max Average Pooling Based Filter for Impulse Noise Removal” in further view of Reisser et al. (US20210073650A1) and in further view of Zeiler et al.’s “Adaptive Deconvolutional Networks for Mid and High Level Feature Learning”.
Regarding claim 5, the rejections of claim 4 are incorporated. Yu et al. further recites:
“a binary argmax/argmin acquisition function, configured to obtain the one-hot vector representing an argmax or argmin of a previous pooling function” (Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; Figure 2 displays that in addition to finding the maximum values of the “max pooling input”, the max pooling layer also outputs “pooling indices” which represent the indices of each entry in the max pooling output, which is then passed as an “unpooling input” thus “unpooling” includes obtaining “the one-hot vector representing an argmax or argmin of a previous pooling function” in which the “pooling indices” correspond to “the one-hot vector”);
“a multiplication function configured to multiply each entry in the one-hot vector by the input value to produce a product one-hot vector” (Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; Figure 2 displays “pooling indices” which correspond to “the one-hot vector”);
Yu et al. in view of Nakahara et al. and in view of Satti et al. fail to recite:
“a multiplication function configured to multiply each entry in the one-hot vector by the input value to produce a product one-hot vector”; and
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor”.
However, Reisser et al. recites:
“a multiplication function configured to multiply each entry in the one-hot vector by the input value to produce a product one-hot vector” (Reisser et al [0029] “ The CIM array 306 comprises c columns and r rows of CIM cells 314(1)(1)-314(c)(r), wherein each CIM cell 314(i)(j) is configured to store a corresponding weight value and multiply it with a received input value. Note that, as noted elsewhere herein, the multiplication of binary weight and input values may be performed using an XNOR operation”; “XNOR operation” corresponds to “a multiplication function”; since the input value is multiplied with a CIM array which includes binary weights and is a type of one-hot vector, this means that the “multiplication of binary weight and input value” results in “a product one-hot vector”); and
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor” (Reisser et al [0029] “ The CIM array 306 comprises c columns and r rows of CIM cells 314(1)(1)-314(c)(r), wherein each CIM cell 314(i)(j) is configured to store a corresponding weight value and multiply it with a received input value. Note that, as noted elsewhere herein, the multiplication of binary weight and input values may be performed using an XNOR operation”; since the input value is multiplied with a CIM array which includes binary weights and is a type of one-hot vector, this means that the “multiplication of binary weight and input value” results in “a product one-hot vector”)
Yu et al., Nakahara et al., Satti et al, and Reisser et al. are all analogous to the claimed invention because they are all in the same field of binarizing pooling indices. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have utilized the pooling indices disclosed by Yu et al. and Nakahara et al. for a multiplication function to generate a product one-hot vector as disclosed by Reisser et al. Doing so would allow for power and time savings (Reisser et al. [0017]).
Yu et al., Nakahara et al., Satti et al., and Reisser et al. fail to teach:
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor”
However, Zeiler et al. teaches:
“a deconvolution function configured to process the product one-hot vector, using a binary constant filter, to generate an output tensor”(Zeiler et al. “Unpooling: In the convnet, the max pooling operation is non-invertible, however we can obtain an approximate inverse by recording the locations of the maxima within each pooling region in a set of switch variables. In the deconvnet, the unpooling operation uses these switches to place the reconstructions from the layer above into appropriate locations, preserving the structure of the stimulus”; Fig. 1
PNG
media_image21.png
567
672
media_image21.png
Greyscale
since “unpooling” is performed during “deconvnet” which is deconvolution, using “a set of switch variable” and Fig. 1 displays that the “switches” as a matrix, this corresponds to “a deconvolution function configured to process…using a binary constant filter” in which “a set of switches” correspond to “a binary constant filter”; Fig. 1 illustrates that the deconvolution results in “unpooled maps” which corresponds to “an output tensor”);
Yu et al., Nakahara et al., Satti et al., Reisser et al., and Zeiler et al. are all analogous to the claimed invention because they are all in the same field of implementing pooling and unpooling in CNNs. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the deconvolution function taught by Yu et al. with a deconvolution function that uses a binary filter as taught by Zeiler et al. in order to reconstruct input data. Hence, this would be a simple substitution of one known element (deconvolution function) with another (deconvolution function with binary filters) to obtain predictable results (reconstructed input data) (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Regarding claim 6, the rejections of claim 5 are incorporated. Nakahara et al. further mentions:
“the multiplication function is performed using an element-wise multiplication operation” (Nakahara et al. Table 1 defines a binarized multiplication at a pooling layer:
PNG
media_image4.png
248
344
media_image4.png
Greyscale
in which “x” corresponds to an “element” in the input data);
Yu et al. and Nakahara et al. are both analogous to the claimed invention because they are in the same field of implementing CNNs through pooling. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have substituted the pooling function disclosed by Yu et al. with a max pooling function disclosed by Nakahara et al. in order to obtain pooled output. Thus, this would be a simple substitution of one known element (pooling function) with another (max pooling function) to obtain predictable results (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Regarding claim 7, the rejections of claim 5 are incorporated. Yu et al. further teaches:
“wherein the deconvolution function is performed using a deconvolution operation” (Yu et al. Section II “For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption….”; Yu et al. Abstract “However, CNNs for semantic segmentation usually contain some symmetrical encoders and decoders, corresponding to the down-sampling process (e.g., pooling, convolution) and the up-sampling process (e.g., unpooling, deconvolution)”; since “unpooling” and “deconvolution” are synonymous, “unpooling operation” corresponds to “a deconvolution operation”)
Regarding claim 8, the rejections of claim 5 are incorporated. Zeiler et al. further mentions:
“wherein the deconvolution function is a grouped deconvolution function”(Zeiler et al. Section 2
PNG
media_image22.png
286
593
media_image22.png
Greyscale
; since the deconvolution is performed on “each of the 2D feature maps…with filters…”, this corresponds to “a grouped deconvolution”);
Yu et al., Nakahara et al., Satti et al., Reisser et al., and Zeiler et al. are all analogous to the claimed invention because they are all in the same field of implementing pooling and unpooling in CNNs. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the deconvolution function taught by Yu et al. with a grouped deconvolution function that uses a binary filter as taught by Zeiler et al. in order to reconstruct input data. Hence, this would be a simple substitution of one known element (deconvolution function) with another (grouped deconvolution function with binary filters) to obtain predictable results (reconstructed input data) (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Nakahara et al.’s “A Fully Connected Layer Elimination for a Binarized Convolutional Neural Network on an FPGA” in view of Yu et al.’s “Optimizing FPGA-based Convolutional Encoder-Decoder Architecture for Semantic Segmentation”, in further view of Satti et al.’s “Min-Max Average Pooling Based Filter for Impulse Noise Removal” in further view of Zhang et al. (US20200294184A1), in further view of user Louis in Stackoverflow’s “If function, if value in vector > 0, then replace with number” and in further view of Stackoverflow’s “How can I find the index of the highest value in a vector, defaulting to the greater index if there are two "greatest" indices?”.
Regarding claim 12, the rejections of claim 10 are incorporated. Yu et al. further discloses:
“a binary maximum/minimum operation, applied to the first input tensor, to produce a binary tensor of the same spatial size as the first input tensor,”( Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; Figure 2 displays that in addition to finding the maximum values of the “max pooling input”, the max pooling layer also outputs “pooling indices” which represent the indices of each entry in the max pooling output, and “our method” represent the indices as binary values, thus this corresponds to “a binary tensor”; Yu et al. also defines their “method” using Equation (4):
PNG
media_image12.png
129
499
media_image12.png
Greyscale
which converts the indices into binary values as shown in “our method” in Figure 2, thus this corresponds to “a binary maximum/minimum operation”) “wherein each element in the binary tensor:”
“corresponds to a different element of the first input tensor”( Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; Figure 2 illustrates the “pooling indices” which correspond to “a binary vector”; since the “pooling indices” contain index values for the elements “1”, “3”, “5” and “7”, thus it “corresponds to a different element of the first input tensor” in which “max pooling input” corresponds to “the first input tensor”);
“contains a binary value indicating whether or not the corresponding element of the first input tensor has a value equal to the maximum/minimum value contained in the first input tensor” Yu et al. Figure 2
PNG
media_image10.png
353
604
media_image10.png
Greyscale
; Yu et al. Section II “Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2x2 pooling window”; Figure 2 illustrates the “pooling indices” which correspond to “a binary vector”; since the pooling indices are “encoded with 2 bits” and each index “stores the relative position of the maximum value”, the indices themselves indicate that the corresponding element it points to “has a value equal to the maximum/minimum value contained in the first input tensor” in which each “2x2 pooling window” and “max pooling input” correspond to “first input tensor”) ; and
“an integer index operation, applied to the binary tensor, that identifies one or more indexes of the binary tensor, the identified one or more indexes being indexes of the one or more elements of the binary tensor that have a binary value that indicates the corresponding element of the first input tensor has a value equal to the maximum/minimum value contained in the first input tensor” (Yu et al. Section II “The unpooling kernel shown in Figure 3 can be replaced with other kernels for up-sampling such as deconvolution kernel to support different CNNs… For unpooling operation, the pooling indices generated by the first encoder will be used in the last decoder, which will result in massive memory consumption. In the Caffe [9] framework widely used on CPU and GPU, indices store the absolute position of the maximum value in the feature map before pooling, which requires multiple bits. Memory consumption is a stiff challenge for embedded platforms with limited storage resources. Therefore, we use a similar approach [10] as shown in Figure 2. Since the pooling windows of SegNet-basic are all with size 2x2, indices can be encoded with 2 bits, which stores the relative position of the maximum value in a 2x2 pooling window”; Yu et al. Fig 2
PNG
media_image19.png
343
715
media_image19.png
Greyscale
; since the pooling “indices…store[s] the relative position of the maximum value” for each “2x2 pooling window” and is “encoded with 2 bits” , this means that the pooling indices are “indexes of the one or more elements of the binary tensor that have a binary value that indicates the corresponding element of the first input tensor has a value equal to the maximum/minimum value contained in the first input tensor” the “max pooling input” corresponds to “first input tensor”; since the “unpooling operation” uses the pooling indices for the decoder, this corresponds to “an integer index operation, applied to the binary tensor”)
Yu et al. and Nakahara et al. are both analogous to the claimed invention because they are in the same field of implementing CNNs through pooling. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have substituted the pooling output disclosed by Nakahara et al. with a binary tensor disclosed by Yu et al. in order to obtain pooled output. Thus, this would be a simple substitution of one known element (pooling output) with another (binary tensor) to obtain predictable results (MPEP 2143 I. (B) Simple substitution of one known element for another to obtain predictable results).
Nakahara et al., Yu et al., Satti et al., Zhang et al., and Louis all fail to recite:
“a tie elimination operation, applied to the identified indexes, that selects a single one of the one or more identified indexes to provide the output of the argmax or argmin function”
However, Stackoverflow recites:
“a tie elimination operation, applied to the identified indexes, that selects a single one of the one or more identified indexes to provide the output of the argmax or argmin function” (Stackoverflow
PNG
media_image23.png
891
778
media_image23.png
Greyscale
)
Nakahara et al., Yu et al., Satti et al., Zhang et al., Louis, and Stackoverflow are all analogous to the claimed invention because they are all in the same field of performing vector operations. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the operations of Nakahara et al. and Yu et al. to further include a tie elimination operation as taught in Stackoverflow. Thus, this would be combining prior art elements (neural network elementary operations) according to known methods to yield predictable results (modified vector) (MPEP 2141(III) (A) Combining prior art elements according to known methods to yield predictable results).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Lau et al. (US 2018/0189238 A1) describes a max pooling matrix processing.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TEWODROS E MENGISTU whose telephone number is (571)270-7714. The examiner can normally be reached Mon-Fri 9:30-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ABDULLAH KAWSAR can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TEWODROS E MENGISTU/Examiner, Art Unit 2127