DETAILED ACTION
Claims 1-13 and 15-20 are pending.
Claims 1-11, 14-16, and 19-20 have been examined.
Claims 12-13 and 17-18 are withdrawn.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on March 3, 2025, has been entered.
Specification
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
The disclosure submitted on September 26, 2024, is objected to because of the following informalities:
In paragraph 72, line 7, replace “instruction” with --instructions--.
In paragraph 72, line 8, replace “writing” with --write--.
In paragraph 72, line 9, replace “filling” with --fill--.
In paragraph 72, 3rd to last line, replace “moving” with --move--.
Appropriate correction is required.
Claim Objections/Recommendations
In claim 1, on page 2, 7th to last line, the examiner recommends deleting “and” since there are still two more steps.
In claim 1, 2nd to last line, the examiner recommends inserting --the-- before “off-chip”.
Claim 12, which will be rejoined when claim 1 if allowable (from MPEP 821.04, “The propriety of a restriction requirement should be reconsidered when all the claims directed to the elected invention are in condition for allowance, and the nonelected invention(s) should be considered for rejoinder.”), is objected to because line 3 is grammatically incorrect and must be reworded. The examiner recommends replacing “to” with --with--.
In claim 12, is the first phase in line 3 the same as that in claim 1? If so, the examiner recommends claiming it as “the first phase”.
Claim 13, which will be rejoined when claim 1 if allowable, is objected to because line 3 is grammatically incorrect and must be reworded. The examiner recommends replacing “to” with --with--.
In claim 15, 2nd to last paragraph, the examiner recommends inserting --the-- before “off-chip”.
Appropriate correction is required.
Claim Interpretation
The following is a quotation of MPEP 2111.04(II):
“The broadest reasonable interpretation of a method (or process) claim having contingent limitations requires only those steps that must be performed and does not include steps that are not required to be performed because the condition(s) precedent are not met.”
“The broadest reasonable interpretation of a system (or apparatus or product) claim having structure that performs a function, which only needs to occur if a condition precedent is met, requires structure for performing the function should the condition occur. The system claim interpretation differs from a method claim interpretation because the claimed structure must be present in the system regardless of whether the condition is met and the function is actually performed.”
Since the examiner’s indication of allowable subject matter, there has been an increased emphasis within the Office on recognizing and properly interpreting contingent limitations. Such has an effect on allowability in this application.
Regarding claim 1 (and similarly claim 15), in each paragraph beginning with “upon completion…”, the associated “providing…” steps are not required to be performed when the completion does not occur (e.g. due to error, system failure, power outage, etc.). The examiner recommends preceding each of these paragraphs with a positively-reciting step of completing the respective phase. For instance, prior to the first “upon completion…” paragraph, applicant may insert a new paragraph worded as --completing the first phase;--. This then forces the subsequent “providing…” to occur. A similar --completing the second phase;-- paragraph may also be provided.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-13 and 15-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The claims recite the following limitations for which there is a lack of antecedent basis:
In claim 1, 2nd to last line, “the processed data blocks” since there are processed data blocks in the two previous “providing…” paragraphs and in the 3rd to last line. If applicant is referring to the remaining data blocks that are processed, applicant could insert --remaining-- after “processed” in the 2nd to last line.
In claim 1, last line, “the data blocks within the segment of the data tensor”, because there are such blocks in line 3, and also in lines 1-2 of the last paragraph. If applicant is referring to those in line 3, applicant could insert --identified-- before “data blocks” in the last line.
In claim 7, last line, “the segment”. There is a segment in claim 1 and another segment in claim 7, line 3. Please clarify which applicant is referring to.
In claim 12, last paragraph, “the independent segments”.
In claim 13, last paragraph, “the independent segments”.
In claim 15, both instances of “the fused phase of the data tensor”. Applicant never initially established that this phase was of the data tensor.
In claim 15, 5th to last line, “the processed data blocks” since there are processed data blocks in two the previous “providing…” paragraphs and in the 6th to last line. If applicant is referring to the remaining data blocks that are processed, applicant could insert --remaining-- after “processed” in the 5th to last line.
In claim 15, 4th to last line, “the data blocks within the segment of the data tensor”, because there are such blocks in line 5, and also in lines 1-2 of the 2nd to last paragraph. If applicant is referring to those in line 5, applicant could insert
--identified-- before “data blocks” in the 4th to last line.
All dependent claims are rejected due to their dependence on an indefinite claim.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-11, 15-16, and 19-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Jiao et al., 2021/0089611 (as cited by applicant).
Referring to claim 1, Jiao has taught a method, comprising:
identifying a multi-phase algorithm to perform on a data tensor (see FIG.7, 701, which is a convolutional neural network algorithm with multiple phases that operate on a tensor (paragraph [0106]));
identifying data blocks within a segment of the data tensor (this is the nature of convolution for instance, where a filter is convolved on different blocks of an image input);
determining a size of an on-chip memory (from paragraphs[0042] and [0061], on-chip memory, e.g. FIG.4, 4022, is loaded with data. As the memory exists physically, its size (i.e., N X-bit blocks) has been previously determined (e.g. during design/fabrication)), wherein the size is equal to a number of the data blocks that fit in the on-chip memory (the size is finite and thus equal to an amount of data that can fit therein);
providing instructions to fill the on-chip memory with the number of the data blocks for the size (eventually all data is to be processed by the neural network and, thus, over time, it is loaded into on-chip memory via instructions to control DMA);
providing, to a processor with multiple hardware execution units, instructions to simultaneously process different phases of the multi-phase algorithm using the multiple hardware execution units of the processor (see FIG.7, 703e, and paragraphs [0072] and [0078], among others, which describe overlapping (simultaneous) processing of instructions to carry out multiple phases. For instance, while one execution unit is performing a 7x7 convolution phase, another execution unit is performing a 3x3 pooling phase); and
providing instructions to process a first phase of the multi-phase algorithm on a first portion of the data blocks (again, from the portions cited above, instructions would be provided to perform a first phase);
(this limitation is contingent and not required to be taught by the prior art. If the first phase is not completed (e.g. perhaps a power outage or system failure occurs that prevents the completion, this “providing…” is never performed);
providing instructions to process a second phase of the multi-phase algorithm on a second portion of the data blocks (again, from the portions cited above, instructions would be provided to perform a second phase); and
(this limitation is contingent and not required to be taught by the prior art. If the second phase is not completed (e.g. perhaps a power outage or system failure occurs that prevents the completion, this “providing…” is never performed);
providing instructions to continue to process any remaining data blocks within the segment of the data tensor using the on-chip memory and move the processed data blocks to off-chip memory until the data blocks within the segment of the data tensor are processed (see paragraphs [0061]-[0064]. After data blocks are processed, output is moved to off-chip memory).
Referring to claim 2, Jiao has taught the method of claim 1, wherein each execution unit of the multiple hardware execution units handles a subset of an instruction set of the processor (the different phases are implemented with different subsets of instructions, e.g. performing convolution requires different instructions than performing pooling. Thus, each execution unit handles its own subset).
Referring to claim 3, Jiao has taught the method of claim 1, further comprising: continuing to provide instructions to process the different phases of the multi-phase algorithm on independent segments of the data tensor until each phase of the multi-phase algorithm is processed in order on each segment of the data tensor (this is how convolutional neural networks work, with convolution including repeated operations on different segments of the tensor, and pooling including the same. Thus, instructions will be provided until the phases are performed in order on each tensor segment).
Referring to claim 4, Jiao has taught the method of claim 1, wherein the multiple hardware execution units include heterogeneous execution units where each execution unit performs an operation (see FIG.9, and note there is a convolution unit and also a separate pooling unit. One is performing the 7x7 CONV phase while another is performing a 3x3 POOL stage).
Referring to claim 5, Jiao has taught the method of claim 1, wherein a subset of the execution units are used to perform operations for a phase of the multi-phase algorithm, wherein the subset of the execution units includes two or more execution units (a convolution unit would include at least two sub-execution units, e.g. a multiplier and an adder to carry out the multiplication and addition operations required as a part of convolution).
Referring to claim 6, Jiao has taught the method of claim 1, wherein a same subset of execution units are used to perform operations for multiple phases of the multi-phase algorithm (from FIGl.9, there is a convolution (CONV) execution unit. This is a subset of all execution unit, and is used to perform multiple convolution phases of FIG.7).
Referring to claim 7, Jiao has taught the method of claim 1, wherein the multi-phase algorithm includes a data dependency between phases of the multi-phase algorithm (from FIG.7, the dashed lines indicate a dependency of a phase on a previous phase) that requires processing of data within an entire segment of the data tensor for a phase of the multi-phase algorithm before moving to a next phase of the multi-phase algorithm for the segment of the data tensor (From FIG.7, phase 701-3 depends on phase 701-2, but can start because the data on which it depends is calculated earlier in phase 701-3, thereby satisfying the dependency. Alternatively, or in addition, since the Concat phase 701-7 depends on parent phases 701-4 and 701-5, the Concat phase must wait until both of its parent phases finish processing the tensor (as seen in 703e). In general, not that, due to dependency, a dependent phase cannot start until after tensor data is processed by the parent phase).
Referring to claim 8, Jiao has taught the method of claim 1, wherein a data dependency exists among phases of the multi-phase algorithm that requires performance of one operation being dependent on performance of another operation (this is clear per the convolution neural network algorithm. For instance, from FIG.7, data must be convolved before it can be pooled. Thus, as shown in 703e, pooling cannot start until some data has been convolved).
Referring to claim 9, Jiao has taught the method of claim 1, further comprising:
generating a fused phase by combining a plurality of phases of the multi-phase algorithm together (again, see FIG.7, 703e, where the 7x7 CONV and 3x3 POOL phases are fused); and
providing instructions to concurrently process the fused phase on independent segments of the data tensor using the multiple hardware execution units until the fused phase is processed on each segment of the data tensor (as shown by the overlap of phases in 703e, multiple execution units are controlled by instructions to carry out the two phases of the fused phase concurrently until the data tensor is entirely processed).
Referring to claim 10, Jiao has taught the method of claim 1, further comprising:
generating a fused phase by combining all of phases of the multi-phase algorithm together (from FIG.7, phases 701-1 and 701-2 may be considered the multi-phase algorithm for initially convolving and pooling); and
providing instructions to concurrently process the fused phase on the independent segments of the data tensor using the multiple hardware execution units until the fused phase is processed on each segment of the data tensor (as shown by the overlap of these two phases in 703e, multiple execution units are controlled by instructions to carry out the two phases of the fused phase concurrently until the data tensor is entirely processed).
Referring to claim 11, Jiao has taught the method of claim 1, further comprising:
generating column blocks of the data tensor by combining a plurality of segments of the data tensor together (from FIG.6B, the data tensor includes an image, which is at least a 2D tensor. As is known, performing convolution on an image generates column blocks by combining elements of the tensor together (via dot product)); and
providing instructions to concurrently process different phases of the multi-phase algorithm on independent column blocks of the data tensor (the tensor is modified as it goes through the timeline of 703e, where instructions are provided to concurrently process different phases on blocks of the tensor).
Referring to claim 15, Jiao has taught a method, comprising:
identifying a multi-phase algorithm to perform on a data tensor (see FIG.7, 701, which is a convolutional neural network algorithm with multiple phases that operate on a tensor (paragraph [0106]));
creating a fused phase by combining a plurality of phases of the multi-phase algorithm together (see FIG.7, 703e, which fuses the 7x7 CONV and 3x3 POOL phases into a phase that may be processed concurrently. Without fusion, they are processed separately, as shown in any of 703a-d);
identifying data blocks within a segment of the data tensor (this is the nature of convolution for instance, where a filter is convolved on different blocks of an image input);
determining a size of an on-chip memory (from paragraphs[0042] and [0061], on-chip memory, e.g. FIG.4, 4022, is loaded with data. As the memory exists physically, its size (i.e., N X-bit blocks) has been previously determined (e.g. during design/fabrication)), wherein the size is equal to a number of the data blocks that fit in the on-chip memory (the size is finite and thus equal to an amount of data that can fit therein);
providing instructions to fill the on-chip memory with the number of the data blocks for the size (eventually all data is to be processed by the neural network and, thus, over time, it is loaded into on-chip memory via instructions to control DMA);
providing, to a processor with multiple hardware execution units, instructions to simultaneously process the fused phase of the data tensor using the multiple hardware execution units of the processor (again, from 703e, multiple execution units will be instructed to concurrently carry out the aforementioned two phases on independent segments of the tensor. The convolution phase operates on remaining portions of the tensor, while the pooling phase pools portions of the tensor already processed in the convolution phase);
providing instructions to process a first phase of the multi-phase algorithm on a first portion of the data blocks (again, from the portions cited above, instructions would be provided to perform a first phase);
(this limitation is contingent and not required to be taught by the prior art. If the first phase is not completed (e.g. perhaps a power outage or system failure prevents the completion), this “providing…” is never performed);
providing instructions to process a second phase of the multi-phase algorithm on a second portion of the data blocks (again, from the portions cited above, instructions would be provided to perform a second phase); and
(this limitation is contingent and not required to be taught by the prior art. If the second phase is not completed (e.g. perhaps a power outage or system failure prevents the completion), this “providing…” is never performed);
providing instructions to continue to process any remaining data blocks within the segment of the data tensor using the on-chip memory and move the processed data blocks to off-chip memory until the data blocks within the segment of the data tensor are processed (see paragraphs [0061]-[0064]. After data blocks are processed, output is moved to off-chip memory).
continuing to provide, to the processor, instructions to process the fused phase of the data tensor until each phase of the multi-phase algorithm is processed in order on each segment of the data tensor (from 703-e, the fused phase is continued until the tensor is fully convolved and pooled by 701-2 and 701-3).
Referring to claim 16, Jiao has taught the method of claim 15, wherein the fused phase includes each phase of the multi-phase algorithm (from FIG.7, phases 701-1 and 701-2 may be considered the multi-phase algorithm for initially convolving and pooling. The fusing involves both of these phases).
Referring to claim 19, Jiao has taught the method of claim 15, wherein columns of the data tensor represent different segments of data in the data tensor (from FIG.6B, a tensor is an image which contains columns, which represent different data segments in the tensor. For instance, the leftmost column in an image represents the left edge of the tensor, whereas the rightmost column would represent the right edge).
Referring to claim 20, Jiao has taught the method of claim 15, wherein the multiple hardware execution units execute operations of the fused phase (again, from 703-e, multiple execution units execute convolution and pooling operations).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to David J. Huisman whose telephone number is 571-272-4168. The examiner can normally be reached on Monday-Friday, 9:00 am-5:30 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta, can be reached at 571-270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/David J. Huisman/Primary Examiner, Art Unit 2183