DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114.
Applicant's submission filed on 06/09/2026 has been entered.
Response to Argument
Applicant's arguments filed 06/09/2026 ("Arguments/Remarks") have been fully considered but they are not persuasive.
Argument – 1: (page: 9 – 10) applicant contends: “Amended claim 1 explicitly recites a multi-stage technical pipeline: (i) parsing a model based on architecture detection; (ii) fusing operators according to data flow granularity; (iii) reconstructing and rewriting the second operator according to the design principles of the target data flow architecture (e.g., Specification at paragraphs [0043]-[0044]); and (iv) deploying the converted model. These are not "mental processes" or "pencil and paper" tasks. As noted in the Specification at paragraphs [0031]-[0046], the fusion of operators and the reconstruction of computation graphs based on specific hardware-level "design principles" are complex, machine-level operations that are fundamentally dependent on the underlying hardware's physical execution constraints. A human analyst cannot mentally manage these high-dimensional, hardware-specific graph transformations.”
Regarding the above argument, the Examiner respectfully notes that the amended limitations still contains abstract idea, such as:
“parsing, [ ], a target deep learning model into an intermediate representation of an instruction set computation graph in response to detecting that the target deep learning model is based on an instruction set architecture” - Under the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves observing characteristics of a model, determining whether it is based on an instruction set architecture and organizing the model into intermediate representation.
Fusing, [ ], a first operator in the intermediate representation of the instruction set computation graph into a second operator in an intermediate representation of a data flow computation graph according to the operator granularity of a data flow - Under the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves evaluating operators and their granularity and deciding how they should grouped or combined within another representation, so the patent eligible analysis proceeds to Step 2A Prong 2. See the other abstract ideas in 35 USC § 101 35 U.S.C. 101 section.
In addition, the specification describes converting an instruction set computational graph into a data flow computational graph and combining operators to enable more efficient execution, including concurrent operation. However, merely dividing or grouping operators for simultaneous execution does not, by itself, demonstrate a technological improvement. These concepts are characterized at a high level of data processing. Further, the cited paragraphs ([0031] – [0046]), describes converting an instruction set computing graph into data flow computational graph and fusing operators for concurrent processing of operations and further indicates that such operator fusion may improve processing efficiency, and reduce computation time. However, merely dividing or grouping operators for simultaneous execution does not, by itself demonstrate a technological improvement or improve functionality. Further, the amended claims do not recite hardware incompatibility or adaption to incompatible hardware platforms. Rather, the alleged improvements are bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art.
Applicant’s arguments (pgs. 11 – 19) with respect to amended claim(s) have been considered but are moot, because arguments/remarks are directed to amended claim limitations that were not previously examined by the examiner. The rejections are noted in the current office action to address amended claim limitations.
As to the remaining dependent claims, applicant argue that they are allowable due to their respective direct and indirect dependencies upon one of the aforementioned Independent claims. The Examiner respectfully disagrees, Independent claims were not allowable as stated in the paragraph above in this “Response to Arguments” section in this office action.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1 – 4 and 7 – 12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e. an abstract idea) without significantly more.
In step 1, of the 101-analysis set forth in the MPEP 2106, the examiner has determined
that the following limitations recite a process that, under the broadest reasonable interpretation, falls within one or more statutory categories (processes).
In step 2A prong 1, of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following limitations recite a process that, under broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components:
Regarding claim 1:
Parsing [ ] a target deep learning model into an intermediate representation of an instruction set computation graph, in response to detecting that the target deep learning model is based on an instruction set architecture;
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves observing characteristics of a model, determining whether it is based on an instruction set architecture and organizing the model into intermediate representation. See (MPEP 2106.04)).
Fusing, [ ], a first operator in the intermediate representation of the instruction set computation graph into a second operator in an intermediate representation of a data flow computation graph according to the operator granularity of a data flow
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves evaluating operators and their granularity and deciding how they should grouped or combined within another representation. See (MPEP 2106.04)).
reconstructing and rewriting, [ ], the second operator in the intermediate representation of the data flow computation graph according to a design principle of a data flow architecture operating the target deep learning model, to adjust the intermediate representation of the data flow computation graph to an intermediate representation of a customized architecture
(i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves comparing an existing representation with requirements of a customized architecture and modifying the representation based on that comparison. See (MPEP 2106.04)).
If the claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process, but for the recitation of generic computer components, then it falls within the mental process. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2 of the 101-analysis, set forth in MPEP 2106, the examiner has determined that
the following additional elements do not integrate this judicial exception into a practical application:
A computer-implemented conversion method for a deep learning model on a computer comprising a processor, the method comprises
Deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation which does not amount to more than a recitation of the words "apply it" (or an equivalent), such as mere instructions to implement an abstract idea on a computer. See MPEP 2106.05(f)).
… using the processor,…
Deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation which does not amount to more than a recitation of the words "apply it" (or an equivalent), such as mere instructions to implement an abstract idea on a computer. See MPEP 2106.05(f)).
wherein the instruction set computation graph defines types of operator and an operation rule between operators of the target deep learning model;
Deemed insufficient to transform the judicial exception to a patentable invention because the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
obtaining [ ] a converted target data flow network model corresponding to the target deep learning model according to the intermediate representation of the customized architecture;
Deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity, See MPEP (2106.05(g))).
wherein the intermediate representation of the customized architecture comprises types of operator and a connection relationship between operators of the target data flow network model.
Deemed insufficient to transform the judicial exception to a patentable invention because the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
deploying, using the processor, the converted target data flow network model in the data flow architecture.
Deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity, See MPEP (2106.05(g))).
In Step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the
claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception:
Regarding limitation (I and II) recite mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f).
Regarding limitation (IV and VI), additional elements considered extra/post solution activity, as analyzed above, are activity that are well-understood routine and conventional, specifically: the courts have recognized the computer functions as well‐understood, routine, and conventional functions.
Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II).
Regarding limitation (III and V), additional elements are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because they generally link the judicial exception to the technology environment, see MPEP 2106.05(h).
As analyzed above, the additional elements, analyzed above, do not integrate the noted judicial exception into a practical application because they do not impose any meaningful limits on practicing the abstract idea. Therefore, the claim is directed to an abstract idea.
Regarding claim 7,
The rest of the limitations recite similar subject matter as claim 1, so are rejected under the same rationale.
A conversion apparatus for a deep learning model, comprising: a storage apparatus configured to store one or more programs
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement:
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
In Step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the
claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception:
Regarding limitation (I and II), recite mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f).
Regarding claim 9,
The rest of the limitations recite similar subject matter as claim 1, so are rejected under the same rationale.
A server, comprising: one or more processors, and a storage apparatus configured to store one or more programs;
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement a conversion method for a deep learning model
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
In Step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the
claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception:
Regarding limitation (I and II) recite mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f).
Regarding claim 2, dependent upon claim 1, and fail to resolve the deficiencies identified above by
integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein the target deep learning model comprises a first operator granularity, the intermediate representation of the instruction set computation graph comprises a second operator granularity, and the intermediate representation of the data flow computation graph comprises a third operator granularity
The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 8, recite similar subject matter as claim 2, so is rejected under the same rationale.
Regarding claim 3, dependent upon claim 2, and fail to resolve the deficiencies identified above by
integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein the first operator granularity is the same as the second operator granularity
The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 11, recite similar subject matter as claim 3, so is rejected under the same rationale.
Regarding claim 4, dependent upon claim 2, and fail to resolve the deficiencies identified above by
integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
wherein the second operator granularity is less than the third operator granularity.
The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim 12, recite similar subject matter as claim 4, so is rejected under the same rationale.
Regarding claim 10, dependent upon claim 1, and fail to resolve the deficiencies identified above by
integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites:
A non-transitory computer-readable storage medium storing a computer program
Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f).
The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above.
wherein the computer program when executed by a processor, implements the conversion method for a deep learning model according to claim 1
The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h).
Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 7 and 9 – 10 are rejected under 35 U.S.C. 103 as being unpatentable over Brady et al., Pub. No.: US20190391796A1, in view of Zhang Pub. No.: US10318259B2, Reference – A, Pub. No.: CN111160551A (google translation) and Milojicic et al., Pub. No.: US20200097440A1.
Regarding claim 1, Brady teaches: A computer-implemented conversion method for a deep learning model on a computer comprising a processor, the method comprises:
(Brady, “[0030] In some embodiments, an example compiler (e.g., 105), such as an example neural network compiler such as discussed herein, as well as other components, may be implemented in software stored in memory 215, and operate on the processor 210 [A computer-implemented conversion method for a deep learning model on a computer comprising a processor].”)
parsing, using the processor, a target deep learning model into an intermediate representation of an instruction set computation graph,
(Brady, “[0094] … For instance, the barrier task objects may be inserted into control flows of the intermediate representation modeled by the control model. For instance, the compiler may parse [parsing, using the processor] the control flows represented in the intermediate representation and identify opportunities for the use of hardware barrier resources of the target device (e.g., by identifying dependencies between operations/tasks in the control flow) [a target deep learning model into an intermediate representation of an instruction set computation graph]…”)
in response to detecting that the target deep learning model is based on an instruction set architecture,
(Brady, “[0037] … Each memory slice 412 a-h may be associated with a corresponding one of SHAVE processors (305 a-h). Further, each SHAVE processors (305 a-h) can also include an instruction unit (e.g., 408) into which instructions may be loaded. A particular embodiment in which the processor includes a SHAVE, the SHAVE can include one or more of a reduced instruction set computer (RISC), a digital signal processor (DSP), a very long instruction word (VLIW), and/or a graphics processing unit (GPU) [in response to detecting that the target deep learning model is based on an instruction set architecture]…”)
fusing, using the processor, a first operator in the intermediate representation of the instruction set computation graph into a second operator in an intermediate representation of a data flow computation graph according to the operator granularity of a data flow.
(Brady, “[0073] … Adaptation passes 1236 may be compilation passes, which identify opportunities (independent of the target hardware) to modify the neural network graph itself and potentially simplify and optimize operation and data flows associated with the neural network, such as through fusion compilation passes (e.g., to combine two operations into a single operation) or replacement compilation passes (e.g., replace operations with functionally equivalent and more efficient or adaptable replacement operations), among other examples [fusing, using the processor, a first operator in the intermediate representation of the instruction set computation graph into a second operator in an intermediate representation of a data flow computation graph according to the operator granularity of a data flow]. Such compilation passes may identify hardware-agnostic opportunities, rooted in the underlying mathematics of the operations to be performed to implement the neural network, to generate a pared, more efficient version of the neural network (and reflect these modifications in a transformation of the intermediate representation graph).”)
reconstructing and rewriting, using the processor, the second operator in the intermediate representation of the data flow computation graph
(Brady, “[0052] … A set of sub-models (e.g., 1005, 1010, 1015) may be generated and encapsulated within the intermediate representation 140 to provide a configurable representation of a mathematical structure (e.g., the computation model of the intermediate representation) of the neural network described in graph 110 [reconstructing and rewriting, using the processor, the second operator in the intermediate representation of the data flow computation graph], for instance, in the form of one or more computation graphs from which a binary may be constructed, among other example implementations…”)
according to a design principle of a data flow architecture operating the target deep learning model,
(Brady, “[0104] … A set of compilation passes may be performed, based at least in part on a target descriptor identifying the particular resources of a target computing device that is to implement the neural network [to a design principle of a data flow architecture operating the target deep learning model]. Each compilation pass may transform the intermediate representation of the neural network at some level (e.g., changing certain sub-model graphs of the intermediate representation) to realize optimizations or modifications determined through the compilation pass…”)
obtaining, using the processor, a converted target data flow network model corresponding to the target deep learning model according to the intermediate representation of the customized architecture,
(Brady, “[0104] … A set of compilation passes may be performed, based at least in part on a target descriptor identifying the particular resources of a target computing device that is to implement the neural network [obtaining, using the processor, a converted target data flow network model corresponding to the target deep learning model]. Each compilation pass may transform the intermediate representation of the neural network at some level (e.g., changing certain sub-model graphs of the intermediate representation) [according to the intermediate representation of the customized architecture] to realize optimizations or modifications determined through the compilation pass…”)
wherein the intermediate representation of the customized architecture comprises types of operator and a connection relationship between operators of the target data flow network model; and
(Brady, “[0073] … Adaptation passes 1236 may be compilation passes, which identify opportunities (independent of the target hardware) to modify the neural network graph itself and potentially simplify and optimize operation and data flows associated with the neural network [and a connection relationship between operators of the target data flow network model], such as through fusion compilation passes (e.g., to combine two operations into a single operation) or replacement compilation passes (e.g., replace operations with functionally equivalent and more efficient or adaptable replacement operations), among other examples [wherein the intermediate representation of the customized architecture comprises types of operator].
Brady does not teach:
wherein the instruction set computation graph defines types of operator and an operation rule between operators of the target deep learning model;
to adjust the intermediate representation of the data flow computation graph to an intermediate representation of a customized architecture;
deploying, using the processor, the converted target data flow network model in the data flow architecture.
Reference – A teaches:
wherein the instruction set computation graph defines types of operator and an operation rule between operators of the target deep learning model;
(Reference - A, "[0104] Step 2): The general-purpose processor checks the operators in the original subgraph according to the rules of the operators [wherein the instruction set computation graph defines types of operator and an operation rule between operators] in the learning library of the artificial intelligence processor, and performs a second division of the original subgraph based on the check results to obtain the target subgraph [of the target deep learning model].")
Reference – A and Brady are related to the same field of endeavor (i.e.: machine learning framework). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Reference – A with teachings of Brady to add generating executable binary instructions for an artificial intelligence processor based on the fusion operator’s operation instructions. (Reference – A, Abstract).
Brady in view of Reference – A do not teach:
to adjust the intermediate representation of the data flow computation graph to an intermediate representation of a customized architecture;
deploying, using the processor, the converted target data flow network model in the data flow architecture.
Zhang teaches:
to adjust the intermediate representation of the data flow computation graph to an intermediate representation of a customized architecture;
(Zhang, (col. 4 line [51 – 61]), “That control flow program is illustrated as source program 40. As described in greater detail below, when compiler 30 executes, compiler 30 may use data flow convertor 32 to convert source program 40 into data flow program 54. For instance, as illustrated, compiler 30 may copy source program 40 from NVS 14 into RAM 12, and compiler 30 may then use data flow convertor 32 to convert source program 40 into data flow program 54 [to adjust the intermediate representation of the data flow computation graph to an intermediate representation of a customized architecture]. Data flow program may then execute on DFP 20. In addition or alternatively, compiler 30 may copy data flow program 54 from RAM 12 to NVS 14 for future utilization.”)
Zhang, Brady and Reference – A are related to the same field of endeavor (i.e.: machine learning framework). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Zhang with teachings of Brady and Reference – A to add automatic transfer from a control flow representation to a data flow representation before code generation to enable more efficient scheduling. (Zhang, Abstract).
Brady in view of Reference – A and Zhang do not teach:
deploying, using the processor, the converted target data flow network model in the data flow architecture.
Milojicic teaches:
deploying, using the processor, the converted target data flow network model in the data flow architecture.
(Milojicic, “[0013] Programming approaches for computing in memory are generally achieved through rigid approaches including, for example, neural networks having fixed weight programming, ternary content-accessible memory, or associative memory. Implementations disclosed herein may provide data flow models [the converted target data flow network model in the data flow architecture] for computing in memory that allow computing in memory to achieve increased processing speeds, as well as deploy optimized hardware on the fly [deploying, using the processor]. Such implementations may thereby provide high degrees of programmability and reconfigurability for computing in memory applications resulting in improved performance, reduced energy, and easier use.”)
Milojicic, Brady, Reference – A and Zhang are related to the same field of endeavor (i.e.: machine learning framework). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Milojicic with teachings of Brady, Reference – A and Zhang to add an execution architecture in which data is processed within a specialized computing unit to improve processing efficiency by reducing data movement and enhancing execution performance. (Milojicic, Abstract).
Regarding claim 7, Brady teaches: A conversion apparatus for a deep learning model, comprising:
a storage apparatus configured to store one or more programs; wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement:
(Brady, “[0029] In some implementations, an example system 205 may have memory 215 such as a computer readable medium, flash memory, a magnetic disk drive, an optical drive, a programmable read-only memory (PROM), and/or a read-only memory (ROM) [a storage apparatus configured to store one or more programs]. The system 205 may be configured with one or more processors 210 [when executed by the one or more processors, cause the one or more processors to implement] that process instructions and run software that may be stored in memory 215. The processor 205 can also communicate with the memory 215 and interfaces 220 to communicate with other devices. The processor 210 can be any applicable processor such as a system-on-a-chip that combines a CPU, an application processor, and flash memory, or a reduced instruction set computing (RISC) processor.”)
The rest of the limitations are analogous to claim 1, so are rejected under similar rationale.
Regarding claim 9, Brady teaches: A server, comprising:
(Brady, “[0033] In some implementations, an example system 205 may be implemented as a computer device, such as a personal computing device, mobile computing device, server computing system (e.g., a rack scale, blade server, or other server computer), among other examples [A server].”)
one or more processors, and a storage apparatus configured to store one or more programs; wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement a conversion method for a deep learning model
(Brady, “[0029] In some implementations, an example system 205 may have memory 215 such as a computer readable medium, flash memory, a magnetic disk drive, an optical drive, a programmable read-only memory (PROM), and/or a read-only memory (ROM) [one or more processors, and a storage apparatus configured to store one or more programs; wherein the one or more programs]. The system 205 may be configured with one or more processors 210 [wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement a conversion method for a deep learning mode] that process instructions and run software that may be stored in memory 215. The processor 205 can also communicate with the memory 215 and interfaces 220 to communicate with other devices. The processor 210 can be any applicable processor such as a system-on-a-chip that combines a CPU, an application processor, and flash memory, or a reduced instruction set computing (RISC) processor.”)
The rest of the limitations are analogous to claim 1, so are rejected under similar rationale.
Regarding claim 10, Brady in view of Reference – A, Zhang and Milojicic, teach the method of claim 1.
Brady further teaches: A non-transitory computer-readable storage medium storing a computer program, wherein the computer program when executed by a processor, implements the conversion method for a deep learning model according to claim 1.
(Brady, “[0030] In some embodiments, an example compiler (e.g., 105), such as an example neural network compiler such as discussed herein, as well as other components, may be implemented in software stored in memory 215, and operate on the processor 210. The memory 215 can be a non-transitory computer readable medium, flash memory, a magnetic disk drive, an optical drive, a programmable read-only memory (PROM), a read-only memory (ROM), or any other memory or combination of memories [A non-transitory computer-readable storage medium storing a computer program,]. The software can run on a processor capable of executing computer instructions or computer code [wherein the computer program when executed by a processor, implements the conversion method for a deep learning model according to claim 1].”)
Claim(s) 2 – 3, 8 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Brady in view of Reference – A, Zhang, Milojicic and in further view of Du et al., Pub. No.: US20200089535A1.
Regarding claim 2, Brady in view of Reference – A, Zhang, Milojicic teach the method of claim 1.
Brady in view of Reference – A, Zhang, Milojicic do not teach:
wherein the target deep learning model comprises a first operator granularity, the intermediate representation of the instruction set computation graph comprises a second operator granularity, and the intermediate representation of the data flow computation graph comprises a third operator granularity
Du teaches:
wherein the target deep learning model comprises a first operator granularity, the intermediate representation of the instruction set computation graph comprises a second operator granularity, and the intermediate representation of the data flow computation graph comprises a third operator granularity
(Du, “[0010] In some embodiment, the granularity task segmentation unit includes at least one of the following units: a first granularity task segmentation unit configured to take the whole task as one of the subtasks; [a first operator granularity, the intermediate representation of the instruction set computation graph] a second granularity task segmentation unit configured to divide sample data associated with the task into one or more subset of sample data [comprises a second operator granularity, and the intermediate representation of the data flow computation graph], and identify a computation of each subset of sample data as one of the subtasks; a third granularity task segmentation unit configured to segment the task according to layer types of a neural network [comprises a third operator granularity], where computation for layers of the same layer type is identified as one of the subtasks; fourth granularity task segmentation unit configured to segment the task according to an interlayer structure of the neural network, wherein computation for multiple adjacent layers is identified as one of the subtasks; and a fifth granularity task segmentation unit configured to segment the task according to intra-layer structures of the neural network to segment computation types in each of the layers of the neural network into subtasks.”)
Du, Brady, Reference – A, Zhang, Milojicic are related to the same field of endeavor (i.e.: deep learning framework). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Du with teachings of Brady, Reference – A, Zhang, Milojicic to improve processing performance and minimize overhead by aligning sub-model partitioning with fine-grained task segmentation and hardware configuration. (Du, Abstract).
Claim 8 recites analogous limitation as claim 2, so is rejected under the same rationale.
Regarding claim 3, Brady in view of Reference – A, Zhang, Milojicic and Du teach the method of claim 2.
Du further teaches: wherein the first operator granularity is the same as the second operator granularity.
(Du, “[0010] In some embodiment, the granularity task segmentation unit includes at least one of the following units: a first granularity task segmentation unit [wherein the first operator granularity] configured to take the whole task as one of the subtasks; a second granularity task segmentation unit [as the second operator granularity] configured to divide sample data associated with the task into one or more subset of sample data, and identify a computation of each subset of sample data as one of the subtasks; a third granularity task segmentation unit configured to segment the task according to layer types of a neural network, where computation for layers of the same layer type [is the same as] is identified as one of the subtasks; fourth granularity task segmentation unit configured to segment the task according to an interlayer structure of the neural network, wherein computation for multiple adjacent layers is identified as one of the subtasks; and a fifth granularity task segmentation unit configured to segment the task according to intra-layer structures of the neural network to segment computation types in each of the layers of the neural network into subtasks.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Du with teachings of Brady, Reference – A, Zhang, Milojicic for the same reasons disclosed for claim 3.
Claim 11 recites analogous limitation as claim 3, so is rejected under the same rationale.
Claim(s) 4 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Brady in view of Reference – A, Zhang, Milojicic, Du and in further view of Cohen et al., Pub. No.: US20190102671A1.
Regarding claim 4, Brady in view of Reference – A, Zhang, Milojicic and Du teach the method of claim 2.
Brady in view of Reference – A, Zhang, Milojicic, Du do not teach:
wherein the second operator granularity is less than the third operator granularity.
Cohen teaches:
wherein the second operator granularity is less than the third operator granularity
(Cohen, “[0327] There are further disclosed one or more tangible, non-transitory computer-readable mediums, wherein the two or more intermediate operators are Op1, Op2, and Op3, wherein the output operator is assigned a first value if Op1 [wherein the second operator granularity] (i.e.: Op1 corresponds to the second operator granularity) is less than Op3 [the third operator granularity] (i.e.: Op3 corresponds to the third operator granularity), a second value is Op1 is between Op3 and Op2, and a third value otherwise.”)
Cohen, Brady, Reference – A, Zhang, Milojicic and Du are related to the same field of endeavor (i.e.: deep learning framework). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Cohen with teachings of Brady, Reference – A, Zhang, Milojicic, Du to add pipeline sub-models across devices to achieve faster and more efficient CNN computations within each hardware unit (Cohen, Abstract).
Claim 12 recites analogous limitation as claim 4, so is rejected under the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Rotem, et al., "Glow: Graph lowering compiler techniques for neural networks." (2018).
Rotem provide a useful compiler toolkit that will allow hardware developers to focus on implementing efficient acceleration hardware, each of which likely differ in capabilities, and use Glow for automating compilation tasks such as instruction selection, memory allocation and graph scheduling.
Chadha, et al., "Performance Analysis of Accelerated Linear Algebra Compiler for TensorFlow." 2017.
Chadha analyze the performance of XLA compilation tool on machine learning algorithms like Convolutional Neural Networks, Long Short Term Memory and custom control flow graphs.
Any inquiry concerning this communication or earlier communications from the examiner
should be directed to MATIYAS T MARU whose telephone number is (571)270-0902 or via email: matiyas.maru@uspto.gov. The examiner can normally be reached Monday 8:00am - Friday 4:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a
USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to
use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor,
Michelle Bechtold can be reached on (571)431-0762. The fax phone number for the organization were this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from
Patent Center. Unpublished application information in Patent Center is available to registered users.
To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit
https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and
https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional
questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like
assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA)
or 571-272-1000.
/M.T.M./ Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148