DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response the application filed 04/25/2024 and the preliminary amendment filed 07/09/2024. In the amendment, claims 1-20 were cancelled, claims 21-40 were added, and no claims were amended. As such, claims 21-40 are pending and have been examined. Claims 21-40 are rejected.
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. The present application is a continuation of Application No. 17/874,876 filed on 07/27/2022 (now U.S. patent number 12,001,944 B2), which is a continuation of Application No. 15/581,152 filed on 04/28/2017 (now U.S. patent number 11,410,024 B2).
Information Disclosure Statement
Acknowledgment is made of the information disclosure statement filed 07/09/2024, which complies with 37 CFR 1.97. As such, the information disclosure statement has been placed in the application file and the information referred to therein has been considered by the examiner.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference characters not mentioned in the description:
Reference character 859 shown in Figure 8C is not found in the detailed description.
Reference characters 2112 and 2116 shown in Figure 21 are not found in the detailed description.
Reference characters 1450A, 550N and 560N shown in Figure 22 are not found in the detailed description.
Reference characters 1731, 1750, 1754 and 1778 shown in Figure 25 are not found in the detailed description.
Reference characters 3120A-B, 3125A-B and 3130A-B shown in Figure 31 are not found in the detailed description.
The drawings are further objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference signs mentioned in the description:
“2554”, “2578” and “2531” are mentioned in description of FIG. 25 in paragraphs 300, 301 and 302 and these reference signs do not appear in FIG. 25; and
“3020A-3020B”, “3025A-3025B”, and “3030A-3030B” are mentioned in the description of FIG. 31 in paragraph 330 and these reference signs do not appear in FIG. 31.
The drawings are also objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “873” has been used to designate both “process 873” and communications between processes 871 and 873 in FIG. 8D, and because reference character “1108” has been used to designate both “Cache Memory 1108” and “I/O Hub 1108” in FIG. 11. It appears that the second occurrence of “873” in FIG. 8D should be “875” (see, e.g., the description of FIG. 8C in paragraph 189 which states “In the illustrated embodiment, lower precision weight vector and wait for global weight are communicated back and forth 875 between process 871 and process 873.”).
Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
In paragraphs 337, 344 and 351, the recitations of “to determine optimal point” in are grammatically incorrect and appear to be missing the word “an” between “optimal” and “point”. Appropriate correction is required.
The specification is also objected to because reference character 859 shown in Figure 8C is not described in applicant’s specification (see, e.g., paragraphs 186-190 describing FIG. 8C). Appropriate correction is required.
The specification is additionally objected to because the description of FIG. 8C in paragraphs 189-190 includes the reference character “875” which does not appear in FIG. 8D. As indicated above in the objections to the drawings, reference character 873 is used in FIG. 8C to designate both “process 873” and communications between processes 871 and 873.
The specification is further objected to because reference characters 2112 and 2116 shown in Figure 21, and 1450A, 550N and 560N shown in Figure 22 are not described in applicant’s specification (see, e.g., paragraphs 265-271 describing FIG. 21 and paragraphs 272-275 describing FIG. 22). Appropriate correction is required.
The specification is also objected to because the description of FIG. 22 in paragraph 275 includes the reference characters “2250A-2250N, 2260A-2260N” for sub-cores which do not appear in FIG. 22. As indicated above in the objections to the drawings, reference characters 1450A, 550N and 560N are used in Figure 22 for sub-cores, but are not found in the detailed description. Appropriate correction is required.
The specification is further objected to because reference characters 1731, 1750, 1754 and 1778 shown in Figure 25 are not described in applicant’s specification (see, e.g., paragraphs 293-304 describing FIG. 25). Appropriate correction is required.
The specification is additionally objected to because the description of FIG. 25 in paragraphs 300, 301 and 302 includes the reference characters “2554”, “2578” and “2531”, which do not appear in FIG. 25. As indicated above in the objections to the drawings, reference characters 1731, 1750, 1754 and 1778 are used in FIG. 25. Appropriate correction is required.
The specification is also objected to because reference characters 3120A-B, 3125A-B and 3130A-B shown in Figure 31 are not described in applicant’s specification (see, e.g., paragraphs 330-331 describing FIG. 31). Appropriate correction is required.
The specification is additionally objected to because the description of FIG. 31 in paragraph 330 includes the reference characters “3020A-3020B”, “3025A-3025B”, and “3030A-3030B”, which do not appear in FIG. 31. As indicated above in the objections to the drawings, reference characters 3120A-B, 3125A-B and 3130A-B are used in FIG. 31. Appropriate correction is required.
Applicant is reminded of the proper content of an abstract of the disclosure.
A patent abstract is a concise statement of the technical disclosure of the patent and should include that which is new in the art to which the invention pertains. The abstract should not refer to purported merits or speculative applications of the invention and should not compare the invention with the prior art.
If the patent is of a basic nature, the entire technical disclosure may be new in the art, and the abstract should be directed to the entire disclosure. If the patent is in the nature of an improvement in an old apparatus, process, product, or composition, the abstract should include the technical disclosure of the improvement. The abstract should also mention by way of example any preferred modifications or alternatives.
Where applicable, the abstract should include the following: (1) if a machine or apparatus, its organization and operation; (2) if an article, its method of making; (3) if a chemical compound, its identity and use; (4) if a mixture, its ingredients; (5) if a process, the steps.
Extensive mechanical and design details of an apparatus should not be included in the abstract. The abstract should be in narrative form and generally limited to a single paragraph within the range of 50 to 150 words in length.
See MPEP § 608.01(b) for guidelines for the preparation of patent abstracts.
The abstract of the disclosure is objected to because the recitation of “to determine optimal point” in the second sentence of the abstract is grammatically incorrect and appears to be missing the word “an” between “optimal” and “point”. Correction is required. See MPEP § 608.01(b).
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 3-8, 10-15 and 17-20 of U.S. Patent No. 11,410,024. Although the claims at issue are not identical, they are not patentably distinct from each other because all of the limitations of claims 21-40 in the present application are covered by claims 1, 3-8, 10-15 and 17-20 of U.S. Patent No. 11,410,024 (please see the table below).
Regarding independent claims 21, 28 and 35, claims 1, 8 and 15 of U.S. Patent No. 11,410,024 teach the claimed invention as shown in the table below.
Instant Application No. 18/646,021 (as amended 07/09/2024)
U.S. Patent No. 11,410,024
21. An apparatus comprising:
a graphics processor to:
cause a neural network application to utilize a library comprising machine
learning primitives, wherein the machine learning primitives are usable to analyze
patterns observed in a distributed gradient synchronization implemented by the neural network application;
determine, using the machine learning primitives of the library, a point to apply
frequency scaling in the graphics processor without degrading performance of the neural network application, the point determined based on analysis of the patterns generated by the distributed gradient synchronization; and
determine, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.
1. An apparatus comprising:
a graphics processor to:
detect one or more sets of data from one or more sources over one or more networks;
cause a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;
implement, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure;
determine, using the machine learning primitives of the library as implemented by the neural network application, a point to apply frequency scaling in the graphics processor that does not degrade performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure; and
determine, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.
22. The apparatus of claim 21, wherein the point is determined through the
distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a
root of the tree-like structure.
1. An apparatus comprising: …
the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure; …
the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure
23. The apparatus of claim 21, wherein the graphics processor is further to introduce sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
3. The apparatus of claim 1, wherein the graphics processor is further operable to introduce a sparse matrix representation for weights to overlap communication and computation across the one or more nodes associated with the neural network application to reduce communication costs.
24. The apparatus of claim 21, wherein the graphics processor is further to
automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
4. The apparatus of claim 1, wherein the graphics processor is further to automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
25. The apparatus of claim 24, wherein the graphics processor is further to provide one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to
a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
5. The apparatus of claim 4, wherein the graphics processor is further to provide one or more of successful execution information obtained from successful execution of programs and failed execution information obtained from failed execution of programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
26. The apparatus of claim 21, wherein the graphics processor is further to perform local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication.
6. The apparatus of claim 1, wherein the graphics processor is further to perform local error propagation by computing high precision and low precision for local weights and compute local errors at each of the one or more nodes, wherein performing the local error propagation further comprises facilitating weight synchronization across the one or more nodes to track the local errors for accuracy and reduced communication.
27. The apparatus of claim 21, wherein the apparatus comprises an autonomous
machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor,
wherein the graphics processor is co-located with an application processor on a common semiconductor package.
7. The apparatus of claim 1, wherein the apparatus comprises an autonomous machine comprising one or more of a vehicle, a device, or an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
28. A method comprising:
causing a neural network application to utilize a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze patterns observed in a distributed gradient synchronization implemented by the neural network application;
determining, using the machine learning primitives of the library, a point to apply
frequency scaling in a computing device hosting the neural network application without degrading performance of the neural network application at the computing device, the point determined based on analysis of the patterns generated by the distributed gradient
synchronization; and
determining, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.
8. A method comprising:
detecting, by a graphics processor, one or more sets of data from one or more sources over one or more networks;
causing, by the graphics processor, a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;
implementing, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure;
determining, using the machine learning primitives of the library as implemented by the neural network application, a point to apply frequency scaling in the graphics processor that does not degrade performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure; and
determining, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.
29. The method of claim 28, wherein the point is determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a root of the tree-like structure.
8. A method comprising: …
the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure …
the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure
30. The method of claim 28, further comprising introducing sparse matrix
representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
10. The method of claim 8, further comprising introducing sparse matrix representation for weights to overlap communication and computation across the one or more nodes associated with the neural network application to reduce communication costs.
31. The method of claim 28, further comprising automatically analyzing failed
execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
11. The method of claim 8, further comprising automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
32. The method of claim 31, further comprising providing one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
12. The method of claim 11, further comprising providing one or more of successful execution information obtained from successful execution of programs and failed execution information obtained from failed execution of programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
33. The method of claim 28, further comprising performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication.
13. The method of claim 8, further comprising performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of the one or more nodes, wherein performing the local error propagation further comprises facilitating weight synchronization across the one or more nodes to track the local errors for accuracy and reduced communication.
34. The method of claim 28, wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including a graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
14. The method of claim 8, wherein the graphics processor is part of an autonomous machine comprising one or more of a vehicle, a device, or an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
35. A non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:
causing a neural network application to utilize a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze patterns observed in a distributed gradient synchronization implemented by the neural network application;
determining, using the machine learning primitives of the library, a point to apply
frequency scaling in the computing device without degrading performance of the neural network application, the point determined based on analysis of the patterns generated by the distributed
gradient synchronization; and
determining, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.
15. At least one non-transitory machine-readable medium comprising instructions that when executed by a local computing device, cause the local computing device to perform operations comprising:
detecting, by a graphics processor of the local computing device, one or more sets of data from one or more sources over one or more networks;
causing, by the graphics processor, a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;
implementing, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure;
determining, using the machine learning primitives of the library as implemented by the neural network application, a point to apply frequency scaling in the graphics processor that does not degrade performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure; and
determining, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.
36. The non-transitory machine-readable medium of claim 35, wherein the point is
determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and
communicate up to a root of the tree-like structure.
15. At least one non-transitory machine-readable medium …
implementing, using the neural network application, the distributed gradient synchronization using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure; …
the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization implemented via the tree structure
37. The non-transitory-machine-readable medium of claim 35, wherein the operations further comprise introducing sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
17. The non-transitory machine-readable medium of claim 15, wherein the operations further comprise introducing sparse matrix representation for weights to overlap communication and computation across the one or more nodes associated with neural network application to reduce communication costs.
38. The non-transitory-machine-readable medium of claim 35, wherein the operations further comprise automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance
counters.
18. The non-transitory machine-readable medium of claim 15, wherein the operations further comprise automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
39. The non-transitory-machine-readable medium of claim 38, wherein the operations further comprise providing one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
19. The non-transitory machine-readable medium of claim 18, wherein the operations further comprise providing one or more of successful execution information obtained from successful execution of programs and failed execution information obtained from failed execution of programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
40. The non-transitory-machine-readable medium of claim 35, wherein the operations further comprise performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the
neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and
reduced communication, wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous
machine comprises one or more processors including a graphics processor, wherein the graphics
processor is co-located with an application processor on a common semiconductor package.
20. The non-transitory machine-readable medium of claim 15, wherein the operations further comprise performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of the one or more nodes, wherein performing the local error propagation further comprises facilitating weight synchronization across the one or more nodes to track the local errors for accuracy and reduced communication, wherein the computing device comprises an autonomous machine comprising one or more of a vehicle, a device, or an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
As shown in the table above, claims 1, 3-8, 10-15 and 17-20 of U.S. Patent No. 11,410,024 encompass similar subject matter as independent claims 21-40 of the instant application.
The table above uses the underlined text to highlight the similarities between the two applications and claims within, while the non-underlined text signifies the difference between the claim sets. Upon review of the non-underlined text, a person of ordinary skill in the art would have concluded that the scope of the claim inventions are obvious variants of each other. To further expand on the above table, each claim listed above is either directly from the reference patent and/or obvious as explained below. Claims not directly referenced in the explanation below are believed to be adequately explained by the comparison chart above.
Regarding instant claim 21, the instant claim is substantially identical to claim 1 of U.S. Patent No. 11,410,024. The examiner points out the similarities between claim 21 of this application and claim 1 of the reference patent in this office action for clarity but the reasoning, motivation(s), and rationale apply for claims 28 and 35 as well.
Regarding instant claim 21, this claim is directed to an “apparatus comprising: a graphics processor to” perform operations substantially identical to operations recited in the “apparatus” in claim 1 of the reference patent 11,410,024, except that claim 1 of the reference patent additionally recites “detect one or more sets of data from one or more sources over one or more networks” and “using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure” and instant claim 21 does not. Also, claim 1 of the reference patent recites “analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application”, “the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization” and “the library is to account for skew characteristics associated with the distributed gradient synchronization”, whereas instant claim 21 recites “analyze patterns observed in a distributed gradient synchronization implemented by the neural network application”, “the point determined based on analysis of the patterns generated by the distributed gradient synchronization” and “the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.” However, as shown in the table above, the operations of instant claim 21 are largely identical to the other operations recited in claim 1 of the reference patent 11,410,024. That is, instant claim 21 is a broader version of claim 1 of the reference patent and independent claim 1 of the reference patent encompasses similar subject matter as instant claim 21.
Similarly, instant independent claims 28 and 35 are directed to a method and non-transitory machine-readable medium, respectively that perform steps and operations substantially identical to steps and operations recited in the method and non-transitory machine-readable medium in claims 8 and 15, respectively, of the reference patent 11,410,024, except that claims 8 and 15 of the reference patent additionally recite, using respective similar language, “detecting one or more sets of data from one or more sources over one or more networks” and “using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure” and instant claims 28 and 35 do not. Claims 8 and 15 of the reference patent also recite “analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application”, “the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization” and “the library is to account for skew characteristics associated with the distributed gradient synchronization”, whereas instant claims 28 and 35 recite “analyze patterns observed in a distributed gradient synchronization implemented by the neural network application”, “the point determined based on analysis of the patterns generated by the distributed gradient synchronization” and “the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.” That is, instant claims 28 and 35 are broader versions of claims 8 and 15 of the reference patent and claims 8 and 15 of the reference patent encompass similar subject matter as instant claims 28 and 35.
Therefore, instant application claims 21, 28 and 35 are taught by claims 1, 8 and 15 of U.S. Patent No. 11,410,024.
Regarding dependent claims 22-27, 29-34, and 36-40, which each depend directly or indirectly from independent claims 21, 28 and 35, respectively, claims 1, 8 and 15 of U.S. Patent No. 11,410,024 teach the limitations of claims 21, 28 and 35 as discussed above and shown in the table above. Claims 1 and 3-7, 8 and 10-15, and 17-20 of U.S. Patent No. 11,410,024 further teach the limitations of dependent claims 22-27, 29-34, and 36-40 of the instant application (please see the table above).
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12,001,944. Although the claims at issue are not identical, they are not patentably distinct from each other because all of the limitations of claims 21-40 in the present application are covered by claims 1-20 of U.S. Patent No. 12,001,944 (please see the table below).
Regarding independent claims 21, 28 and 35, claims 1, 8 and 15 of U.S. Patent No. 12,001,944 teach the claimed invention as shown in the table below.
Instant Application No. 18/646,021 (as amended 07/09/2024)
U.S. Patent No. 12,001,944
21. An apparatus comprising:
a graphics processor to:
cause a neural network application to utilize a library comprising machine
learning primitives, wherein the machine learning primitives are usable to analyze
patterns observed in a distributed gradient synchronization implemented by the neural network application;
determine, using the machine learning primitives of the library, a point to apply
frequency scaling in the graphics processor without degrading performance of the neural network application, the point determined based on analysis of the patterns generated by the distributed gradient synchronization; and
determine, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.
1. An apparatus comprising:
a graphics processor to:
cause a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;
determine, using the machine learning primitives of the library, a point to apply frequency scaling in the graphics processor without degrading performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization; and
determine, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.
22. The apparatus of claim 21, wherein the point is determined through the
distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a
root of the tree-like structure.
2. The apparatus of claim 1, wherein the point is determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a root of the tree-like structure.
23. The apparatus of claim 21, wherein the graphics processor is further to introduce sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
3. The apparatus of claim 1, wherein the graphics processor is further to introduce sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
24. The apparatus of claim 21, wherein the graphics processor is further to
automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
4. The apparatus of claim 1, wherein the graphics processor is further to automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
25. The apparatus of claim 24, wherein the graphics processor is further to provide one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to
a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
5. The apparatus of claim 4, wherein the graphics processor is further to provide one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
26. The apparatus of claim 21, wherein the graphics processor is further to perform local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication.
6. The apparatus of claim 1, wherein the graphics processor is further to perform local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication.
27. The apparatus of claim 21, wherein the apparatus comprises an autonomous
machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor,
wherein the graphics processor is co-located with an application processor on a common semiconductor package.
7. The apparatus of claim 1, wherein the apparatus comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
28. A method comprising:
causing a neural network application to utilize a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze patterns observed in a distributed gradient synchronization implemented by the neural network application;
determining, using the machine learning primitives of the library, a point to apply
frequency scaling in a computing device hosting the neural network application without degrading performance of the neural network application at the computing device, the point determined based on analysis of the patterns generated by the distributed gradient
synchronization; and
determining, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.
8. A method comprising:
causing a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;
determining, using the machine learning primitives of the library, a point to apply frequency scaling in a computing device hosting the neural network application without degrading performance of the neural network application at the computing device, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization; and
determining, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.
29. The method of claim 28, wherein the point is determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a root of the tree-like structure.
9. The method of claim 8, wherein the point is determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a root of the tree-like structure.
30. The method of claim 28, further comprising introducing sparse matrix
representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
10. The method of claim 8, further comprising introducing sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
31. The method of claim 28, further comprising automatically analyzing failed
execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
11. The method of claim 8, further comprising automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
32. The method of claim 31, further comprising providing one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
12. The method of claim 11, further comprising providing one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
33. The method of claim 28, further comprising performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication.
13. The method of claim 8, further comprising performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication.
34. The method of claim 28, wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including a graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
14. The method of claim 8, wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including a graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
35. A non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:
causing a neural network application to utilize a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze patterns observed in a distributed gradient synchronization implemented by the neural network application;
determining, using the machine learning primitives of the library, a point to apply
frequency scaling in the computing device without degrading performance of the neural network application, the point determined based on analysis of the patterns generated by the distributed
gradient synchronization; and
determining, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.
15. A non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:
causing a neural network application to implement a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application;
determining, using the machine learning primitives of the library, a point to apply frequency scaling in the computing device without degrading performance of the neural network application, the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization; and
determining, using the library as implemented by the neural network application, a core frequency of the frequency scaling applied at the point, wherein the library is to account for skew characteristics associated with the distributed gradient synchronization to decide the core frequency.
36. The non-transitory machine-readable medium of claim 35, wherein the point is
determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and
communicate up to a root of the tree-like structure.
16. The non-transitory machine-readable medium of claim 15, wherein the point is determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a root of the tree-like structure.
37. The non-transitory-machine-readable medium of claim 35, wherein the operations further comprise introducing sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
17. The non-transitory machine-readable medium of claim 15, wherein the operations further comprise introducing sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
38. The non-transitory-machine-readable medium of claim 35, wherein the operations further comprise automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance
counters.
18. The non-transitory machine-readable medium of claim 15, wherein the operations further comprise automatically analyzing failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
39. The non-transitory-machine-readable medium of claim 38, wherein the operations further comprise providing one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
19. The non-transitory machine-readable medium of claim 18, wherein the operations further comprise providing one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
40. The non-transitory-machine-readable medium of claim 35, wherein the operations further comprise performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the
neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and
reduced communication, wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous
machine comprises one or more processors including a graphics processor, wherein the graphics
processor is co-located with an application processor on a common semiconductor package.
20. The non-transitory machine-readable medium of claim 15, wherein the operations further comprise performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication, wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including a graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
As shown in the table above, claims 1-20 of U.S. Patent No. 12,001,944 encompass similar subject matter as independent claims 21-40 of the instant application.
The table above uses the underlined text to highlight the similarities between the two applications and claims within, while the non-underlined text signifies the difference between the claim sets. Upon review of the non-underlined text, a person of ordinary skill in the art would have concluded that the scope of the claim inventions are obvious variants of each other. To further expand on the above table, each claim listed above is either directly from the reference patent and/or obvious as explained below. Claims not directly referenced in the explanation below are believed to be adequately explained by the comparison chart above.
Regarding instant claim 21, the instant claim is substantially identical to claim 1 of U.S. Patent No. 12,001,944. The examiner points out the similarities between claim 21 of this application and claim 1 of the reference patent in this office action for clarity but the reasoning, motivation(s), and rationale apply for claims 28 and 35 as well.
Regarding instant claim 21, this claim is directed to an “apparatus comprising: a graphics processor to” perform operations substantially identical to operations recited in the “apparatus” in claim 1 of the reference patent 12,001,944, except that claim 1 of the reference patent additionally recites “analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application”, “the point determined based on analysis of the skew pattern generated by the distributed gradient synchronization” and “the library is to account for skew characteristics associated with the distributed gradient synchronization”, whereas instant claim 21 recites “analyze patterns observed in a distributed gradient synchronization implemented by the neural network application”, “the point determined based on analysis of the patterns generated by the distributed gradient synchronization” and “the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.” However, as shown in the table above, the operations of instant claim 21 are largely identical to the other operations recited in claim 1 of the reference patent 12,001,944. That is, instant claim 21 is a slightly broader version of claim 1 of the reference patent and independent claim 1 of the reference patent encompasses similar subject matter as instant claim 21.
Similarly, instant independent claims 28 and 35 are directed to a method and non-transitory machine-readable medium, respectively that perform steps and operations substantially identical to steps and operations recited in the method and non-transitory machine-readable medium in claims 8 and 15, respectively, of the reference patent 12,001,944, except that claims 8 and 15 of the reference patent additionally recite, using respective similar language, “detecting one or more sets of data from one or more sources over one or more networks” and “using a tree structure such that local weight vectors start at one or more nodes represented as leaves of the tree structure and communicate up to a root of the tree structure” and instant claims 28 and 35 do not. Claims 8 and 15 of the reference patent also recite “analyze a skew pattern observed in a distributed gradient synchronization implemented by the neural network application” whereas instant claims 28 and 35 recite “analyze patterns observed in a distributed gradient synchronization implemented by the neural network application”. That is, instant claims 28 and 35 are slightly broader versions of claims 8 and 15 of the reference patent and claims 8 and 15 of the reference patent encompass similar subject matter as instant claims 28 and 35.
Therefore, instant application claims 21, 28 and 35 are taught by claims 1, 8 and 15 of U.S. Patent No. 12,001,944.
Regarding dependent claims 22-27, 29-34, and 36-40, which each depend directly or indirectly from independent claims 21, 28 and 35, respectively, claims 1, 8 and 15 of U.S. Patent No. 12,001,944 teach the limitations of claims 21, 28 and 35 as discussed above and shown in the table above. Claims 2-7, 9-14 and 16-20 of U.S. Patent No. 12,001,944 further teach the limitations of dependent claims 22-27, 29-34, and 36-40 of the instant application (please see the table above).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 21-22, 28-29 and 35-36 are rejected are rejected under 35 U.S.C. 103 as being unpatentable over Roblek et al. (U.S. Patent Application Pub. No. 2017/0330586 A1, cited in applicants’ information disclosure statement filed 7/09/2022, hereinafter “Roblek”) in view of non-patent literature Dettmers, Tim ("8-bit approximations for parallelism in deep learning." arXiv preprint arXiv:1511.04561 (2015). pp. 1 -14, cited in applicants’ information disclosure statement filed 7/09/2022hereinafter “Dettmers”) and Tokui et al. (U.S. Patent Application Pub. No. 2018/0349772 A1, cited in applicants’ information disclosure statement filed 7/09/2022hereinafter “Tokui”), and further in view of non-patent literature Lambert et al. ("Adaptive Frequency Neural Networks for Dynamic Pulse and Metre Perception." ISMIR. Schloss Dagstuhl LZI, 2016, cited in applicants’ information disclosure statement filed 7/09/2022hereinafter “Lambert”). Tokui was filed on April 27, 2018 as a national stage application of PCT application no. PCT/JP2016/004027 filed September 2, 2016, and this date is before the effective filing date of the present application, April 28, 2017. Therefore, Tokui constitutes prior art under 35 U.S.C. 102(a)(2).
With respect to claim 21, Roblek discloses the invention as claimed including an apparatus comprising a processor (see, e.g., paragraphs 104, “data processing apparatus" encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers” [i.e., an apparatus comprising a processor]) … to: cause a neural network application to utilize a library … determine, using … the library, a point to apply frequency scaling in the … processor without degrading performance of the neural network application (see, e.g., paragraphs 102, “The trained parameters define an optimal convolutional mapping of frequency domain features to multi -scaled frequency domain features (step 706). The other neural network layers in the cascaded convolutional neural network system, e.g., the output layer, are able to select and use appropriate features from a concatenated convolutional neural network stage output, enabling the neural network system to tailor and optimize the convolutional mapping of frequency domain features to multi -scaled frequency domain features to the given task” [i.e., the trained parameters form a library which is utilized by the convolutional neural network/neural network application to define or determine the optimal logarithmic convolutional mapping/optimal point for applying the frequency domain features to multi-scaled frequency domain features (frequency scaling)] and 28, “However, important information may be lost during the mapping process and a hardcoded fixed-scale mapping may not provide an optimal mapping of frequency domain features for a given task. Therefore, the accuracy and performance of an audio classification system receiving the mapped frequency domain features may be reduced” [i.e., when an optimal mapping is not provided, then the performance of an audio classification system is reduced, as such, the optimal mapping inherently provides a non-reduced or non-degraded/without degrading performance of the classification system]).
Although Roblek substantially discloses the claimed invention, Roblek is not relied on to explicitly disclose a graphics processor to: … determine … a point to apply frequency scaling in the graphics processor.
In the same field, analogous art Dettmers teaches a graphics processor to: … determine … a point to apply frequency scaling in the graphics processor … the point determined based on analysis of the patterns generated by the distributed gradient synchronization (see, e.g., pages 2, 6 and 7-8, sec. 2.1, 3.4 and 4.1, “In data parallelism, the model is kept constant for all GPUs while each GPU is fed with a different mini-batch. After each pass the gradients are exchanged, i.e. synchronized with each GPU … Scaling limitations: Current GPU implementations are optimized for larger matrices, hence data parallelism does not scale indefinitely”, “this scheme yields good scaling” [i.e., a GPU/graphics processor used for a distributed gradient synchronization for frequency scaling in the GPU] and “In our work we show that we can use 8-bit gradients for the parameter updates without degrading performance. However, dynamic fixed point data types can also be used for end-to-end training and as such a combination of both methods might yield optimal performance … Although our 8-bit data type with dynamic binary tree achieves better approximation, it cannot be used in fixed point computation and thus remains useful solely as an intermediate approximate representation” [i.e., the 8-bit gradients are used for updates which keep the most optimal performance (optimal point determined based on analysis of patterns generated by the distributed gradient synchronization]).
Roblek and Dettmers are analogous art because they are both directed to gradient synchronizations within a neural network.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Roblek to incorporate the teachings of Dettmers in order to modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek to incorporate the gradient synchronization for an optimal performance through the updating of parameters within a GPU of Dettmers.
Doing so would enable Roblek to use “8-bit approximation [that] is able to circumnavigate problems with large batch sizes for GPU clusters and thus improves convergence rates in convolutional networks”, as suggested by Dettmers (See, e.g., Dettmers, page 2, 3rd bullet point).
Although Roblek in view of Dettmers substantially teaches the claimed invention, Roblek in view of Dettmers is not relied on to teach a library comprising machine learning primitives, wherein the machine learning primitives are usable to analyze patterns observed in a distributed gradient synchronization implemented by the neural network application and using the machine learning primitives of the library.
In the same field, analogous art Tokui teaches the library comprising machine learning primitives1 (see, e.g., paragraphs 67-68, “method for using calculation libraries such as Caffe … and Theano (http://deeplearning.net/software/theano/). … According to these libraries, by using a dedicated Mini programming language to describe the loss function as a combination of prepared primitives” [i.e., libraries include machine learning primitives]),
wherein the machine learning primitives are usable to analyze patterns observed in a distributed gradient synchronization implemented by the neural network application and using the machine learning primitives of the library (see, e.g., paragraph 68, “According to these libraries, by using a dedicated Mini programming language … as a combination of prepared primitives, it is possible to automatically obtain a gradient function of the loss function, too. This is because a gradient of each primitive is defined, and therefore a gradient of the entire combination can be also obtained by automatic differentiation. … by using this Mini programming language, the neural network can perform learning by the gradient method by using a gradient function” [i.e., the primitives are useable and used to analyze patterns observed in a distributed gradient method/function/synchronization implemented/performed by the neural network application/function]).
Roblek, Dettmers and Tokui are analogous art because they are directed to gradient synchronizations in a neural network and performing machine learning “by the gradient method by using a gradient function” within a neural network (See, e.g., Tokui, paragraph 68).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Roblek in view of Dettmers to incorporate the teachings of Tokui in order to provide “a dedicated Mini programming language to describe the loss function as a combination of prepared primitives” where “a gradient of each primitive is defined” (See, e.g., Tokui, paragraph 68). Doing so would have allowed Roblek in view of Dettmers to “automatically obtain a gradient function of the loss function” by using automatic differentiation to obtain “a gradient of the entire combination”, as suggested by Tokui (See, e.g., Tokui, paragraph 68).
Although Roblek in view of Dettmers and Tokui substantially teaches the claimed invention, Roblek in view of Dettmers and Tokui is not relied on to teach analyze patterns observed in a distributed gradient synchronization implemented by the neural network application; and
determine, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency.
In the same field, analogous art Lambert teaches analyze patterns observed in a distributed gradient synchronization implemented by the neural network application (see, e.g., FIG. 4 and page 63, right col., paragraph 3, “We have introduced this rule to ensure the AFNN retains a spread of frequencies (and thus metrical structure) across the gradient. The force is relative to natural frequency, and can be scaled through the ϵh parameter. By balancing the adaptive (ϵf) and elastic (ϵh) parameters, the oscillator frequency is able to entrain to a greater range of frequencies, whilst also returning to its natural frequency (ω0) when the stimulus is removed. Figure 4 shows the frequencies adapting over time in the AFNN under sinusoidal input” [i.e., the gradients as part of the Adaptive Frequency Neural Network/AFNN correspond to the patterns associated with/observed in a gradient synchronization implemented by the AFNN/neural network application]); and
determine, using the library, a core frequency of the frequency scaling applied at the point, wherein the library is to account for characteristics associated with the patterns generated by the distributed gradient synchronization to decide the core frequency (see, e.g., page 63, right col., Par. 3, “We have introduced this rule to ensure the AFNN retains a spread of frequencies (and thus metrical structure) across the gradient. The force is relative to natural frequency, and can be scaled through the ϵh parameter. By balancing the adaptive (ϵf) and elastic (ϵh) parameters, the oscillator frequency is able to entrain to a greater range of frequencies, whilst also returning to its natural frequency (ω0) when the stimulus is removed. Figure 4 shows the frequencies adapting over time in the AFNN under sinusoidal input” [i.e., the gradients as part of the Adaptive Frequency Neural network corresponds to the characteristics associated with the gradient synchronization, which balances to the natural frequency (core frequency)]).
Roblek, Dettmers, Tokui and Lambert are analogous art because they are directed to gradient synchronizations within a neural network and performing machine learning “by the gradient method by using a gradient function” within a neural network (See, e.g., Tokui, paragraph 68).
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek in view of Dettmers and Tokui to incorporate the balancing to the natural frequency by the adaptive frequency neural network of Lambert.
Doing so would have “significantly improve[d] responses of AFNNs compared to GFNNs to stimuli with both steady and varying pulse frequencies”, which “leads us to believe that AFNNs could replace the linear filtering methods commonly used in beat tracking and tempo estimation systems, and lead to more accurate methods”, as suggested by Lambert (See, e.g., Lambert, page 60, Abstract).
With respect to independent claim 28, claim 28 is substantially similar to claim 21 and therefore is rejected on the same ground as claim 21, discussed above. In particular, claim 28 is a method claim with steps that corresponds to the operations performed by the apparatus of claim 21.
In addition, Roblek further discloses a method (see, e.g., paragraph 42, “This specification describes methods for learning variable size convolutions on a linear spectrogram”).
With respect to independent claim 35, this claim is substantially similar to claim 21 and therefore is rejected on the same ground as claim 21, discussed above. Claim 35 is a machine-readable medium claim that performs operations that correspond operations performed by the apparatus of claim 21.
In addition, Roblek further discloses A non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations (see, e.g., paragraph 103, “the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, … The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them”).
Regarding claims 22, 29 and 36, as discussed above, Roblek in view of Dettmers, Tokui and Lambert teaches the apparatus of claim 21, the method of claim 28, and the machine-readable medium of claim 35.
Regarding claim 22, although Roblek substantially discloses the claimed invention, Roblek is not relied on to explicitly disclose wherein the point is determined through the distributed gradient synchronization using a tree-like structure such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a root of the tree-like structure.
In the same field, analogous art Dettmers teaches wherein the point is determined through the distributed gradient synchronization using a tree-like structure (see, e.g., pages 2 and 7-8, sec. 2.1 and 4.1, “In data parallelism, the model is kept constant for all GPUs while each GPU is fed with a different mini-batch. After each pass the gradients are exchanged, i.e. synchronized with each GPU” [i.e., gradient synchronization with a GPU (graphics processor)], “In our work we show that we can use 8-bit gradients for the parameter updates without degrading performance. However, dynamic fixed point data types can also be used for end-to-end training and as such a combination of both methods might yield optimal performance … Although our 8-bit data type with dynamic binary tree achieves better approximation, it cannot be used in fixed point computation and thus remains useful solely as an intermediate approximate representation” [i.e., the 8-bit gradients are used for updates which keep the most optimal performance (optimal point) wherein the 8-bit gradients are within a binary tree (a tree structure)]) such that local weight vectors start at one or more nodes represented as leaves of the tree-like structure and communicate up to a root of the tree-like structure (see, e.g., pages 4 and 7 - sec. 3.1 and 4.1, and 14 - paragraph 1, “In order to decrease this error, we can use the bits of the mantissa to represent a binary tree with interval (0.1, 1) which is bisected according to the route taken through the tree; the children thus represent the start and end points for intervals in a bisection method. With this method we can cover a broader range of numbers with the mantissa and can thus reduce the average relative error” [i.e., the mantissa includes bits that present a binary tree (tree structure) with an interval to take route through the tree (including leaves and roots communicated throughout the binary tree)], “Dynamic fixed point data types are data types which use all their bits for the mantissa and have a dynamic exponent which is kept for collection of numbers (matrix, vector) and is adjusted during run-time … In our work we show that we can use 8-bit gradients for the parameter updates without degrading performance. However, dynamic fixed point data types can also be used for end-to-end training and as such a combination of both methods might yield optimal performance” [i.e., the end-to-end train or run-time corresponds to the route taken through the tree wherein the vectors are part of the tree as a dynamic exponent of the bits in the mantissa which are considered to be the weights of the bits within the nodes], “We do this messaging scheme twice: Once to distribute the raw gradients, twice so we distribute the accumulated gradient to all nodes. The total time for this gradient synchronization scheme is about 1.9ms. This shows, as in the 4 GPU case, that there is no bottleneck in the data parallelism part of convolutional layers and thus 8-bit or 1-bit quantization will not improve performance in these layers” [i.e., the gradient is distributed to all nodes of the tree during gradient synchronization]).
Roblek and Dettmers are analogous art because they are directed to gradient synchronizations within a neural network.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Roblek to incorporate the teachings of Dettmers in order to modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek in view of Dettmers to incorporate the gradient synchronization while the binary tree is traversed through by routes wherein bits are given weights by exponents (vectors) of Dettmers.
Doing so would enable Roblek to use “8-bit approximation [that] is able to circumnavigate problems with large batch sizes for GPU clusters and thus improves convergence rates in convolutional networks”, as suggested by Dettmers (See, e.g., Dettmers, page 2, 3rd bullet point).
Regarding claim 29, claim 29 is substantially similar to claim 22 and therefore is rejected on the same ground as claim 22, discussed above. In particular, claim 29 is a method claim that corresponds to the apparatus of claim 22.
Regarding claim 36, this claim is substantially similar to claim 22 and therefore is rejected on the same ground as claim 22. In particular, claim 36 is a machine-readable medium claim that corresponds to the apparatus of claim 22.
Claims 23, 30 and 37 are rejected under 35 U.S.C. 103 as being unpatentable over Roblek in view of Dettmers, Tokui and Lambert as applied to claims 21, 28 and 35 above, and further in view of Chen et al. (US 2016/0293167, cited in applicants’ information disclosure statement filed 7/09/2022, hereinafter “Chen”).
Regarding claims 23, 30 and 37, as discussed above, as discussed above, Roblek in view of Dettmers, Tokui and Lambert teaches the apparatus of claim 21, the method of claim 28, and the machine-readable medium of claim 35.
Regarding claim 23, although Roblek substantially discloses the claimed invention, Roblek is not relied on to explicitly disclose wherein the graphics processor is further to introduce sparse matrix representation.
In the same field, analogous art Dettmers teaches wherein the graphics processor is further to introduce sparse matrix representation (see, e.g., page 2 sec. 2.1, “Scaling limitations: Current GPU implementations are optimized for larger matrices, hence data parallelism does not scale indefinitely due to slow matrix operations (especially matrix multiplication) for small mini-batch sizes (< 128 per GPU)” [i.e., the GPU implementations are optimized for larger matrices or sparse matrices]).
Roblek and Dettmers are analogous art because they are directed to gradient synchronizations within a neural network.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Roblek to incorporate the teachings of Dettmers in order to modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek to incorporate the GPU implementations optimized for larger matrices of Dettmers.
Doing so would enable Roblek to use “8-bit approximation [that] is able to circumnavigate problems with large batch sizes for GPU clusters and thus improves convergence rates in convolutional networks”, as suggested by Dettmers (See, e.g., Dettmers, page 2, 3rd bullet point).
Although Roblek in view of Dettmers, Tokui and Lambert substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui and Lambert is not relied on to teach wherein the … sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs.
In the same field, analogous art Chen teaches wherein the … sparse matrix representation for weights to overlap communication and computation across multiple nodes associated with the neural network application to reduce communication costs (see, e.g., Fig. 3 and paragraph 68, “Here k denotes the number of nodes of the rest of the hidden layers in the network. Note by comparing (2) and (3) that the variables flcn and n offer finer control over the number of parameters in the network. The first two hidden layers are influenced by flcn while remaining hidden layers have k2 weights. One interpretation of local connections is that they enforce patch-based sparse matrices when training; given the sparse filters in the first fully-connected hidden layer, e.g., as illustrated in FIG. 3, local connections are a natural fit” [i.e., the sparse matrices in each layer of the neural network depicted in Fig. 3 corresponds to the sparse matrix representation including non-zero weights which are patch-based connections of the multiple nodes (each box of a layer) of the convolution neural network]; see, e.g., paragraph 64, “This is important because parallel SIMD operations may be heavily relied upon in implementations of the techniques described herein to efficiently compute neural nets using small dense matrices rather than large, and sparse matrices. In some examples, LCN and CNN layers may be leveraged to take advantage of the sparse and local nature of the DNN to constrain the model size while improving performance” [i.e., using small dense matrices and sparse matrices improves the performance of the CNN and the LCN (as part of the DNN) which decreases or reduces costs]; see, e.g., paragraph 43, “An experimental system for implementing the techniques described herein used a fully-connected Deep Neural Network ("DNN") to extract a speaker-discriminative feature, or "d-vector", from each utterance. Utterance d-vectors were incrementally computed frame by frame, and improved latency by avoiding the computational costs associated with the latent variables of a factor analysis model, which occurred after utterance completion” [i.e., the techniques (sparse matrices) used on the DNN improves latency by avoiding computational costs (communication costs)]).
Roblek, Dettmers, Tokui, Lambert and Chen are analogous art because they are directed to matrix representations of nodes or parameters of convolutional neural networks.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek in view of Dettmers, Tokui and Lambert to incorporate the overlapping and product of parameters of the convolutional neural network through sparse matrices of Chen.
Doing so would enable Roblek in view of Dettmers, Tokui and Lambert to “reduce the total model footprint, for example, to 30% of the original size compared to a baseline fully-connected DNN, generally with reduced latency and minimal impact in performance” and “improved latency by avoiding the computational costs”, as suggested by Chen (See, e.g., Chen, paragraphs 4 and 44).
Regarding claim 30, claim 30 is substantially similar to claim 23 and therefore is rejected on the same ground as claim 23, discussed above. In particular, claim 30 is a method claim that corresponds to the apparatus of claim 23.
Regarding claim 37, this claim is substantially similar to claim 23 and therefore is rejected on the same ground as claim 23. In particular, claim 37 is a machine-readable medium claim that corresponds to the apparatus of claim 23.
Claims 24, 31 and 38 are rejected under 35 U.S.C. 103 as being unpatentable over Roblek in view of Dettmers, Tokui and Lambert as applied to claims 21, 28 and 35 above, and further in view of Britt et al. (U.S. Patent Application Pub. No. 2008/0092181, cited in applicants’ information disclosure statement filed 7/09/2022, hereinafter “Britt”).
Regarding claim 24, as discussed above, Roblek in view of Dettmers, Tokui and Lambert teaches the apparatus of claim 21.
Although Roblek substantially discloses the claimed invention, Roblek is not relied on to explicitly disclose wherein the graphics processor is further to …
In the same field, analogous art Dettmers teaches wherein the graphics processor is further to (see, e.g., page 2 sec. 2.1, “Scaling limitations: Current GPU implementations are optimized for larger matrices, hence data parallelism does not scale indefinitely due to slow matrix operations (especially matrix multiplication) for small mini-batch sizes (< 128 per GPU)” [i.e., the GPU implementations are optimized for larger matrices or sparse matrices]).
Roblek and Dettmers are analogous art because they are directed to gradient synchronizations within a neural network.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Roblek to incorporate the teachings of Dettmers in order to modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek to incorporate the gradient synchronization for an optimal performance through the updating of parameters within a GPU of Dettmers.
Doing so would enable Roblek to use “8-bit approximation [that] is able to circumnavigate problems with large batch sizes for GPU clusters and thus improves convergence rates in convolutional networks”, as suggested by Dettmers (See, e.g., Dettmers, page 2, 3rd bullet point).
Although Roblek in view of Dettmers, Tokui and Lambert substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui and Lambert is not relied on to teach wherein the … processor is further to automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters.
In the same field, analogous art Britt teaches wherein the … processor (see, e.g., paragraph 36, “In one embodiment, the apparatus comprises: a processor; a storage device in data communication with the processor”) is further to automatically analyze failed execution of programs including or relevant to the neural network application to obtain insights on one or more faults of hardware performance counters (see, e.g., paragraph 136, “In one embodiment, the storage devices 204 each comprise a redundant array (e.g., RAID) device, and when coupled with the fault tolerance, self monitoring, self-healing, and automatic communication channel fail-over (in the event of a hardware or software failure or loss of channel) of the illustrated architecture 200, provide a highly redundant and reliable configuration” [i.e., redundant array device corresponds to the debugging logic which is coupled with the automatic communication channel fail-over which detects for failed executions (failure or loss of channel) within the hardware for performance]; see, e.g., paragraph 98, “As used herein, the term " speech recognition" refers to any methodology or technique by which human or other speech can be interpreted and converted to an electronic or data format or signals related thereto … Phoneme/word recognition, if used, may be based on HMM (hidden Markov modeling), although other processes such as, without limitation, DTW (Dynamic Time Warping) or NNs (Neural Networks) may be used” [i.e., the program associated with the neural network application is the speech recognition system]).
Roblek, Dettmers, Tokui, Lambert and Britt are analogous art because they are directed to the analysis of speech (see, e.g., Tokui, paragraphs 119 and 131) and audio using neural networks.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek in view of Dettmers, Tokui and Lambert to incorporate the detection of failure of programs within a hardware of Britt.
Doing so would enable of Roblek in view of Dettmers, Tokui and Lambert to “provide a highly redundant and reliable configuration”, as suggested by Britt (See, e.g., Britt, paragraph 136]).
Regarding claim 31, this claim is substantially similar to claim 24 and therefore is rejected on the same ground as claim 24, discussed above. In particular, claim 31 is a method claim that corresponds to the apparatus of claim 24.
Regarding claim 38, claim 38 is substantially similar to claim 24 and therefore is rejected on the same ground as claim 24, discussed above. In particular, claim 38 is a machine-readable medium claim that corresponds to the apparatus of claim 24.
Claims 25, 32 and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Roblek in view Dettmers, Tokui, Lambert and Britt as applied to claims 24, 31, and 38 above, and further in view of non-patent literature Wong et al. (Wong, W. Eric, et al. "Effective software fault localization using an RBF neural network." IEEE Transactions on Reliability 61.1 (March 2012): 149-169, cited in applicants’ information disclosure statement filed 7/09/2022, hereinafter “Wong”).
Regarding claim 25, as discussed above, Roblek in view of Dettmers, Tokui, Lambert and Britt teaches the apparatus of claim 24.
Although Roblek substantially discloses the claimed invention, Roblek is not relied on to explicitly disclose wherein the graphics processor is further to provide one or more of successful execution information.
In the same field, analogous art Dettmers teaches wherein the graphics processor is further to provide one or more of successful execution information (see, e.g., page 2 sec. 2, “To understand the properties of a successful parallel deep learning algorithm, it is necessary to understand how the communication between GPUs works and what the bottlenecks for both model and data parallelized deep learning architectures are” [i.e., the successful parallel deep learning algorithms correspond to the successful execution information done by the GPU (graphic processor)].
Roblek and Dettmers are analogous art because they are both directed to gradient synchronizations within a neural network.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Roblek to incorporate the teachings of Dettmers in order to modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek to incorporate the gradient synchronization for an optimal performance through the updating of parameters within a GPU of Dettmers.
Doing so would enable Roblek to use “8-bit approximation [that] is able to circumnavigate problems with large batch sizes for GPU clusters and thus improves convergence rates in convolutional networks”, as suggested by Dettmers (See, e.g., Dettmers, page 2, 3rd bullet point).
Although Roblek in view of Dettmers, Tokui, Lambert and Britt substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui, Lambert and Britt is not relied on to teach wherein the … processor is further to provide one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval.
In the same field, analogous art Wong teaches wherein the … processor is further to provide one or more of successful execution information obtained from successful execution of the programs and failed execution information obtained from the failed execution of the programs to a trained network model (see, e.g., page 150, left col., paragraph 2, “A typical RBF neural network has a three-layer feed-forward structure that can be trained to learn an input-output relationship based on a data set. In this paper, the input is the statement coverage of a test case which indicates how the program is executed by the test case, and the output is the result (success or failure) of the corresponding program execution. Once the network has been trained, the coverage of a virtual test case with only one statement covered1 is used as an input to compute the suspiciousness of the corresponding statement in terms of its likelihood of containing bugs” [i.e., the RBF neural network is used to learn the input-output relationship in which a test case can indicate the successful or failed execution of a program and use that information as input to the trained neural network]) model to seek out one or more of the hardware performance counters that are regarded as faulty or outside a range of approval (see, e.g., page 157, right col. last paragraph, “For a fair comparison, we compute the effectiveness of both techniques (RBF, and Crosstab) using the same data. Note that statistics such as fault revealing behavior and statement coverage of each test can vary under different compilers, operating systems, and hardware platforms” [i.e., the RBF neural network is used to seek faults within a system or hardware platform]).
Roblek, Dettmers, Tokui, Lambert, Britt, and Wong are analogous art because they are directed to the optimization or yielding of a best performance of neural network.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature with failure detection of programs of Roblek in view of Dettmers, Tokui and Lambert and further in view of Britt to incorporate the output of success or failure of the execution of programs used as input to the trained RBF neural network of Wong.
Doing so would enable Roblek in view of Dettmers, Tokui and Lambert and further in view of Britt to be “more effective at locating bugs, in that a relatively smaller amount of code needs to be examined to find bugs, compared to other state of the art contemporary techniques”, as suggested by Wong (See, e.g., Wong, page 149, Introduction, paragraph 1).
Regarding claim 32, this claim is substantially similar to claim 25 and therefore is rejected on the same ground as claim 25, discussed above. In particular, claim 32 is a method claim that corresponds to the apparatus of claim 25.
Regarding Claim 39, claim 39 is substantially similar to claim 25 and therefore is rejected on the same ground as claim 25, discussed above. In particular, claim 39 is a machine-readable medium claim that corresponds to the apparatus of claim 25.
Claims 26 and 33 are rejected under 35 U.S.C. 103 as being unpatentable over Roblek in view of Dettmers, Tokui and Lambert as applied to claims 1 and 8 above, in view of Sugiura et al. (U.S. Patent Application Pub. No. 2019/0095757 A1, cited in applicants’ information disclosure statement filed 7/09/2022 hereinafter “Sugiura”) and further in view of non-patent literature Anwar et al. ("Fixed point optimization of deep convolutional neural networks for object recognition." 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2015, cited in applicants’ information disclosure statement filed 7/09/2022, hereinafter “Anwar”). Sugiura was filed on October 31, 2018 as a national stage application of PCT application no. PCT/JP2017/044295 filed December 11, 2017, and this date is before the effective filing date of the present application, April 28, 2017. Therefore, Sugiura constitutes prior art under 35 U.S.C. 102(a)(2).
Regarding claims 26 and 33, as discussed above, Roblek in view of Dettmers, Tokui and Lambert teaches the apparatus of claim 21 and the method of claim 28.
Regarding claim 26, although Roblek substantially discloses the claimed invention, Roblek is not relied on to explicitly disclose wherein the graphics processor is further to perform local error propagation.
In the same field, analogous art Dettmers teaches wherein the graphics processor is further to perform local error propagation (see, e.g., page 6 sec. 3.4, “Since we only had one GPU available for the following experiments, we simulated training on a large GPU cluster by only using the pure 8-bit approximation gradient component by training on a single GPU – so no 32-bit gradients or activations where used. On MNIST, we found that the best test error of all four approximation techniques static tree, dynamic tree, linear quantization, and mantissa did not differ significantly from the test error of 32-bit training for both data parallelism F(4, 4) = 0.71, p = 0.59, and model parallelism F(4, 4) = 0.54, p = 0.71 (F-test assumptions were satisfied); also the 99% confidence intervals did overlap for all techniques” [i.e., the GPU corresponds to the graphics processor that is used to test error of the approximation techniques (local error propagation of the binary tree)]).
Roblek and Dettmers are analogous art because they are directed to gradient synchronizations within a neural network.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature using matrix representation for parameters of the neural network of Roblek in view of Dettmers to incorporate the error testing of the bits in the binary tree of Dettmers.
Doing so would enable Roblek to use “8-bit approximation [that] is able to circumnavigate problems with large batch sizes for GPU clusters and thus improves convergence rates in convolutional networks”, as suggested by Dettmers (See, e.g., Dettmers, page 2, 3rd bullet point).
Although Roblek in view of Dettmers, Tokui and Lambert substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui and Lambert is not relied on to teach further … to perform local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication.
In the same field, analogous art Sugiura teaches the apparatus further … to perform local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application (see, e.g., paragraph 49, “In the learning processing of deep learning, a weight value of the coupling (synaptic coupling) between nodes configuring the neural network is updated using a known algorithm (for example, in a reverse error propagation method, adjust and update the weight value so as to reduce the error from the correct at the output layer, or the like). An aggregate of the weight values between the nodes on which the learning process is completed is called a "learned model". By applying the learned model to a neural network having the same configuration as the neural network used in the learning process (setting as the weight value of inter -node coupling), it is possible to output correct data with a constant precision as output data (recognition result) when inputting unknown input data, i.e., new input data not used in learning processing, into the neural network” [i.e., the weight value between nodes corresponds to the local weights being updated using error propagation to reduce error at the output layer of the neural network which provides constant precision with inputting unknown data]), wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication (see, e.g., paragraph 49, “In the learning processing of deep learning, a weight value of the coupling (synaptic coupling) between nodes configuring the neural network is updated using a known algorithm (for example, in a reverse error propagation method, adjust and update the weight value so as to reduce the error from the correct at the output layer, or the like). An aggregate of the weight values between the nodes on which the learning process is completed is called a "learned model". By applying the learned model to a neural network having the same configuration as the neural network used in the learning process (setting as the weight value of inter -node coupling), it is possible to output correct data with a constant precision as output data (recognition result) when inputting unknown input data, i.e., new input data not used in learning processing, into the neural network” [i.e., the aggregating of weight values corresponds to the weight synchronization across the nodes of the neural network to reduce the error (reduced communication) by tracking and outputting correct data (accuracy)]).
Roblek, Dettmers, Tokui, Lambert and Sugiura are analogous art because they are directed to using techniques of back propagation on a neural network.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature with failure detection of programs of Roblek in view of Dettmers, Tokui and Lambert to incorporate the error propagation at multiple nodes and aggregating the weights of each node to reduce errors and output correct data of Sugiura.
Doing so would enable Roblek in view of Dettmers, Tokui and Lambert to “adjust and update the weight value so as to reduce the error from the correct at the output layer”, as suggested by Sugiura (See, see, e.g., Sugiura, paragraph 49).
Although Roblek in view of Dettmers, Tokui, Lambert and Sugiura substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui and Lambert and Sugiura is not relied on to teach perform local error propagation by computing high precision and low precision for local weights.
In the same field, analogous art Anwar teaches perform local error propagation by computing high precision and low precision for local weights (see, e.g., page 1133, sec. 3, “During training we keep parameters in both high and low precision. We set aside 5000 training samples for validation … We start with a high precision pre trained network and obtain a quantized network using L2 error minimization. Then the inputs are fed forward via the network with the low precision weights … The output error is back propagated via low precision weights. The computed change in weights is added to the high precision weights. Thus we obtain new high precision weights. This process is iterated for several mini-batches and epochs. During training the selection of mini-batch size is important. Generally CNN employs the stochastic gradient descent (SGD) algorithm, where conventionally the minibatch size is one and weights are updated after each sample” [i.e., the output error of the CNN is back propagated by the use of low precision weights wherein the computed change in weights is added to high precision weights]).
Roblek, Dettmers, Tokui, Lambert, Sugiura, and Anwar are analogous art because they are directed to using techniques of back propagation in a neural network.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature with failure detection of programs and error propagation of Roblek in view of Dettmers, Tokui, Lambert and further in view of Sugiura to incorporate the error propagation of computing high and low precision weights of Anwar.
Doing so would enable Roblek in view of Dettmers, Tokui, Lambert and further in view of Sugiura to induce “sparsity in the network which reduces the effective number of network parameters and improves generalization” and “reduces the required memory storage by a factor of 1/10 and achieves better classification results than the high precision networks”, as suggested by Anwar (See, e.g., Anwar, Abstract).
Regarding claim 33, claim 33 is substantially similar to claim 26 and therefore is rejected on the same ground as claim 26, discussed above. In particular, claim 33 is a method claim that corresponds to the apparatus of claim 26.
Claims 27 and 34 are rejected under 35 U.S.C. 103 as being unpatentable over Roblek in view of Dettmers, Tokui and Lambert as applied to claims 1 and 8 above, and further in view of Ray et al. (U.S. Patent Application Pub. No. 2018/0293102, cited in applicants’ information disclosure statement filed 7/09/2022, hereinafter “Ray”).
Regarding claims 27 and 34, as discussed above, Roblek in view of Dettmers, Tokui and Lambert teach the apparatus of claim 21 and the method of claim 28.
Regarding claim 27, although Roblek in view of Dettmers, Tokui and Lambert substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui and Lambert is not relied on to teach wherein the apparatus comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
In the same field, analogous art Ray teaches wherein the apparatus comprises an autonomous machine including one or more of a vehicle, a device, and an equipment (see, e.g., paragraph 144, “Computing device 600 may further include (without limitations) an autonomous machine or an artificially intelligent agent, such as a mechanical agent or machine, an electronics agent or machine, a virtual agent or machine, an electro-mechanical agent or machine, etc.” [i.e., the computing device corresponds to the autonomous machine including a mechanical agent, electronics agent, electro-mechanical agent or machine – one or more of a device and an equipment]), wherein the autonomous machine comprises one or more processors including the graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package (see, e.g., paragraph 353, “Example 7 includes the subject matter of Examples 1-6, wherein the graphics processor is co-located with an application processor on a common semiconductor package”).
Roblek, Dettmers, Tokui, Lambert and Ray are analogous art because they are directed to the analysis of speech and audio using neural networks.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature of Roblek in view of Dettmers, Tokui, Lambert to incorporate the autonomous machine including a graphics processor coupled with a semiconductor of Ray.
Doing so would allow manufacturers of the systems of Roblek in view of Dettmers, Tokui, Lambert to “maximize the amount of parallel processing in the graphics pipeline” and “attempt to execute program instructions synchronously together as often as possible to increase processing efficiency” wherein the “efficiency provided by parallel machine learning algorithm implementations allows the use of high capacity networks and enables those networks to be trained on larger datasets”, as suggested by Ray (See, e.g., Ray, paragraph 4).
Regarding claim 34, this claim is substantially similar to claim 27 and therefore is rejected on the same ground as claim 27, discussed above. In particular, claim 34 is a method claim that corresponds to the apparatus of claim 27.
Claim 40 is rejected under 35 U.S.C. 103 as being unpatentable over Roblek in view Dettmers, Tokui and Lambert as applied to claim 35 above, in view of Sugiura in view of Anwar, and further in view of Ray.
Regarding claim 40, as discussed above, Roblek in view of Dettmers, Tokui and Lambert teach the machine-readable medium of claim 35.
Although Roblek in view of Dettmers, Tokui and Lambert substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui and Lambert is not relied on to teach wherein the operations further comprise performing local error propagation by computing high precision and low precision for local weights and compute local errors at each of multiple nodes associated with the neural network application, wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication, wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including a graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
In the same field, analogous art Sugiura teaches wherein the operations further comprise performing local error propagation by … compute local errors at each of multiple nodes (see, e.g., paragraph 49, “In the learning processing of deep learning, a weight value of the coupling (synaptic coupling) between nodes configuring the neural network is updated using a known algorithm (for example, in a reverse error propagation method, adjust and update the weight value so as to reduce the error from the correct at the output layer, or the like). An aggregate of the weight values between the nodes on which the learning process is completed is called a "learned model". By applying the learned model to a neural network having the same configuration as the neural network used in the learning process (setting as the weight value of inter -node coupling), it is possible to output correct data with a constant precision as output data (recognition result) when inputting unknown input data, i.e., new input data not used in learning processing, into the neural network” [i.e., the weight value between nodes corresponds to the local weights being updated using error propagation to reduce error at the output layer of the neural network which provides constant precision with inputting unknown data]), wherein performing the local error propagation includes facilitating weight synchronization across the multiple nodes to track the local errors for accuracy and reduced communication (see, e.g., paragraph 49, “In the learning processing of deep learning, a weight value of the coupling (synaptic coupling) between nodes configuring the neural network is updated using a known algorithm (for example, in a reverse error propagation method, adjust and update the weight value so as to reduce the error from the correct at the output layer, or the like). An aggregate of the weight values between the nodes on which the learning process is completed is called a "learned model". By applying the learned model to a neural network having the same configuration as the neural network used in the learning process (setting as the weight value of inter -node coupling), it is possible to output correct data with a constant precision as output data (recognition result) when inputting unknown input data, i.e., new input data not used in learning processing, into the neural network” [i.e., the aggregating of weight values corresponds to the weight synchronization across the nodes of the neural network to reduce the error (reduced communication) by tracking and outputting correct data (accuracy)]).
Roblek, Dettmers Tokui, Lambert, and Sugiura are analogous art because they are each directed to using techniques of back propagation on a neural network.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the machine-readable medium of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature with failure detection of programs of Roblek in view of Dettmers, Tokui and Lambert to incorporate the error propagation at multiple nodes and aggregating the weights of each node to reduce errors and output correct data of Sugiura.
Doing so would enable Roblek in view of Dettmers, Tokui and Lambert to “adjust and update the weight value so as to reduce the error from the correct at the output layer” (See, e.g., Sugiura, paragraph 49).
Although Roblek in view of Dettmers, Tokui and Sugiura substantially teaches the perform local error propagation by computing high precision and low precision for local weights.
In the same field, analogous art Anwar teaches perform local error propagation by computing high precision and low precision for local weights (see, e.g., page 1133, sec. 3, “During training we keep parameters in both high and low precision. We set aside 5000 training samples for validation…We start with a high precision pre trained network and obtain a quantized network using L2 error minimization. Then the inputs are fed forward via the network with the low precision weights …The output error is back propagated via low precision weights. The computed change in weights is added to the high precision weights. Thus we obtain new high precision weights. This process is iterated for several mini-batches and epochs. During training the selection of mini-batch size is important. Generally CNN employs the stochastic gradient descent (SGD) algorithm, where conventionally the minibatch size is one and weights are updated after each sample” [i.e., the output error of the CNN is back propagated by the use of low precision weights wherein the computed change in weights is added to high precision weights]).
Roblek, Dettmers, Tokui, Lambert, Sugiura and Anwar are analogous art because they are each directed to using techniques of back propagation on a neural network.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature with failure detection of programs and error propagation of Roblek, Dettmers, Tokui and Lambert, and further in view of Sugiura to incorporate the error propagation of computing high and low precision weights of Anwar.
Doing so would enable Roblek, Dettmers, Tokui and Lambert, and further in view of Sugiura to induce “sparsity in the network which reduces the effective number of network parameters and improves generalization” and “reduces the required memory storage by a factor of 1/10 and achieves better classification results than the high precision networks”, as suggested by Anwar (See, e.g., Anwar, Abstract).
Although Roblek in view of Dettmers, Tokui, Sugiura and Anwar substantially teaches the claimed invention, Roblek in view of Dettmers, Tokui, Sugiura and Anwar is not relied on to teach wherein the computing device comprises an autonomous machine including one or more of a vehicle, a device, and an equipment, wherein the autonomous machine comprises one or more processors including a graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package.
In the same field, analogous art Ray teaches wherein the computing device comprises an autonomous machine comprising one or more of a vehicle, a device, and an equipment (see, e.g., paragraph 144, “Computing device 600 may further include (without limitations) an autonomous machine or an artificially intelligent agent, such as a mechanical agent or machine, an electronics agent or machine, a virtual agent or machine, an electro-mechanical agent or machine, etc.” [i.e., the computing device corresponds to the autonomous machine including a mechanical agent, electronics agent, electro-mechanical agent or machine - one or more of a device and an equipment]), wherein the autonomous machine comprises one or more processors including a graphics processor, wherein the graphics processor is co-located with an application processor on a common semiconductor package (see, e.g., paragraph 353, “Example 7 includes the subject matter of Examples 1-6, wherein the graphics processor is co-located with an application processor on a common semiconductor package”).
Roblek, Dettmers, Tokui, Lambert, Sugiura, Anwar, and Ray are analogous art because they are each directed to the analysis of speech (see, e.g., Tokui, paragraphs 119 and 131) and audio using neural networks.
It would have been obvious to a person having ordinary skill in the art before the effective filing date modify the apparatus of a convolutional neural network system which uses an optimal mapping for a multi-scaled frequency domain feature of Roblek in view of Dettmers, Tokui, Lambert, Sugiura and further in view Anwar to incorporate the autonomous machine including a graphics processor coupled with a semiconductor of Ray.
Doing so would allow manufacturers of a systems such as those of Roblek in view of Dettmers, Tokui, Lambert, Sugiura and further in view Anwar to “maximize the amount of parallel processing in the graphics pipeline” and “attempt to execute program instructions synchronously together as often as possible to increase processing efficiency” wherein the “efficiency provided by parallel machine learning algorithm implementations allows the use of high capacity networks and enables those networks to be trained on larger datasets”, as suggested by Ray (See, e.g., Ray, paragraph 4).
Conclusion
The prior art made of record, listed on form PTO-892, and not relied upon, is considered pertinent to applicant's disclosure.
The examiner requests, in response to this office action, support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line no(s) in the specification and/or drawing figure(s). This will assist the examiner in prosecuting the application.
When responding to this office action, Applicant is advised to clearly point out the patentable novelty which he or she thinks the claims present, in view of the state of the art disclosed by the reference cited or the objections made. He or she must also show how the amendments avoid such references or objections See 37 CFR 1.111 (c).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RANDY K BALDWIN whose telephone number is (571)270-5222. The examiner can normally be reached on Mon - Fri 9:00-6:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RANDALL K. BALDWIN/Primary Examiner, Art Unit 2125
1 Paragraph 199 of applicant’s specification states “Machine learning primitives are basic operations that are commonly performed by machine learning algorithms.” Therefore, a “library comprising machine learning primitives”, under the broadest reasonable interpretation (BRI), is a library that includes operations, such as programs or functions, that can be performed by machine learning algorithms)