Detailed Action
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
2. Claims 1-20 are pending.
Claim Analysis - 35 USC § 101
3. The invention as recited in independent claims 1, 15, and 20 pertains to distributing (pipelined) layers of a machine learning model among a plurality of individual devices connected via a computer network, determining which layers interconnect between the plurality of individual devices, and then reducing the parameters in these interconnection layers to affect the communication between the plurality of individual devices while the machine learning model is executing. The Specification makes clear that the improvement at least includes reduced latency in the network, optimized communication between the devices, and thus improved efficiency of the distributed model. Therefore, even if some aspects of the claim limitations were somehow to be construed as a judicial exception, nevertheless it would be well integrated into the network system, and the claimed invention would amount to significantly more than a judicial exception. Thus a 101 rejection would not be applicable.
Claim Rejections - 35 USC § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
5. Claim(s) 1-3, 5-6, 9-11, 15-16, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Banitalebi et al “Banitalebi” (WO 2021174370 A1) and Shi et al “Shi” (Improving Device Edge Cooperative Inference of Deep Learning via 2-Step Pruning).
(Please see the attached copies of Banitalebi and Shi that number paragraphs and pages in the same manner as that used in this Action).
6. Regarding claim 1, Banitalebi shows a method, comprising: distributing, by a controller device, a plurality of pipelined layers of a machine learning model among a plurality of individual devices connected via a computer network (Figure 1, para 10-13, 37-38 show the computing system distributing the interconnected data processing and modelling layers of the neural network model among a plurality of devices connected over a network, para 111-116 show the partitioning of the fully trained neural network that can be executed on different computing platform devices); determining, by the controller device, a subset of the plurality of pipelined layers that are fringe layers that interconnect between the plurality of individual devices via the computer network (para 43-44, 77, 116-119 show determining which of the distributed layers interconnect between the specific devices within the network. This is what the claimed invention defines as “fringe” and so these layers are thus fringe layers); and executing, by the controller device, the machine learning model with an input passed through the plurality of pipelined layers to produce an output, wherein communication between the plurality of individual devices is based on the fringe layers (para 45, 64, 94, 114 show running the neural network model with an input passed through the distributed layers to produce an output, and communication between the plurality of devices is based on the interconnecting/fringe layers). Banitalebi does not show generating, by the controller device, compressed fringe layers by reducing parameters of the fringe layers, such that the communication between the plurality of individual devices is based on the compressed fringe layers. Shi however does show generating, by the controller device, compressed fringe layers by reducing parameters of the fringe layers such that the communication between the plurality of individual devices is based on the compressed fringe layers (page 4 column 2 para 2-3, page 5 column 2 para 3-4, page 6 column 1 para 1-2 show compressing the device interconnection layers by pruning the parameters in those layers). It would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to have generated compressed fringe layers, by reducing parameters of the fringe layers, in the distributed neural network model of Banitalebi such that the communication between the plurality of individual devices would be based on the compressed fringe layers, because it would provide an efficient way to reduce network traffic and thus latency between devices (Banitalebi para 114, Shi page 4 column 2 para 3).
7. Regarding claim 2, in addition to that mentioned for claim 1, Shi shows progressively compressing the compressed fringe layers through iterations of executing the machine learning model until a desired reduction in one or both of bandwidth or latency of the communication between the plurality of individual devices is achieved (Shi page 2 column 2 para 4-5, page 3 column 1 para 1-4 and column 2 para 1-2 show progressive pruning iterations of a series of generated neural network models. This is done to reach a particular accuracy as well as latency constraint). Motivation to combine Banitalebi with Shi is the same as that mentioned for claim 1, namely because it would provide an efficient way to reduce network traffic and thus latency between the plurality of devices.
8. Regarding claim 3, in addition to that mentioned for claim 1, generating the compressed fringe layers is based on performance of the computer network (Banitalebi para 39, 111-114 show partitioning and generating the layers that interconnect between devices/platforms based on achieving lower latency over the network. Shi page 2 column 1 para 1-2, page 3 column 2 para 2-5, page 4 column 2 para 2-3 show that partitioning the layers and compressing those interconnecting devices based on network latency – motivation to combine Banitalebi with Shi is the same as that mentioned for claim 1, namely because it would provide an efficient way to reduce network traffic and thus latency between the plurality of devices).
9. Regarding claim 5, in addition to that mentioned for claim 1, Shi shows reducing parameters of the fringe layers consequently reduces information transmission for the communication between the plurality of individual devices, thereby reducing one or both of bandwidth or latency of the communication between the plurality of individual devices (please note the alternative language – Shi page 4 column 2 para 2-3, page 5 column 2 para 3-4, page 6 column 1 para 1-2 show pruning the parameters in the device interconnection layers reduces information transmission and communication between the devices, which then reduces latency of communication between the plurality of individual devices, which page 3 column 2 para 2-5 show for example may be mobile devices and edge devices). Motivation to combine Banitalebi with Shi is the same as that mentioned for claim 1, namely because it would provide an efficient way to reduce network traffic and thus latency between the plurality of devices.
10. Regarding claim 6, in addition to that mentioned for claim 1, Shi further shows reducing the parameters of the fringe layers by progressively pruning the parameters of the fringe layers based on an accuracy of the machine learning model reaching a minimum accuracy threshold for the machine learning model (Shi page 2 column 1 para 1, page 3 column 2 para 2-5, page 5 column 2 para 1-3 show tuning and progressively pruning the parameters in the device interconnection layers based on the accuracy of the deep neural network model reaching a minimum accuracy constraint threshold). Note that Banitalebi para 48, 87, 91, 94, 114 also show distributing the neural network model among the plurality of devices in order achieve a threshold accuracy. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to progressively reduce/prune the parameters of the fringe layers based on an accuracy of the machine learning model reaching a minimum accuracy threshold for the machine learning model as is done in Shi, in the distributed neural network model of Banitalebi, because it would provide an efficient way to maintain a threshold accuracy constraint for the neural network model.
11. Regarding claim 9, in addition to that mentioned for claim 1, Banitalebi shows generating the fringe layers is based on automated system-based control (para 40 shows the model splitting and layer distribution of the neural network model is automatically processed by the system. Para 41 shows this includes the device interconnection layers and is based on network and communication constraint thresholds between the devices).
12. Regarding claim 10, in addition to that mentioned for claim 9, Banitalebi shows the automated system-based control is based on one or more constraints of the computer network (para 41 shows the controlling of the layer generation, including the device interconnection layers, is based on network and communication constraint thresholds between the devices. Para 40 shows this is automatically processed by the system).
13. Regarding claim 11, please note the alternative language. In addition to that mentioned for claim 1, note claim 1 shows generating the fringe layers as explained above. Shi shows compressing the fringe layers, to thus generate the compressed fringe layers, by reducing parameters of the fringe layers according to the control option of a minimum accuracy threshold for the machine learning model (Shi page 2 column 1 para 1, page 3 column 2 para 2-5, page 5 column 2 para 1-3 show pruning the parameters in the device interconnection layers based on the accuracy of the deep neural network model reaching a minimum accuracy constraint threshold). Note that Banitalebi para 48, 87, 91, 94, 114 also show distributing the neural network model among the plurality of devices in order achieve a threshold accuracy. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to reduce/prune the parameters of the fringe layers based on an accuracy of the machine learning model reaching a minimum accuracy threshold for the machine learning model as is done in Shi, in the distributed neural network model of Banitalebi, because it would provide an efficient way to maintain a threshold accuracy constraint for the neural network model while reducing latency between the plurality of devices.
14. Claims 15-16 and 18 show the same limitations as claims 1-2 and 6 respectively, and are rejected for the same reasons. In addition, note that Banitalebi para 98, 101 show the one or more network interfaces to communicate with a network, para 96-98 and 101 show a processor coupled to the one or more network interfaces and configured to execute one or more processes; and para 99 shows a memory configured to store a process that is executable by the processor.
15. Claim 20 shows the same limitations as claim 1 and is rejected for the same reasons. In addition, note that Banitalebi shows the tangible non-transitory computer readable storage medium storing program instructions that cause a device to execute a process (para 99 shows the non-transitory hardware memories storing program instructions that cause a processing device to execute a process).
16. Claim(s) 7 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Banitalebi and Shi and Wang et al “Wang” (WO 2023030513 A1).
(Please also see the attached copy of Wang that numbers paragraphs in the same format as that used in this Action).
17. Regarding claim 7, in addition to that mentioned for claim 1, Shi shows reducing the parameters of the fringe layers based on achieving a minimum accuracy threshold for the machine learning model according to a set value of one or both of bandwidth or latency of the communication between the plurality of individual devices (note the alternative recitation – Shi page 4 column 2 para 3, page 5 column 2 para 2-4, page 6 column 1 para 2 show reducing the parameters in the interconnection layers based on achieving a set threshold accuracy according to a bandwidth constraint or according to a latency constraint). Note that Banitalebi para 48, 87, 91, 94, 114 also show distributing the neural network model among the plurality of devices in order achieve a threshold accuracy. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to reduce the parameters of the fringe layers based on achieving a minimum accuracy threshold for the machine learning model according to a set value of one or both of bandwidth or latency of the communication between the plurality of individual devices as is done in Shi, in the distributed neural network model of Banitalebi, because it would provide an efficient way to maintain a threshold accuracy for the neural network model while reducing latency between the plurality of devices. Neither Banitalebi nor Shi explicitly show reducing the parameters using knowledge distillation per se to construct an updated layer architecture for the machine learning model. Wang however does use knowledge distillation to reduce parameters to construct an updated layer architecture for a machine learning model (Wang para 595-598 shows using knowledge distillation to reduce parameters to achieve multi-level compression and optimization of a machine learning model; this results in updated layer architecture for the machine learning model). Wang para 592, 597-599 show this is done to meet accuracy requirements while also optimizing speed (and thus reducing latency). Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to reduce the parameters using knowledge distillation to thus construct an updated layer architecture for the machine learning model as is done in Wang, in the distributed neural network model of Banitalebi especially as modified by Shi, because it would provide an efficient way to maintain a threshold accuracy for the neural network model while maintaining optimized speed/reducing latency between the plurality of devices.
18. Claim 19 shows the same limitations as claim 7 and is rejected for the same reasons.
19. Claim(s) 8, 13, and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Banitalebi and Shi and Kim et al “Kim” (US 2023/0252293 A1).
20. Regarding claim 8, in addition to that mentioned for claim 1, neither Banitalebi nor Shi explicitly show that generating the compressed fringe layers is based on receiving user-based control through a user interface, but Banitalebi para 46, 63, 72 do show user controlled constraints to modify the neural network model. Furthermore, Kim shows generating compressed layers for a neural network based on receiving user-based control through a user interface (Kim para 20, 58-61 show user inputted parameter values and block selections via a user interface to control and configure the compression of neural network blocks, and this generates compressed blocks for the neural network model. Para 80, 109 show these blocks may be layers). Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to generate the compressed layers based on receiving user-based control through a user interface as is done in Kim, in the distributed neural network model of Banitalebi especially as modified by Shi, to thus generate the compressed fringe layers, because it would provide an efficient way for a user to generate compressed fringe layers and modify the neural network model.
21. Regarding claim 13, in addition to that mentioned for claim 1, neither Banitalebi nor Shi explicitly show displaying a graphical user interface based on comparative performance of the communication between the plurality of individual devices according to the compressed fringe layers, but Shi page 2 column 1 para 2, page 3 column 2 para 2-5, Figures 4-5 show outputting comparative performance of the communication such as latency between the plurality of devices according to the interconnection/fringe layers as they are compressed/pruned. Furthermore, Kim shows displaying a graphical user interface based on comparative performance of the communication between the plurality of devices according to the compressed blocks of the neural network (Kim para 49, 54, 58, 152 show output comparing latency performance of communication between the devices among the device farm, Kim para 141, 148, 151, 160-162 show displaying output including the latencies on a user interface according to the compressed blocks between devices, and para 80, 109 show these blocks may be layers). Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to display a graphical user interface based on comparative performance of the communication between the plurality of devices according to the compressed blocks/layers of the neural network as is done in Kim, in the distributed neural network model of Banitalebi as modified by Shi, to thus display a graphical user interface based on comparative performance of the communication between the plurality of individual devices according to the compressed fringe layers, because it would provide an efficient way to output comparative performance of the communication/latency between the plurality of devices according to the fringe layers as they are compressed. Doing so would help a user to then input appropriate values and modifications to reduce latency between the devices and reduce network traffic.
22. Regarding claim 14, in addition to that mentioned for claim 13, Shi shows the comparative performance is based on progressive iterations of compression of the fringe layers and is selected from a group consisting of: connection-based bandwidth reduction; connection-based latency reduction; global bandwidth reduction; global latency reduction; and accuracy of the machine learning model (note the alternative recitation – Shi page 2 column 2 para 4-5, page 3 column 1 para 1-4 and column 2 para 1-4 show progressive pruning iterations of the layers including the interconnection layers. Page 3 column 2 para 2-5, Figures 4-5 show the connection based latency reduction is then compared as the layers are compressed). That this would be used to display a graphical user interface is obvious in view of Kim, as explained for claim 13.
23. Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Banitalebi and Shi and Gorokhov et al “Gorokhov” (WO 2019197855 A1).
(Please also see the attached copy of Gorokhov that numbers paragraphs in the same manner as that used in this Action).
24. Regarding claim 12, in addition to that mentioned for claim 1, neither Banitalebi nor Shi explicitly show the generating of the compressed fringe layers occurs at serving time per se. Gorokhov however does indeed generate compressed layers of a neural network at serving time (para 89-90, 96, 101, 139 show pruning/compressing layers of a neural network in real time (on the fly) which is as it is being served). It would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention to generate the compressed layers in real time as is done in Gorokhov, in the distributed neural network model of Banitalebi as modified by Shi, to thus generate the compressed fringe layers at serving time, because it would provide an efficient way to maintain accuracy while reducing latency in the network. Doing so would only reduce parameters in the fringe layers and prune the neural network while it is actually being trained at serving time (Gorokhov para 139).
Allowable Subject Matter
25. Claims 4 and 17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The prior art alone and/or in combination does not show preventing changes to parameters of all the plurality of pipelined layers that are not fringe layers, in combination with all of the limitations of the independent claims through dependency.
Conclusion
26. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
a) Arikawa (US 2023/0385603 A1) reduces parameters in a neural network to satisfy a performance constraint.
b) D’Ercoli (US 2020/0210834 A1) uses peer to peer routing between computational nodes in a partitioned neural network.
27. Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEVEN PAUL SAX whose telephone number is (571)272-4072. The examiner can normally be reached Monday - Friday, 9:30 - 6:00 Est.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed, can be reached at 571-272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/STEVEN P SAX/ Primary Examiner, Art Unit 2146