Prosecution Insights
Last updated: October 02, 2026
Application No. 17/723,216

COMPRESSION AS A SOLUTION FOR CONGESTION CONTROL ON AI WORKLOADS

Final Rejection §101§103§112
Filed
Apr 18, 2022
Examiner
WONG, WILLIAM
Art Unit
2144
Tech Center
2100 — Computer Architecture & Software
Assignee
Intel Corporation
OA Round
4 (Final)
31%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
58%
With Interview

Examiner Intelligence

Grants only 31% of cases
31%
Career Allowance Rate
125 granted / 407 resolved
-24.3% vs TC avg
Strong +28% interview lift
Without
With
+27.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
22 currently pending
Career history
439
Total Applications
across all art units

Statute-Specific Performance

§101
11.9%
-28.1% vs TC avg
§103
47.4%
+7.4% vs TC avg
§102
13.0%
-27.0% vs TC avg
§112
23.4%
-16.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 407 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to communications filed on 06/29/2026. Claims 1-25 are pending and have been examined. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-25 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 1 is amended to recite “determine, based upon (1) machine learning model accuracy change data, (2) network bandwidth utilization change data, (3) compression compute time data, and (4) the network telemetry data, whether to selectively compress Tensor data generated by the plurality of compute nodes… the machine learning model accuracy change data is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism: and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes”. However, the specification does not support the above features. The specification only generally describes “variable compression ratios are used to compress the gradient data based on hardware network events to prevent network congestion… variable compression techniques also consider cost/benefit tradeoffs… compressing data to reduce the network traffic involves a trade-off between machine learning model accuracy loss, network bandwidth utilization reduction and the compute time required for the compression” (e.g. in paragraphs 65-66), “both homogeneous and heterogeneous compute nodes may be used in a distributed system” (e.g. in paragraph 35) and “ML/AI models may be processed by distributed systems using data parallelism, model parallelism, or a combination of the two. A simple example of data parallelism is shown in Figure 2, which shows a distributed system 200 including a pair of compute nodes 202 and 204 connected in communication via a switch” (e.g. in paragraph 36). It does not describe any particular data of a feature that is associated with multiple machine learning models in multiple distributed systems with both types of nodes. For example, it does not state that data indicating a change in accuracy of a particular model would be associated with models in a different distributed system, particularly not by using data parallelism and model parallelism and with both homogeneous and heterogeneous compute nodes. As such, the claim lacks written description. Claims 11 and 20 have similar issues and thus also fail to comply with the written description requirement. Due at least to their dependency upon claims 1, 11, and 20, dependent claims 2-10, 12-19, 21-25 also fail to comply with the written description requirement. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite an apparatus to receive and determine. The limitation “determine, based upon (1) machine learning model accuracy change data, (2) network bandwidth utilization change data, (3) compression compute time data, and (4) the network telemetry data, whether to selectively compress Tensor data…” as recited in claim 1 is a process, under the broadest reasonable interpretation, covering performance of the limitations in the mind or by pen and paper (See Berkheimer v. HP, Inc., 881 F.3d 1360, 1366, 125 USPQ2d 1649 (Fed. Cir. 2018)) but for the recitation of generic computer components. That is, other than the additional elements noted in the next paragraph below, the limitation “determine, based upon (1) machine learning model accuracy change data, (2) network bandwidth utilization change data, (3) compression compute time data, and (4) the network telemetry data, whether to selectively compress Tensor data…” in the context of the claim encompasses the user making judgements. If a claimed limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “mental processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, the claim recites additional elements. The claim recites “circuitry”, but this is recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using a generic computer component (e.g. See MPEP 2106.05(f)). The limitations “apparatus to be communicatively coupled to a network or fabric to which a plurality of compute nodes is coupled via at least one switch and at least one network interface controllers (NIC), the plurality of compute nodes performing distributed training of multiple instances of an Artificial Intelligence (AI) model that includes exchanging Tensor data amongst the plurality of compute nodes” and “Tensor data generated by the plurality of compute nodes… the machine learning model accuracy change data is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism: and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes” amounts to generally linking the use of the judicial exception to a particular technological environment or field of use (e.g. see MPEP 2106.05(h)). Moreover, the limitation “receive network telemetry data relating to congestion in the network or fabric to which the plurality of compute nodes is coupled… wherein: the network telemetry data comprises NIC event tracking data from the at least one NIC and switch congestion notification data from the at least one switch” is considered as insignificant extra-solution activity (see MPEP 2106.05(g)). Accordingly, the additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are no more than a generic computer component and/or field of use. With respect to “receive network telemetry data relating to congestion in the network or fabric to which the plurality of compute nodes is coupled… wherein: the network telemetry data comprises NIC event tracking data from the at least one NIC and switch congestion notification data from the at least one switch” considered as insignificant extra-solution activity, MPEP 2106.05(d)(II) indicates that mere receiving and transmitting data is a well-understood, routine, and conventional function when it is claimed in a merely generic manner (as it is here; note e.g. Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362, OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015), buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014), etc.). Therefore, the claims are not patent eligible. Regarding claim 2, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim merely further describes determining of times (encompassing the user making judgements) and selectively compressing (encompassing the user making calculations), which are mental steps and does not include any additional elements. Regarding claim 3, the claim does not include any additional elements that are sufficient to amount to significantly more than the judicial exception. For example, the claim merely further describes how the time is determined (encompassing the user making calculations), which is part of the mental steps and does not include any additional elements. Regarding claim 4, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim merely further describes calculating (encompassing the user making calculations), which is part of the mental steps and does not include any additional elements. Regarding claim 5, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim further includes detecting presence of indicia and decompressing data can be performed in the mind or pen and paper. At best, it amounts to generally linking the use of the judicial exception to a particular technological environment or field of use (e.g. see MPEP 2106.05(h)). Regarding claim 6, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim further includes a list of integrated circuits, but this is recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using a generic computer component (e.g. See MPEP 2106.05(f)). Regarding claim 7, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim further includes a network switch chip, but this is recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using a generic computer component (e.g. See MPEP 2106.05(f)). Regarding claim 8, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim further includes a network switch and selectively compressing using the network switch, but this is recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using a generic computer component (e.g. See MPEP 2106.05(f)). Regarding claim 9, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim further logic inputs, but this is considered as insignificant extra-solution activity, MPEP 2106.05(d)(II) indicates that mere receiving/input of data is a well-understood, routine, and conventional function when it is claimed in a merely generic manner (as it is here; note e.g. Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362, OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015), buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014), etc.). Regarding claim 10, the claim does not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception. For example, the claim further describes the apparatus, but this is recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using a generic computer component (e.g. See MPEP 2106.05(f)). Response to Arguments Previous rejections under 35 USC 112 have been withdrawn in view of amendments. With respect to 35 US 101, the amendments do not appear to overcome the rejections. Applicant argues in substance that claims 1, 11, and 20 (it is noted that the rejections under 35 US 101 are directed to claim 1 and its dependent claims 2-10, not claims 11 and 20 and their associated dependent claims) are allegedly clearly directed to solving technological problems. Applicant also cites some example case law, but does not explain the relevance of those examples. As noted in the rejections above, other than the recited additional elements, claim 1 is generally directed to taking data and making a determination based on that data, which are mental steps (and/or performed with pen and paper). More specifically, claim 1 does not actually result in any “solving technological problems”. Claim 1 does not even have a step of compressing the tensor data based on the determining step. The claim only results in a determination, i.e. it does not actually solve any technological problem. A human can receive data and make a decision as to whether to compress data based on the data received. With respect to the additional elements, however as noted above, they amount to no more than mere instructions to amount to generally linking the use of the judicial exception to a particular technological environment or field of use (e.g. see MPEP 2106.05(h)). More specifically, this is merely applied to well-known distributed computer systems. See rejections above for details. As such, applicant’s arguments are not persuasive. Applicant’s arguments with respect to the prior art have been considered but are moot in view of new grounds of rejection. See Chen et al. (US 12242928 B1) below. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 9-14, 20-22, and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Anwar et al. (US 20220156633 A1) in view of Biederman (US 7069342 B1), Zhang et al. (US 20190205759 A1), Zur (US 20060203730 A1), and Chen et al. (US 12242928 B1). As per independent claim 1, Anwar teaches an apparatus to be communicatively coupled to a network or fabric to which a plurality of compute nodes is coupled via at least one network interface controller (NIC) (e.g. in paragraphs 3, 14, 19, and 110, “devices are pieces of hardware in a network… update sent between nodes within a distributed system… adapt the size of each update based on the quality of service (e.g. throughput) of the communication link over which the update is to be sent… network interface 160 facilitates connections between the computing device and one or more other computing devices over a network. For example, the network interface 160 may be an Ethernet network interface, a Wi-Fi network interface, or a cellular network interface”), the plurality of compute nodes performing distributed training of multiple instances of an Artificial Intelligence (Al) model that includes exchanging Tensor data amongst the plurality of compute nodes (e.g. in paragraphs 15, 41, and 56, “training a machine learning model in a distributed system, the distributed system comprising a plurality of nodes that exchange updates to communally train the machine learning model… update to a local model from one or more other nodes in the distributed system, the local model being a locally maintained version of the machine learning model [i.e. multiple local instances corresponding to multiple nodes]… sending an update to the one or more other nodes in the distributed system… training involves adjusting the parameters (the weights) of the neural network to optimize… each update will also include the identifiers for the updates (e.g. the indices for the relevant parameter weights)… each update may specify either the update parameter or the gradient for the updated parameter”, i.e. tensor data), the apparatus comprising: circuitry and logic (e.g. in paragraphs 106-107) to perform operations comprising: receive network telemetry data relating to congestion in the network or fabric to which the plurality of compute nodes is coupled (e.g. in paragraphs 14 and 19, “update sent between nodes within a distributed system… adapt the size of each update based on the quality of service (e.g. throughput [relating to congestion]) of the communication link over which the update is to be sent… monitoring a quality of service of a communication link between the node and the one or more other nodes [i.e. network telemetry data]… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate [relating to congestion]”); and determine, based upon a feature(s) including (4) the network telemetry data, whether to selectively compress Tensor data generated by the plurality of compute nodes (e.g. in paragraphs 19 and 86, “monitoring a quality of service of a communication link between the node and the one or more other nodes… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate or network availability… server adapts whether it sends compressed or full updates… the size of each update is varied based on…the current network quality”), wherein: the feature(s) is associated, at least in part, with multiple machine learning models to be processed (e.g. in paragraphs 38-39, “nodes 20 contribute to training the global model… node stores a local model 22”, and figure 1), but does not specifically teach at least one switch, the feature(s) including (1) machine learning model accuracy change data, (2) network bandwidth utilization change data, and (3) compression compute time data, and wherein: the network telemetry data comprises NIC event tracking data from the at least one NIC and switch congestion notification data from the at least one switch; the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Biederman teaches compression determined based upon features including network bandwidth utilization change data and compression compute time data (e.g. in column 1 lines 27-29, column 2 lines 20-27, column 4 lines 11-32, column 6 lines 18-32, column 7 lines 1-19 and column 8 lines 22-49, “Based on the contents of each packet and, in one embodiment, the degree of network congestion [i.e. network bandwidth utilization] currently being experienced [i.e. changes], packet compression selector 202 either forwards each packet to a selected one of compression units 204, or designates particular packets to be forwarded with their contents uncompressed… data that is relatively easy to compress, e.g., data which is compressed and uncompressed relatively quickly [i.e. compression compute time], is compressed to save the time required to perform compression operations on data which is more difficult to compress, while freeing bandwidth for the transmission of data which is more difficult to compress… the amount of time required to compress… a compression ratio may be adapted based on data content and current or estimated network traffic”, i.e. changes). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Anwar to include the teachings of Biederman because one of ordinary skill in the art would have recognized the benefit of facilitation of assessing and adapting compression to relevant attributes of network quality, but does not specifically teach at least one switch, the feature(s) including (1) machine learning model accuracy change data, and wherein: the network telemetry data comprises NIC event tracking data from the at least one NIC and switch congestion notification data from the at least one switch; the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Zhang teaches compression determined based upon machine learning model accuracy change data (e.g. in paragraph 14, “selecting, from the parameters included in the to-be-compressed layer, the pruning [i.e. a form of compression] number of parameters… stopping execution of the pruning…operations in response to determining that an accuracy of the current trained neural network is lower than a preset accuracy”, i.e. expected machine learning model accuracy change). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Zhang because one of ordinary skill in the art would have recognized the benefit of facilitating accurate models, but does not specifically teach at least one switch and wherein: the network telemetry data comprises NIC event tracking data from the at least one NIC and switch congestion notification data from the at least one switch; the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Zur teaches nodes coupled via at least one switch and receiving network telemetry data comprising NIC event tracking data from the at least one NIC and switch congestion notification data from the at least one switch (e.g. in paragraphs 6, 21, 25-26, and 59, “network may comprise a plurality of end points (EPs) and a plurality of switches and/or routers… switched or routed from their source to their destination… the network switch 106 may experience congestion due to, for example, limitations in the bandwidth of the output network signal 114 or limited buffering capabilities or both. As a result, the network switch 106 may generate a congestion indication 110, which may be communicated to a stack on the network destination device… NDD NIC 202 may be implemented within a network destination device… congestion experience (CE)… filter congestion indications… CE events may be received by the NSD NIC 502”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Zur because one of ordinary skill in the art would have recognized the benefit of facilitating network operation and/or reducing latency (e.g. also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP 2143(B)]), but does not specifically teach the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Chen teaches a feature(s) being associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism, wherein the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes (e.g. in column 2 lines 53-67, column 3 lines 46-59, column 8 lines 22-45, column 13 lines 10-18, and column 21 lines 10-32, “distributed training may be orchestrated with…a fleet of distributed/parallel computing resources in various embodiments… automation techniques described involve aspects of both data parallelism (in which operations on different subsets of a training data set may be performed in parallel at multiple resources) as well as model parallelism (in which computations associated with training several different model versions may be performed in parallel)… various advantages, including…reducing the overall amount of time required for obtaining high-quality versions of machine learning models, (b) enhancing the user experience of clients interacting with machine learning services at which models are trained, e.g., by orchestrating large numbers of training experiments without requiring detailed guidance regarding individual experiments, and/or (c) reducing the amount of computations which may have to be re-performed if individual resources run out of memory, by automatically subdividing down the work of training models with very large data sets into smaller per-batch sub-tasks… several different [i.e. heterogeneous] fleets of compute resources, such as Class-A compute resource fleet 152 [i.e. homogeneous], Class-B compute resource fleet 156 [i.e. homogeneous], and so on… MLS 102 may distribute the training workload based on the performance capabilities of the available compute resources… compute resources may, for example, comprise respective nodes… the compute resources need not be homogeneous [i.e. both homogeneous and heterogeneous]… CPU/GPU utilization levels at the compute resources, the average memory utilization levels at the compute resources, the average or peak network bandwidth usage in the network”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Chen because one of ordinary skill in the art would have recognized the benefit of reducing time required for obtaining quality models and/or reducing computations. As per claim 9, the rejection of claim 1 is incorporated and the combination further teaches wherein the circuitry and logic include network monitor logic having multiple inputs including one or more of a network telemetry data input and congestion notification input (e.g. Anwar, in paragraphs 19 and 21, “monitoring a quality of service of a communication link between the node and the one or more other nodes… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate or network availability… quality of service increases when the throughput increases and the quality of service decreases when the throughput decreases”; Zur, in paragraphs 6, 21, 25-26, and 59, “generate a congestion indication 110, which may be communicated to a stack on the network destination device… NDD NIC 202 may be implemented within a network destination device… congestion experience (CE)… CE events may be received by the NSD NIC 502”). As per claim 10, the rejection of claim 1 is incorporated and the combination further teaches wherein the apparatus comprises one of the plurality of compute nodes (e.g. Anwar, in paragraphs 15 and 26, “receiving an update to a local model from one or more other nodes in the distributed system… and sending an update to the one or more other nodes in the distributed system… may be implemented in any of the nodes, for instance, in one of the workers or in the server. In addition, one node may function as both a worker and a server. When updates are being sent from the server to a plurality of workers, each update to a worker may be based on the quality of service (e.g. throughput) of the communication link to that specific worker”). As per independent claim 11, Anwar teaches a method for training an Artificial Intelligence (AI) model, comprising: implementing respective instances of the Al model on a plurality of compute nodes (e.g. in paragraph 15, “a local model from…other nodes in the distributed system, the local model being a locally maintained version of the machine learning model”) interconnected via a network or fabric and at least one network interface controller (e.g. in paragraphs 3, 14, 19, and 110, “devices are pieces of hardware in a network… update sent between nodes within a distributed system… network interface 160 facilitates connections between the computing device and one or more other computing devices over a network. For example, the network interface 160 may be an Ethernet network interface, a Wi-Fi network interface, or a cellular network interface”); processing respective batches of training data with the respective instances of the Al model at respective compute nodes, the processing including calculation of local model gradient data (e.g. in paragraphs 15, 27, and 56, “a plurality of nodes that exchange updates to communally train the machine learning model… receiving an update to a local model from one or more other nodes in the distributed system, the local model being a locally maintained version of the machine learning model and the update specifying a change to one or more parameters of the local model… training the local model based on training data to obtain the updated local model comprising updated parameters… update may specify…the gradient for the updated parameter”); exchanging local model gradient data amongst the plurality of compute nodes by transmitting the local model gradient data via the network or fabric (e.g. in paragraphs 15 and 56, “a plurality of nodes that exchange updates to communally train the machine learning model… the update specifying a change to one or more parameters of the local model… update may specify…the gradient for the updated parameter”); and updating local weights in the instances of the Al model on the plurality of compute nodes (e.g. Anwar, in paragraphs 15, 41, and 56, “the update specifying a change to one or more parameters of the local model… the parameters (the weights) of the neural network to optimize… each update will also include the identifiers for the updates (e.g. the indices for the relevant parameter weights)… update may specify either the update parameter”), wherein: the local model gradient data are exchanged by applying selective compression to the local model gradient data in consideration of network or fabric quality (e.g. in paragraphs 14, 19, and 86, “update sent between nodes within a distributed system… adapt the size of each update based on the quality of service (e.g. throughput) of the communication link over which the update is to be sent… monitoring a quality of service of a communication link between the node and the one or more other nodes… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate or network availability… server adapts whether it sends compressed or full updates… the size of each update is varied based on the size of the gradients and the current network quality”), the selective compression is to be determined, based at least in part, upon a feature(s) including (4) network telemetry data (e.g. in paragraphs 19 and 86, “monitoring a quality of service of a communication link between the node and the one or more other nodes [i.e. network telemetry data]… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate or network availability… server adapts whether it sends compressed or full updates… the size of each update is varied based on…the current network quality”), wherein: the feature(s) is associated, at least in part, with multiple machine learning models to be processed (e.g. in paragraphs 38-39, “nodes 20 contribute to training the global model… node stores a local model 22”, and figure 1), but does not specifically teach at least one switch, wherein quality includes congestion, the feature(s) including (1) machine learning model accuracy change data, (2) network bandwidth utilization change data, and (3) compression compute time data, and the network telemetry data comprises NIC event tracking data and switch congestion notification data from the at least one switch; the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Biederman teaches interconnection via at least one switch and quality including congestion (e.g. in column 1 lines 20-29, column 2 lines 20-27, column 4 lines 11-32, column 7 lines 1-19 and column 8 lines 22-49, “computer…connected to the Internet…network switching devices… a compression switch… as network congestion decreases, data compression system 200 reduces the compression ratio used for the different compression types to increase the speed of processing… a compression ratio may be adapted based on data content and current or estimated network traffic”) and compression determined based upon network bandwidth utilization change data and compression compute time data (e.g. in column 4 lines 11-32, column 6 lines 18-32, column 7 lines 1-19 and column 8 lines 22-49, “Based on the contents of each packet and, in one embodiment, the degree of network congestion [i.e. network bandwidth utilization] currently being experienced [i.e. changes], packet compression selector 202 either forwards each packet to a selected one of compression units 204, or designates particular packets to be forwarded with their contents uncompressed… data that is relatively easy to compress, e.g., data which is compressed and uncompressed relatively quickly [i.e. compression compute time], is compressed to save the time required to perform compression operations on data which is more difficult to compress, while freeing bandwidth for the transmission of data which is more difficult to compress… the amount of time required to compress… a compression ratio may be adapted based on data content and current or estimated network traffic”, i.e. changes). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Anwar to include the teachings of Biederman because one of ordinary skill in the art would have recognized the benefit of facilitating connection in a network and assessing and adapting to relevant attributes of network quality, but does not specifically teach the feature(s) including (1) machine learning model accuracy change data and wherein the network telemetry data comprises NIC event tracking data and switch congestion notification data from the at least one switch; the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Zhang teaches compression determined based upon machine learning model accuracy change data (e.g. in paragraph 14, “selecting, from the parameters included in the to-be-compressed layer, the pruning [i.e. a form of compression] number of parameters… stopping execution of the pruning…operations in response to determining that an accuracy of the current trained neural network is lower than a preset accuracy”, i.e. expected machine learning model accuracy change). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Zhang because one of ordinary skill in the art would have recognized the benefit of facilitating accurate models, but does not specifically teach wherein the network telemetry data comprises NIC event tracking data and switch congestion notification data from the at least one switch; the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Zur teaches network telemetry data comprising NIC event tracking data and switch congestion notification data from at least one switch (e.g. in paragraphs 6, 21, 25-26, and 59, “network may comprise a plurality of end points (EPs) and a plurality of switches and/or routers… switched or routed from their source to their destination… the network switch 106 may experience congestion due to, for example, limitations in the bandwidth of the output network signal 114 or limited buffering capabilities or both. As a result, the network switch 106 may generate a congestion indication 110, which may be communicated to a stack on the network destination device… NDD NIC 202 may be implemented within a network destination device… congestion experience (CE)… filter congestion indications… CE events may be received by the NSD NIC 502”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Zur because one of ordinary skill in the art would have recognized the benefit of reducing latency (e.g. also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP 2143(B)]), but does not specifically teach the feature(s) is associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism; and the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes. However, Chen teaches a feature(s) being associated, at least in part, with multiple machine learning models to be processed by multiple distributed hardware systems using data parallelism and model parallelism, wherein the multiple distributed hardware systems are comprised in homogeneous and heterogeneous compute nodes (e.g. in column 2 lines 53-67, column 3 lines 46-59, column 8 lines 22-45, column 13 lines 10-18, and column 21 lines 10-32, “distributed training may be orchestrated with…a fleet of distributed/parallel computing resources in various embodiments… automation techniques described involve aspects of both data parallelism (in which operations on different subsets of a training data set may be performed in parallel at multiple resources) as well as model parallelism (in which computations associated with training several different model versions may be performed in parallel)… various advantages, including…reducing the overall amount of time required for obtaining high-quality versions of machine learning models, (b) enhancing the user experience of clients interacting with machine learning services at which models are trained, e.g., by orchestrating large numbers of training experiments without requiring detailed guidance regarding individual experiments, and/or (c) reducing the amount of computations which may have to be re-performed if individual resources run out of memory, by automatically subdividing down the work of training models with very large data sets into smaller per-batch sub-tasks… several different [i.e. heterogeneous] fleets of compute resources, such as Class-A compute resource fleet 152 [i.e. homogeneous], Class-B compute resource fleet 156 [i.e. homogeneous], and so on… MLS 102 may distribute the training workload based on the performance capabilities of the available compute resources… compute resources may, for example, comprise respective nodes… the compute resources need not be homogeneous [i.e. both homogeneous and heterogeneous]… CPU/GPU utilization levels at the compute resources, the average memory utilization levels at the compute resources, the average or peak network bandwidth usage in the network”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Chen because one of ordinary skill in the art would have recognized the benefit of reducing time required for obtaining quality models and/or reducing computations. As per claim 12, the rejection of claim 11 is incorporated and the combination further teaches wherein the at least one switch is coupled to the plurality of compute nodes via a plurality of links (e.g. Anwar, in paragraphs 19 and 86, “a communication link between the node and the one or more other nodes”, i.e. a plurality of links for more other nodes; Biederman, in column 1 lines 20-29, column 2 lines 20-38, and column 9 lines 15-18, “computer…connected to the Internet…network switching devices… a compression switch… output interface forwards the compressed data across the network… determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”), the method further comprising: detecting there is congestion on a link (e.g. Biederman, in column 2 lines 20-38 and column 9 lines 15-18, “determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”); and in response thereto, compressing local model gradient data generated at a source compute node coupled to the link and sending the compressed local model gradient data over the link (e.g. Anwar, in paragraphs 15, 19, and 86, “update specifying a change to one or more parameters of the local model… monitoring a quality of service of a communication link between the node and the one or more other nodes… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate or network availability… sends compressed…updates… the size of each update is varied based on the size of the gradients and the current network quality”; Biederman, in column 2 lines 20-38, “a packet to be compressed… output interface forwards the compressed data across the network”). As per claim 13, the rejection of claim 11 is incorporated and the combination further teaches wherein the at least one switch is coupled to the plurality of compute nodes via a plurality of links (e.g. Anwar, in paragraphs 19 and 86, “a communication link between the node and the one or more other nodes”, i.e. a plurality of links for more other nodes; Biederman, in column 1 lines 20-29, column 2 lines 20-38, and column 9 lines 15-18, “computer…connected to the Internet…network switching devices… a compression switch… output interface forwards the compressed data across the network… determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”), the method further comprising: at the at least one switch, receiving a packet containing local gradient data via a first link from a first compute node and having a second compute node as a destination (e.g. Anwar, in paragraphs 19 and 86, “a communication link between the node and the one or more other nodes”, i.e. a first link from a first node to a second node; Biederman, in column 1 lines 27-29, column 2 lines 20-38, and column 9 lines 15-18, “a packet to be compressed… network switching devices… a compression switch… output interface forwards the compressed data across the network”); determining there is congestion on a second link used to forward the packet to the second compute node (e.g. Anwar, in paragraphs 19 and 86, “a communication link between the node and the one or more other nodes”; Biederman, in column 2 lines 20-38 and column 9 lines 15-18, “a packet to be compressed… output interface forwards the compressed data across the network… determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”); and compressing the local gradient data in the packet prior to forwarding the packet via the second link to the second node (e.g. Anwar, in paragraphs 15, 19, and 86, “update specifying a change to one or more parameters of the local model… monitoring a quality of service of a communication link between the node and the one or more other nodes… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate or network availability… sends compressed…updates… the size of each update is varied based on the size of the gradients and the current network quality”; Biederman, in column 2 lines 20-38, “a packet to be compressed… output interface forwards the compressed data across the network”). As per claim 14, the rejection of claim 11 is incorporated and the combination further teaches wherein the at least one switch is coupled to the plurality of compute nodes via a plurality of links (e.g. Anwar, in paragraphs 19 and 86, “a communication link between the node and the one or more other nodes”, i.e. a plurality of links for more other nodes; Biederman, in column 1 lines 20-29, column 2 lines 20-38, and column 9 lines 15-18, “computer…connected to the Internet…network switching devices… a compression switch… output interface forwards the compressed data across the network… determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”), the method further comprising: at the at least one switch, receiving a packet containing local gradient data via a first link from a first compute node, the packet associated with a broadcast message associated with a broadcast group comprising a plurality of destination compute nodes (e.g. Anwar, in paragraphs 15, 19, and 86, “sending an update to the…more other nodes in the distributed system [i.e. a broadcast group]… a communication link between the node and the one or more other nodes”; Biederman, in column 1 lines 27-29, column 2 lines 20-38, and column 9 lines 15-18, “a packet to be compressed… network switching devices… a compression switch… output interface forwards the compressed data across the network”); determining there is congestion on a second link used to forward the packet to a destination compute node among the plurality of destination compute nodes in the broadcast group (e.g. Anwar, in paragraph 15, “sending an update to the…more other nodes in the distributed system [i.e. a broadcast group]… a communication link between the node and the one or more other nodes”; Biederman, in column 2 lines 20-38 and column 9 lines 15-18, “a packet to be compressed… output interface forwards the compressed data across the network… determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”); and compressing the local gradient data in the packet prior to forwarding the packet via the second link to the destination compute node (e.g. Anwar, in paragraphs 15, 19, and 86, “update specifying a change to one or more parameters of the local model… monitoring a quality of service of a communication link between the node and the one or more other nodes… Quality of service may be represented by any number of parameters, including throughput, bandwidth, signal to noise ratio, channel quality, received signal strength, error rate or network availability… sends compressed…updates… the size of each update is varied based on the size of the gradients and the current network quality”; Biederman, in column 2 lines 20-38, “a packet to be compressed… output interface forwards the compressed data across the network”). Claims 20-21 are the system claims corresponding to method claims 11-12, and are rejected under the same reasons set forth and the combination further teaches a plurality of compute nodes interconnected via a network or fabric, wherein a compute node of the plurality of compute nodes comprises at least one processor coupled to memory and a network or fabric interface coupled to the network or fabric (e.g. Anwar, in paragraphs 19, 86, 104, and 111, “a communication link between the node and the one or more other nodes… computing device 100 includes a bus 110, a processor 120, a memory 130, a persistent storage device 140… and a network interface… realized in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. For instance, hardware may include processors, microprocessors, electronic circuitry, electronic components, integrated circuits, etc. Implementations of the subject matter described in this specification can be realized using one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus”). As per claim 22, the rejection of claim 21 is incorporated and the combination further teaches determine a compression ratio to be applied to the compressed local model gradient data and compress the local model gradient data with the compression ratio that is determined (e.g. Anwar, in paragraphs 15, 56, and 86, “the update specifying a change to one or more parameters of the local model… update may specify…the gradient for the updated parameter… monitoring a quality of service of a communication link between the node and the one or more other nodes… server adapts whether it sends compressed or full updates… the size of each update is varied based on the size of the gradients and the current network quality”; Biederman, in column 2 lines 46-53, “the estimated compression ratio... The compression levels may be dynamically adapted in response to the congestion level of the network”). As per claim 24, the rejection of claim 20 is incorporated and the combination further teaches wherein the plurality of compute nodes comprise multiple processors interconnected via a plurality of input-output (IO) interconnects, wherein the multiple processors comprise one or more of Graphic Processor Units (GPUs), Tensor Processing Units (TPUs), Data Processor Units (DPUs), Infrastructure Processing Units (IPUs), Al processors, Al inference units, and Field Programmable Gate Arrays (FPGAs) (e.g. Anwar, in paragraphs 19, 86, 104, and 111, “a communication link between the node and the one or more other nodes… computing device 100 includes a bus 110, a processor 120, a memory 130, a persistent storage device 140… and a network interface… network interface 160 facilitates connections between the computing device and one or more other computing devices over a network… realized in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. For instance, hardware may include processors, microprocessors, electronic circuitry, electronic components, integrated circuits, etc. Implementations of the subject matter described in this specification can be realized using one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus”). Claims 2, 4-8, 16, 18, and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Anwar et al. (US 20220156633 A1) in view of Biederman (US 7069342 B1), Zhang et al. (US 20190205759 A1), Zur (US 20060203730 A1), and Chen et al. (US 12242928 B1), and further in view of Nagapudi et al. (US 20110058474 A1). As per claim 2, the rejection of claim 1 is incorporated, but the combination does not specifically teach determine, as a function of the telemetry data, a network pause time; determine, for Tensor data to be exchanged between at least two of the plurality of compute nodes, a compute time to compress the Tensor data; and selectively compress the Tensor data when the compute time is less than the network pause time. However, the combination teaches Tensor data (e.g. Anwar, in paragraphs 15, 41, and 56, “training involves adjusting the parameters (the weights) of the neural network to optimize… each update will also include the identifiers for the updates (e.g. the indices for the relevant parameter weights)… each update may specify either the update parameter or the gradient for the updated parameter”, i.e. tensor data) and Nagapudi teaches determine, as a function of the telemetry data, a network pause time (e.g. in paragraph 35, “signal network switch 100 to pause the transmission of data across the link 160”, i.e. pause time), determine, for data to be exchanged between at least two of a plurality of compute nodes, a compute time to compress the data (e.g. in paragraphs 35 and 40, “network switch 100 can enable or disable compression based on the receipt of pause signals… When a pause signal is received, the network switch 100 can enable compression”, i.e. compute time), and selectively compress the data when the compute time is less than the network pause time (e.g. in paragraphs 40-41, “indicating that the congestion on link 160 has cleared or decreased, then the network switch 100 can respond by…suspending compression”, compression only when compute time is less than pause time). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Nagapudi because one of ordinary skill in the art would have recognized the benefit of decreasing end to end latency. As per claim 4, the rejection of claim 2 is incorporated and the combination further teaches wherein: the circuitry and logic are to calculate a compression ratio to be applied to compress Tensor data (e.g. Anwar, in paragraphs 15, 41, and 56, “training involves adjusting the parameters (the weights) of the neural network to optimize… each update will also include the identifiers for the updates (e.g. the indices for the relevant parameter weights)… each update may specify either the update parameter or the gradient for the updated parameter”, i.e. tensor data; Biederman, in column 7 lines 1-19 and column 8 lines 22-49, “as network congestion decreases, data compression system 200 reduces the compression ratio used for the different compression types to increase the speed of processing… a compression ratio may be adapted based on data content and current or estimated network traffic”). As per claim 5, the rejection of claim 4 is incorporated and the combination further teaches wherein the circuitry and logic are to detect presence of compression indicia in Tensor data received by the apparatus identifying a compression type and to decompress received compressed Tensor data when the circuitry and logic detect presence of the compression indicia (e.g. Anwar, in paragraphs 15, 41, and 56, “the parameters (the weights) of the neural network to optimize… each update will also include the identifiers for the updates (e.g. the indices for the relevant parameter weights)… each update may specify either the update parameter or the gradient for the updated parameter”, i.e. tensor data; Biederman, in column 7 lines 1-19 and column 8 lines 22-49, “three compression units 204, e.g., "type A", "type B", and "type C"… compression levels may generally differ based…the amount of time required to compress and decompress data from a compressed format… packets identified as having "type A" content may be compressed using a low compression level”). As per claim 6, the rejection of claim 5 is incorporated and the combination further teaches wherein the circuitry and logic are embedded in an integrated circuit comprising one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), an Al processor, Al inference unit, an Infrastructure Processing Unit (IPU) and a Data Processing Unit (DPU) (e.g. Biederman, in column 1 lines 27-29 and column 4 lines 11-32, “CPU”). As per claim 7, the rejection of claim 4 is incorporated and the combination further teaches wherein the circuitry and logic are implemented in a network switch chip (e.g. Biederman, in column 1 lines 27-29, column 2 lines 20-27, and column 4 lines 11-32, “network switching devices… a compression switch”). As per claim 8, the rejection of claim 4 is incorporated and the combination further teaches wherein the circuitry and logic are implemented in a network switch and wherein the circuitry and logic are further configured to selectively compress Tensor data in packets received at the network switch and broadcast the packets via a plurality of transmit ports (e.g. Anwar, in paragraphs 15, 41, and 56, and 86, “the parameters (the weights) of the neural network to optimize… each update will also include the identifiers for the updates (e.g. the indices for the relevant parameter weights)… each update may specify either the update parameter or the gradient for the updated parameter [i.e. tensor data]… server adapts whether it sends compressed or full updates… the size of each update is varied based on the size of the gradients and the current network quality”; Biederman, in column 1 lines 27-29, column 2 lines 20-27, column 4 lines 11-32 and 39-44, and column 6 lines 50-54, “a packet to be compressed… network switching devices… a compression switch… interfaces 68 may include ports appropriate for communication… information flow may be forwarded based upon…an associated TCP or UDP port number… compression of data may be dynamically adapted both when network congestion increases, and when network congestion decreases”). As per claim 16, the rejection of claim 11 is incorporated and the combination further teaches detecting there is congestion on a link (e.g. Biederman, in column 2 lines 20-38 and column 9 lines 15-18, “a packet to be compressed… output interface forwards the compressed data across the network… determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”) and in response thereto, compressing local model gradient data generated at a source compute node coupled to the link prior to sending the compressed local model gradient data over the link (e.g. Anwar, in paragraphs 15, 19, and 86, “update specifying a change to one or more parameters of the local model… monitoring a quality of service of a communication link between the node and the one or more other nodes… sends compressed…updates… the size of each update is varied based on the size of the gradients and the current network quality”; Biederman, in column 2 lines 20-38, “a packet to be compressed… output interface forwards the compressed data across the network”), but does not specifically teach determining an amount of pause time that would be employed at a source node without compression is greater than an amount of compute time for compressing the local gradient data. However, Nagapudi teaches determining an amount of pause time that would be employed at a source node without compression is greater than an amount of compute time for compressing data (e.g.. in paragraphs 35 and 40, “network switch 100 can enable or disable compression based on the receipt of pause signals… When a pause signal is received, the network switch 100 can enable compression so that when it is allowed to restart transmitting, it can transmit the compressed the data stream”, i.e. pause time greater than compute time). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Nagapudi because one of ordinary skill in the art would have recognized the benefit of decreasing end to end latency. As per claim 18, the rejection of claim 11 is incorporated and the combination further teaches local gradient data generated on a source compute node to be sent over the network or fabric to one or more destination compute nodes (e.g. Anwar, in paragraphs 14, 19, and 86, “update sent between nodes within a distributed system… adapt the size of each update based on the quality of service (e.g. throughput) of the communication link over which the update is to be sent… server adapts whether it sends compressed or full updates… the size of each update is varied based on the size of the gradients and the current network quality”), but does not specifically teach determine, as a function of network telemetry data, a network pause time; determine, for data, a compute time to compress the local gradient data; and selectively compressing the local gradient data when the compute time is less than the network pause time. However, Nagapudi teaches determine, as a function of the telemetry data, a network pause time (e.g. in paragraph 35, “signal network switch 100 to pause the transmission of data across the link 160”, i.e. pause time), determine, for data to be exchanged between at least two of a plurality of compute nodes, a compute time to compress the data (e.g. in paragraphs 35 and 40, “network switch 100 can enable or disable compression based on the receipt of pause signals… When a pause signal is received, the network switch 100 can enable compression”, i.e. compute time), and selectively compress the data when the compute time is less than the network pause time (e.g. in paragraphs 40-41, “indicating that the congestion on link 160 has cleared or decreased, then the network switch 100 can respond by…suspending compression”, compression only when compute time is less than pause time). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Nagapudi because one of ordinary skill in the art would have recognized the benefit of decreasing end to end latency. Claim 23 is the system claim corresponding to method claim 18, and is rejected under the same reasons set forth. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Anwar et al. (US 20220156633 A1) in view of Biederman (US 7069342 B1), Zhang et al. (US 20190205759 A1), Zur (US 20060203730 A1), Chen et al. (US 12242928 B1), and Nagapudi et al. (US 20110058474 A1), and further in view of Yanagihara et al. (US 20030152032 A1). As per claim 3, the rejection of claim 2 is incorporated and the combination further teaches wherein the network pause time is determined as a function of network telemetry data related to congestion (e.g. Nagapudi, in paragraphs 35-36, “network switch 100 can enable or disable compression based on the receipt of pause signals… transmission pausing flow control technique described above…allows the network switch 100 to attempt to reduce congestion”), but does not specifically teach comprising a transmitted packet drop rate. However, Yanagihara teaches congestion being based on data comprising a transmitted packet drop rate (e.g. in paragraph 20, “the congestion information on the network is categorized into a plurality of congestion levels by a queue overflow detection processing performed on a network gateway based on the packet loss rate”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Yanagihara because one of ordinary skill in the art would have recognized the benefit of evaluating relevant features (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP 2143(B)]). Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Anwar et al. (US 20220156633 A1) in view of Biederman (US 7069342 B1), Zhang et al. (US 20190205759 A1), Zur (US 20060203730 A1), and Chen et al. (US 12242928 B1), and further in view of Matthews et al. (US 10931602 B1). As per claim 15, the rejection of claim 14 is incorporated and the combination further teaches forwarding the local gradient data to the plurality of destination compute nodes in the broadcast group and copying a packet containing compressed local gradient data (e.g. Anwar, in paragraphs 15, 19, and 86, “sending an update to the…more other nodes in the distributed system [i.e. a broadcast group]… a communication link between the node and the one or more other nodes”; Biederman, in column 1 lines 27-29, column 2 lines 20-38, and column 9 lines 15-18, “a packet to be compressed… network switching devices… a compression switch… output interface forwards the compressed data across the network”), but does not specifically teach determining transmit ports on at least one switch to be used for forwarding and copying to egress buffers associated with the transmit ports. However, Matthews teaches determining transmit ports on a switch to be used for forwarding data and copying a packet containing data to egress buffers associated with the transmit ports (e.g. in column 23 line 55-66 and column 29 lines 8-20, “Device 500 is generally configured to receive and forward data units 505 to other devices in a network… an integrated circuit, or “chip,” dedicated to performing switching… Once in an egress buffer 544, a data unit 505 (or portion thereof) may be “released” to one or more egress packet processor(s) 550 for processing… replicated to multiple egress queues 545. For instance, a data unit 505 may be linked to separate queues 545 for each of ports 1, 3, and 5”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Matthews because one of ordinary skill in the art would have recognized the benefit of facilitating transmission of data. Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Anwar et al. (US 20220156633 A1) in view of Biederman (US 7069342 B1), Zhang et al. (US 20190205759 A1), Zur (US 20060203730 A1), and Chen et al. (US 12242928 B1), and further in view of Yanagihara et al. (US 20030152032 A1). As per claim 17, the rejection of claim 11 is incorporated and the combination further teaches detecting there is network congestion for a link by at least one of detecting a feature of dropped packets for packets transmitted from a source compute node via the link and detecting a feature of dropped packets for packets received at the source compute node via the link (e.g. Anwar, in paragraphs 19 and 86, “monitoring a quality of service of a communication link between the node and the one or more other nodes… server adapts whether it sends compressed or full updates… the size of each update is varied based on the size of the gradients and the current network quality”; Biederman, in column 1 lines 27-29, column 2 lines 20-44, and column 9 lines 15-18, “output interface forwards the compressed data across the network… Estimating the congestion level may include determining…a number of dropped packets… determines if the degree of congestion or the congestion level in a network or on a network link has changed since the last measurement”), but does not specifically teach wherein the feature includes a rate of dropped packets. However, Yanagihara teaches detecting congestion based on a rate of dropped packets (e.g. in paragraph 20, “the congestion information on the network is categorized into a plurality of congestion levels by a queue overflow detection processing performed on a network gateway based on the packet loss rate”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Yanagihara because one of ordinary skill in the art would have recognized the benefit of evaluating relevant features associated with congestion (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP 2143(B)]). Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Anwar et al. (US 20220156633 A1) in view of Biederman (US 7069342 B1), Zhang et al. (US 20190205759 A1), Zur (US 20060203730 A1), and Chen et al. (US 12242928 B1), and further in view of Bunandar et al. (US 20220172052 A1). As per claim 19, the rejection of claim 11 is incorporated and the combination further teaches wherein the Al model uses a first form of data and compression comprises converting the first form of data to a second form of data (e.g. Anwar, in paragraph 48, “model parameter compression method to reduce the overall amount of data transmitted during the training phase of a distributed system”; compression changes the form of the data to reduced data), but does not specifically teach a first form of data including 32-bit floating point numerical data (FP32) and the second form of data including 16-bit Brain floating point (Bfloat16) data. However, Bunandar teaches a first form of data including 32-bit floating point numerical data (FP32) and a second form of data including 16-bit Brain floating point (Bfloat16) data (e.g. in paragraph 92, “a 32 bit floating-point representation (“float32” or “FP32”)… a 16 bit brain floating-point format (“bfloat16”)”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Bunandar because one of ordinary skill in the art would have recognized the benefit of incorporating well-known forms of data (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP 2143(B)]). Claim 25 is rejected under 35 U.S.C. 103 as being unpatentable over Anwar et al. (US 20220156633 A1) in view of Biederman (US 7069342 B1), Zhang et al. (US 20190205759 A1), Zur (US 20060203730 A1), and Chen et al. (US 12242928 B1), and further in view of Booth et al. (US 20080310422 A1). As per claim 25, the rejection of claim 20 is incorporated, but the combination does not specifically teach wherein the plurality of compute nodes comprise a plurality of servers installed in a rack including a switch to which the plurality of servers are communicatively coupled. However, Booth teaches a plurality of servers installed in a rack including a switch to which the plurality of servers are communicatively coupled (e.g. in paragraphs 23 and 41, “access switch 12 can be implemented as a rack into which "blade" servers are installed and configured”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Booth because one of ordinary skill in the art would have recognized the benefit of incorporating well-known device configurations (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP 2143(B)]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For example, Prakash et al. (US 20190220703 A1) teaches “each worker node 130-1, 130-2,…, 130-N utilizes its local subset of mini-batch data to execute a forward propagation process on the DL model, followed by error backpropagation to compute gradients of the loss with respect to the DL network model parameters” (e.g. in paragraph 29). Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM WONG whose telephone number is (571)270-1399. The examiner can normally be reached Monday-Friday 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, TAMARA KYLE can be reached at (571)272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /W.W/Examiner, Art Unit 2144 09/05/2026 /TAMARA T KYLE/Supervisory Patent Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Show 2 earlier events
Jun 23, 2025
Non-Final Rejection mailed — §101, §103, §112
Sep 19, 2025
Response Filed
Jan 13, 2026
Final Rejection mailed — §101, §103, §112
Mar 30, 2026
Request for Continued Examination
Apr 04, 2026
Response after Non-Final Action
Apr 09, 2026
Non-Final Rejection mailed — §101, §103, §112
Jun 29, 2026
Response Filed
Sep 11, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12736359
MACHINE LEARNING TECHNIQUES FOR TRAVERSAL PATH OPTIMIZATION
4y 4m to grant Granted Sep 15, 2026
Patent 12705009
SHARED AUGMENTED REALITY UNBOXING EXPERIENCE
2y 5m to grant Granted Aug 11, 2026
Patent 12688394
SYSTEM AND METHOD FOR MOLECULAR PROPERTY PREDICTION USING EDGE CONDITIONED IDENTITY MAPPING CONVOLUTION NEURAL NETWORK
4y 9m to grant Granted Jul 21, 2026
Patent 12682258
OPERATIONAL FORECASTING SYSTEM BASED ON ANOMALOUS BEHAVIORS IN COMPLEX SYSTEMS
5y 1m to grant Granted Jul 14, 2026
Patent 12639585
PROACTIVE ALERT AGGREGATION AND CORRELATION MANAGEMENT WITH AUTOMATED SUMMARIZATION
5y 0m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
31%
Grant Probability
58%
With Interview (+27.8%)
4y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 407 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month