DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status
This application claims priority to Chinese PCT Application PCT/CN2022/090818 filed April 30, 2022.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on August 12, 2024 was filed in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim 1, 16, 24 and 28 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US Pat. Pub. 20220271851 to Athul Prasad et al. (hereinafter Prasad).
Regarding claim 1, Prasad teaches A method for wireless communication at a user equipment (UE), comprising:
receiving, from a network node, a configuration for training a machine learning (ML) model for performing channel estimation or reporting channel state information, wherein the configuration indicates a data set for the ML model and one or more learning rate parameters; (Prasad para. [0045] and Fig. 6, below, teach in element 600 configuring a UE for CSI feedback for training.
PNG
media_image1.png
587
724
media_image1.png
Greyscale
Further, Prasad teaches learning rate parameters in Fig. 1 wherein CSI feedback configuration is identified as “periodic or semi-persistent” in Fig. 1 and para. [0029].
PNG
media_image2.png
662
670
media_image2.png
Greyscale
)
and
training, based on the configuration, the ML model, wherein the ML model performs at least one of:
performing, based on the ML model, channel estimation of a channel between the UE and the network node;
or
reporting, based on the ML model, channel state information of the channel between the UE and the network node. (Prasad teaches performing the channel estimation based on the ML model provided in Fig. 6 above, in 620, reporting based on the model when there is a difference between predicted and actual values.) Examiner notes that the “OR” in the claim negates a requirement to reject each element of the alternatives.
Regarding claim 16, Prasad teaches A method for wireless communication, comprising: transmitting a configuration for training, at a user equipment (UE), a machine learning (ML) model for performing channel estimation or reporting channel state information, wherein the configuration indicates a data set for the ML model and one or more learning rate parameters; (Prasad para. [0045] and Fig. 6, below, teach in element 600 configuring a UE for CSI feedback for training.
PNG
media_image1.png
587
724
media_image1.png
Greyscale
Further, Prasad teaches learning rate parameters in Fig. 1 wherein CSI feedback configuration is identified as “periodic or semi-persistent” in Fig. 1 and para. [0029].
PNG
media_image2.png
662
670
media_image2.png
Greyscale
)
and
wherein the ML model is used for at least one of channel estimation of a channel between the UE and a network node or CSI reporting of the channel between the UE and the network node. (Prasad teaches performing the channel estimation based on the ML model provided in Fig. 6 above, in 620, reporting based on the model when there is a difference between predicted and actual values.) Examiner notes that the “OR” in the claim negates a requirement to reject each element of the alternatives.
Regarding claim 24, Prasad teaches A method for wireless communication, comprising: receiving, from a network node, a network-side machine learning (ML) model trained on a reference user equipment (UE)-side ML model for performing channel estimation or reporting channel state information; (Prasad para. [0045] and Fig. 6, below, teach in element 600 configuring a UE for CSI feedback for training.
PNG
media_image1.png
587
724
media_image1.png
Greyscale
Further, Prasad teaches learning rate parameters in Fig. 1 wherein CSI feedback configuration is identified as “periodic or semi-persistent” in Fig. 1 and para. [0029].
PNG
media_image2.png
662
670
media_image2.png
Greyscale
training a UE-side ML model using data received from a server and based on the network-side ML model, wherein the network-side ML model and UE-side ML model comprise at least one of: the network-side ML model used for channel state information (CSI)- reference signal (RS) transmission and the UE-side ML model used for channel estimation; or the network-side ML model used for CSI decoding and the UE-side ML model used for CSI encoding. (Prasad teaches performing the channel estimation based on the ML model provided in Fig. 6 above, in 620, reporting based on the model when there is a difference between predicted and actual values.) Examiner notes that the “OR” in the claim negates a requirement to reject each element of the alternatives.
Regarding claim 28, Prasad teaches A method for wireless communication, comprising:
transmitting a network-side machine learning (ML) model trained on a reference user equipment (UE)-side ML model for performing channel estimation or reporting channel state information; (Prasad para. [0045] and Fig. 6, below, teach in element 600 configuring a UE for CSI feedback for training.
PNG
media_image1.png
587
724
media_image1.png
Greyscale
Further, Prasad teaches learning rate parameters in Fig. 1 wherein CSI feedback configuration is identified as “periodic or semi-persistent” in Fig. 1 and para. [0029].
PNG
media_image2.png
662
670
media_image2.png
Greyscale
)
and
wherein the ML model is used for at least one of channel estimation of the channel between a UE and a network node or CSI reporting of the channel between the UE and the network node. (Prasad teaches performing the channel estimation based on the ML model provided in Fig. 6 above, in 620, reporting based on the model when there is a difference between predicted and actual values.) Examiner notes that the “OR” in the claim negates a requirement to reject each element of the alternatives.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2, 3, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Prasad further in view of US Pat. Pub. 20230351248 to Awn Muhammad et al. (hereinafter Muhammad).
Regarding claim 2, Prasad teaches The method of claim 1 as stated. Prasad does NOT teach wherein the configuration indicates at least one of a batch size or number of epochs to use in training the ML model based on the data set.
In the same field of endeavor, Muhammad teaches wherein the configuration indicates at least one of a batch size or number of epochs to use in training the ML model based on the data set. (Muhammad teaches in paras. [0054] and [0055] that when training a ML model a larger number of epochs enables achieving a high level of accuracy, and a limited number of epochs achieves an acceptable level of accuracy.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Muhammad and Prasad to teach a configuration indicating a number of epochs. Each of Muhammad and Prasad are in the field of wireless communications and ML CSI. One of ordinary skill in the art would have been motivated to combine Muhammad and Prasad in order to leverage AI/ML use by user devices and improve services and processing capabilities as taught in Muhammad para. [0028].
Regarding claim 3, Prasad teaches The method of claim 1 as stated. Prasad does NOT teach wherein the one or more learning rate parameters include at least one of an initial learning rate, a learning rate decaying type, or a decaying ratio.
In the same field of endeavor, Muhammad teaches wherein the one or more learning rate parameters include at least one of an initial learning rate, a learning rate decaying type, or a decaying ratio. (Muhammad teaches in Table 1, that the machine learning capabilities of the UE is provided to enable configuration of the learning rate parameters would include an “initial learning rate” when a UE is capable of machine learning:
PNG
media_image3.png
200
911
media_image3.png
Greyscale
Examiner interprets the model training Y and “full training” designations as teaching an initial learning rate since those UEs capable of ML would have a learning rate capability.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Muhammad and Prasad to teach a configuration indicating a number of epochs. Each of Muhammad and Prasad are in the field of wireless communications and ML CSI. One of ordinary skill in the art would have been motivated to combine Muhammad and Prasad in order to leverage AI/ML use by user devices and improve services and processing capabilities as taught in Muhammad para. [0028].
Regarding claim 17, Prasad teaches The method of claim 16 as stated. Prasad does NOT teach wherein the configuration indicates at least one of a batch size or number of epochs to use in training the ML model based on the data set, wherein the one or more learning rate parameters include at least one of an initial learning rate, a learning rate decaying type, or a decaying ratio, or wherein the configuration indicates a loss function to use in training the ML model.
In the same field of endeavor, Muhammad teaches wherein the configuration indicates at least one of a batch size or number of epochs to use in training the ML model based on the data set,(Muhammad teaches in paras. [0054] and [0055] that when training a ML model a larger number of epochs enables achieving a high level of accuracy, and a limited number of epochs achieves an acceptable level of accuracy.) wherein the one or more learning rate parameters include at least one of an initial learning rate, a learning rate decaying type, or a decaying ratio, or wherein the configuration indicates a loss function to use in training the ML model. (Muhammad teaches in Table 1, that the machine learning capabilities of the UE is provided to enable configuration of the learning rate parameters would include an “initial learning rate” when a UE is capable of machine learning:
PNG
media_image3.png
200
911
media_image3.png
Greyscale
Examiner interprets the model training Y and “full training” designations as teaching an initial learning rate since those UEs capable of ML would have a learning rate capability.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Muhammad and Prasad to teach a configuration indicating a number of epochs. Each of Muhammad and Prasad are in the field of wireless communications and ML CSI. One of ordinary skill in the art would have been motivated to combine Muhammad and Prasad in order to leverage AI/ML use by user devices and improve services and processing capabilities as taught in Muhammad para. [0028].
Claim 4, 5, 7, 8, 18, 20, 22, 25, 26, 29 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of US Pat. Pub. 20240333604 to Yonghe Zhu et al. (hereinafter Zhu).
Regarding claim 4, Prasad teaches The method of claim 1 as stated. Prasad does NOT teach wherein the configuration indicates a loss function used for training the ML model.
In the same field of endeavor, Zhu teaches wherein the configuration indicates a loss function used for training the ML model. (Zhu teaches para. [0073] that the training process may define a loss function that describes the gap or difference between an output value of a neural network and an ideal target value, and teaches that the value of the loss function is less than a threshold value or meets a target.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 5, Prasad teaches The method of claim 1 as stated.
Prasad teaches, wherein training the ML model includes
performing multiple training iterations, (Prasad teaches multiple training iterations as shown in Fig. 6 above) the multiple training iterations comprising:
an initial training iteration including reporting, to the network node and within a timer from receiving the configuration, (Prasad teaches as shown in Fig. 6 and Fig. 1 that reporting is according to a timing configuration)
Prasad does NOT teach a local training gradient associated with training the ML model;
In the same field of endeavor Zhu teaches a local training gradient associated with training the ML model. (Zhu teaches in para. [0091] it may be better ensured that the gradient that is of the model training and that is reported by the terminal to the base station is a gradient in the current round of model training, instead of a gradient in another round of model training, for example, a gradient in a previous round of model training.)
and
one or more remaining training iterations including receiving a global gradient transmitted from the network node, (Zhu para. [0091] further teaches that the training duration may be a global parameter determined by comprehensively considering computing capabilities of the terminals participating in federated learning, complexity of the model, and the like, and then configured by the base station for each terminal)
updating the ML model based on the global gradient, (Zhu teaches in para. [0118] and step 507 in Fig. 5 that the ML model is updated according to an average gradient for all n terminals in a round.)
and
reporting an updated local training gradient associated with the updated ML model within the timer from receiving the global gradient. (Zhu para. [0119]-[0120] teaches terminals participating in federated learning complete the model training within the training duration T at the reporting moment, wherein the terminals report the gradients in the current round of model training to the base station.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 7, Prasad teaches The method of claim 1 as stated. Prasad further teaches receiving a first signaling to trigger a semi-persistent gradient reporting, wherein the semi-persistent gradient reporting comprises a plurality of reporting occasions; (Prasad teaches semi-persistent reporting in Fig. 1 that is triggered as shown above)
Prasad does NOT teach receiving, prior to each reporting occasion, a second signaling conveying global gradients; where a timing gap between the second signaling and a corresponding reporting occasion is greater than a threshold; updating the ML model based on the global gradients; and reporting a local gradient associated with the updated ML model; where the timing gap between the second signaling and the corresponding reporting occasion is not greater than a threshold:
refraining from updating the ML model; and
at least one of: refraining from reporting the local gradient; or reporting an outdated local gradient.
In the same field of endeavor, Zhu teaches receiving, prior to each reporting occasion, a second signaling conveying global gradients; where a timing gap between the second signaling and a corresponding reporting occasion is greater than a threshold; (Zhu teaches in para. [0119] – [0120] n terminals participating in federated learning complete the model training within the training duration T at the reporting moment, the terminals may report the gradients in the current round of model training to the base station, but if the terminal does not complete the model training within the training duration T it no longer reports the gradient in the current round of model training to the base station.)
updating the ML model based on the global gradients; and reporting a local gradient associated with the updated ML model; where the timing gap between the second signaling and the corresponding reporting occasion is not greater than a threshold: (Zhu teaches in para. [0119] – [0120] that terminals participating in federated learning complete the model training within the training duration T at the reporting moment, the terminals may report the gradients in the current round of model training to the base station.)
refraining from updating the ML model; (Zhu teaches that the terminals do not report gradients when training is not complete as taught in para. [0154].
and
at least one of: refraining from reporting the local gradient; or reporting an outdated local gradient. (Zhu teaches in para. [0119] – [0120] teach that the sending of model training that is not completed during a training duration T is OPTIONALLY sent or not sent.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 8, Prasad teaches The method of claim 1 as stated.
Prasad does NOT teach receiving, from the network node, downlink control information indicating resources for reporting an aperiodic local training gradient associated with training the ML model; and reporting, to the network node, the aperiodic local training gradient associated with training the ML model over the resources.
In the same field of endeavor, Zhu teaches receiving, from the network node, downlink control information indicating resources for reporting an aperiodic local training gradient associated with training the ML model; (Zhu para. [0141] teaches a start of an AI/ML training may include sending a DCI specific to AI/ML training. Zhu para. [0142] teaches that the AI/ML training includes “model update information” according to federated learning . Zhu para. [0168] teaches that the model updates are “local gradients” in federated learning. Zhu para. [0153] teaches scheduled reporting and not periodic reporting which is aperiodic.)
and
reporting, to the network node, the aperiodic local training gradient associated with training the ML model over the resources. (Zhu teaches in para. [0144] that a UE transmits its AI model update information to report updated local AI/ML model parameters. Zhu para. [0168] teaches that the model updates are “local gradients” in federated learning. Zhu also teaches in para. [0153] that training gradient updating is according to a schedule and therefore aperiodic.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 18, Prasad teaches The method of claim 16 as stated. Prasad does NOT teach further comprising: receiving, within a timer from transmitting the configuration or within a timer from transmitting a global gradient, a training gradient associated with training the ML model at the UE; and transmitting, for multiple UEs including the UE, an aggregated global training gradient based at least in part on the training gradient and other received training gradients, wherein the aggregated global training gradient is based on the training gradient and other received training gradients received within the timer, and wherein transmitting the aggregated global training gradient is based on the multiple UEs reporting the training gradient and the other received training gradients within the timer or expiration of the timer.
In the same field of endeavor, Zhu teaches receiving, within a timer from transmitting the configuration or within a timer from transmitting a global gradient, a training gradient associated with training the ML model at the UE; (Zhu teaches in Fig. 7B receiving within a timer a training gradient associated with the ML model at the UE “within the training duration”:
PNG
media_image4.png
576
733
media_image4.png
Greyscale
Zhu also teaches “wherein the aggregated global training gradient is based on the training gradient and other received training gradients received within the timer, wherein transmitting the aggregated global training gradient is based on the multiple UEs reporting the training gradient and the other received training gradients within the timer or expiration of the timer. (Zhu Fig. 7B teaches “collect statistics on quantity of UEs that complete the training within the training duration” and Fig. 7C block 706 teaches “calculate an average gradient” which is based on the UES mapped to aggregated global training gradient.”
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 20, Prasad teaches The method of claim 16 as stated. Prasad teaches further comprising transmitting a trigger activating a semi-persistent local gradient reporting, wherein the semi-persistent local gradient reporting comprises a plurality of global gradient reporting instances, (Prasad teaches semi-persistent reporting in Fig. 1 that is triggered as shown above)
and
Prasad does NOT teach further comprising: receiving a training gradient associated with training the ML model at the UE; and transmitting a second signaling conveying an aggregated global gradient used for training the ML model, wherein a time gap between the second signaling and a next reporting occasion is greater than a threshold.
In the same field of endeavor, Zhu teaches receiving a training gradient associated with training the ML model at the UE; (Zhu teaches in Fig. 7C receiving training gradients from UEs:
PNG
media_image5.png
696
1167
media_image5.png
Greyscale
)
and
transmitting a second signaling conveying an aggregated global gradient used for training the ML model, wherein a time gap between the second signaling and a next reporting occasion is greater than a threshold. (Zhu teaches in para. [0119] – [0120] that terminals participating in federated learning complete the model training within the training duration T at the reporting moment, the terminals may report the gradients in the current round of model training to the base station. Zhu Fig. 7B illustrates the threshold for the timing duration:
PNG
media_image4.png
576
733
media_image4.png
Greyscale
)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 22, Prasad teaches The method of claim 16 as stated. Prasad does NOT teach further comprising: receiving, within a timer from transmitting the configuration or within a timer from transmitting a global gradient, an output of a channel state information (CSI) encoder based on training a UE-side ML model at the UE; updating, based on the output of the CSI encoder and other received outputs of other CSI encoders, the ML model for decoding channel state information; and transmitting, to multiple UEs including the UE, an aggregated global training gradient of an output of the ML model.
In the same field of endeavor, Zhu teaches receiving, within a timer from transmitting the configuration or within a timer from transmitting a global gradient, an output [[of a channel state information (CSI) encoder]] based on training a UE-side ML model at the UE; (Zhu teaches in Fig. 7B receiving within a timer a training gradient associated with the ML model at the UE “within the training duration” based on training at the UE:
PNG
media_image4.png
576
733
media_image4.png
Greyscale
Zhu also teaches updating, based on the output of the CSI encoder and other received outputs of other CSI encoders, the ML model for decoding channel state information; (Zhu Fig. 7C illustrates in block 706 “calculate an average gradient” which is based on other received local gradients from UEs:
PNG
media_image5.png
696
1167
media_image5.png
Greyscale
and
transmitting, to multiple UEs including the UE, an aggregated global training gradient of an output of the ML model. (Zhu Fig. 7C illustrates in block 706 transmitting the average gradient mapped to aggregate global training gradient to the UEs).
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 25, Prasad teaches The method of claim 24 as stated. Prasad does NOT teach further comprising transmitting, to the network node, an indication that training the UE-side ML model is completed.
In the same field of endeavor, Zhu teaches receiving, from the UE, an indication that training the UE-side ML model is completed. (Zhu teaches in para. [0008] that a “training completion indication” is sent by the terminal to a node when training of the AI model is complete.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a completion indication. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 26, Prasad teaches The method of claim 24 as stated. Prasad teaches further comprising receiving, from the network node, a configuration for training the UE-side ML model, wherein the configuration indicates a data set for the UE-side ML model and one or more learning rate parameters. (Prasad Fig. 4, block 430 teaches that the network configures each UE with a “learned model” relevant for the UE, which would include learning rate parameters:
PNG
media_image6.png
579
876
media_image6.png
Greyscale
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 29, Prasad teaches The method of claim 28 as stated. Prasad does NOT teach , further comprising receiving, from the UE, an indication that training the UE-side ML model is completed.
In the same field of endeavor, Zhu teaches receiving, from the UE, an indication that training the UE-side ML model is completed. (Zhu teaches in para. [0008] that a “training completion indication” is sent by the terminal to a node when training of the AI model is complete.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a completion indication. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 30, Prasad teaches The method of claim 29 as stated. Prasad does NOT teach further comprising: refining the network-side ML model based on the UE-side ML model as trained using additional data of the UE; and transmitting the refined network-side ML model to one or more UEs.
In the same field of endeavor, Zhu teaches refining the network-side ML model based on the UE-side ML model as trained using additional data of the UE; and transmitting the refined network-side ML model to one or more UEs. (Zhu teaches in Fig. 5B refining the network-side ML model in “calculate an average gradient” in block 507 wherein it is based on additional data of UE, and in the same block “deliver an effective average gradient”:
PNG
media_image7.png
602
782
media_image7.png
Greyscale
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach refining an ML model. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Claims 27 is rejected under 35 U.S.C. 103 as being unpatentable over Prasad further in view of Muhammad, further in view of Zhu.
Regarding claim 27, Prasad teaches The method of claim 26 as stated. Prasad does NOT teach wherein the configuration indicates at least one of a batch size or number of epochs to use in training the UE-side ML model based on the data set, wherein the one or more learning rate parameters include at least one of an initial learning rate, a learning rate decaying type, or a decaying ratio, wherein the configuration indicates the network-side ML model, or wherein the configuration indicates a loss function for training the UE-side ML model.
In the same field of endeavor, Muhammad teaches wherein the configuration indicates at least one of a batch size or number of epochs to use in training the ML model based on the data set, (Muhammad teaches in paras. [0054] and [0055] that when training a ML model a larger number of epochs enables achieving a high level of accuracy, and a limited number of epochs achieves an acceptable level of accuracy.) wherein the one or more learning rate parameters include at least one of an initial learning rate, a learning rate decaying type, or a decaying ratio, or wherein the configuration indicates a loss function to use in training the ML model. (Muhammad teaches in Table 1, that the machine learning capabilities of the UE is provided to enable configuration of the learning rate parameters would include an “initial learning rate” when a UE is capable of machine learning:
PNG
media_image3.png
200
911
media_image3.png
Greyscale
Examiner interprets the model training Y and “full training” designations as teaching an initial learning rate since those UEs capable of ML would have a learning rate capability.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Muhammad and Prasad to teach a configuration indicating a number of epochs. Each of Muhammad and Prasad are in the field of wireless communications and ML CSI. One of ordinary skill in the art would have been motivated to combine Muhammad and Prasad in order to leverage AI/ML use by user devices and improve services and processing capabilities as taught in Muhammad para. [0028].
Claims 6 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Zhu further in view of US Pat. Pub. 202401193 to Gary Boudreau et al. (hereinafter Boudreau) **
Regarding claim 6, Prasad in view of Zhu teaches The method of claim 5 as stated. Prasad does NOT teach further comprising reporting, to the network node, a timestamp associated with the local training gradient, wherein the timestamp corresponds to at least one of a slot index of receiving the configuration, a slot index of receiving the global gradient, an iteration index of when the ML model is trained, or a timestamp received in the configuration.
In the same field of endeavor, Boudreau teaches reporting, to the network node, a timestamp associated with the local training gradient, wherein the timestamp corresponds to at least one of a slot index of receiving the configuration, a slot index of receiving the global gradient, an iteration index of when the ML model is trained, or a timestamp received in the configuration (Boudreau teaches applying Federated Learning (see para. [0005])to CSI (see para. [0027]) and teaches in para. [0078] :
PNG
media_image8.png
278
438
media_image8.png
Greyscale
As shown, the time stamps are adjusted and associated with the gradients sent to the master node).
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Boudreau to teach reporting of time stamps to a network node. Each of Prasad and Boudreau are in the field of wireless communications and machine learning. One of ordinary skill in the art would have been motivated to combine Prasad and Boudreau in order to take full advantage of both the timely local and delayed global information, while allowing gradient descent at both the network edge and control center for improved system performance as taught in Boudreau para. [0006].
Regarding claim 19, Prasad teaches The method of claim 16 as stated. Prasad does NOT teach receiving a training gradient associated with training the ML model along with a timestamp; and transmitting, for multiple UEs including the UE, an aggregated training gradient based at least in part on applying a weight to the training gradient based on the timestamp, wherein the timestamp corresponds to at least one of a slot index of transmitting the configuration, or a slot index of transmitting a global gradient, an iteration index of when the ML model is trained, or a timestamp transmitted in the configuration.
In the same field of endeavor, Boudreau teaches receiving a training gradient associated with training the ML model along with a timestamp; and transmitting, for multiple UEs including the UE, an aggregated training gradient based at least in part on applying a weight to the training gradient based on the timestamp, wherein the timestamp corresponds to at least one of a slot index of transmitting the configuration, or a slot index of transmitting a global gradient, an iteration index of when the ML model is trained, or a timestamp transmitted in the configuration (Boudreau teaches applying Federated Learning (see para. [0005] to CSI (see para. [0027]) and teaches in para. [0078] :
PNG
media_image8.png
278
438
media_image8.png
Greyscale
As shown, the time stamps are adjusted and associated with the gradients sent to the master node).
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Boudreau to teach reporting of time stamps to a network node. Each of Prasad and Boudreau are in the field of wireless communications and machine learning. One of ordinary skill in the art would have been motivated to combine Prasad and Boudreau in order to take full advantage of both the timely local and delayed global information, while allowing gradient descent at both the network edge and control center for improved system performance as taught in Boudreau para. [0006].
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Zhu further in view of 3GPP TS 38.214 V16.8.0 (2021-12) (hereinafter 38.214)
Regarding claim 9, Prasad in view of Zhu teaches The method of claim 8 as stated. Prasad does NOT teach wherein training the ML model is based at least in part on a minimum timing gap between receiving the downlink control information and the resources for reporting the aperiodic local training gradient.
In the same field of endeavor, Zhu teaches wherein training the ML model is based at least in part on [[a minimum timing gap between receiving the downlink control information and the resources for]] reporting the aperiodic local training gradient. (Zhu para. [0085] teaches a method for training an AI model in a wireless network wherein a same time-frequency resource and a same reporting moment may be allocated to terminals participating in federated learning. The terminals participating in federated learning reports gradients of a trained AI model at the same reporting moment by using the same time-frequency resource.)
Neither Zhu nor Prasad teach “a minimum timing gap between receiving the downlink control information and the resources”.
In the analogous field of 3GPP 5G Wireless Standards, 38.214 teaches “a minimum timing gap between receiving the downlink control information and the resources”. (38.214 teaches in Section 5.2.1 page 50, that CSI reporting is triggered by DCI. Further, Section 5.2.1.5.1 “Triggering/activation of CSI Reports and CSI-RS teaches on page 59 teaches a “scheduling offset” mapped to the timing gap which is “between the last symbol of the PDCCH carrying the triggering DCI and the first symbol of the aperiodic CSI-RS resources”. When smaller than the UE reported threshold beamSwitchTiming… or smaller than 48, then the UE provides a value.)
It would have been obvious to one of ordinary skill in the art to combine Prasad and 38.214 to teach a minimum timing gap between receiving the downlink control information and the resources as applied to a machine learning model. Each of Prasad and 38.214 teach CSI reporting. One of ordinary skill would have been motivated to combine Prasad and 38.214 in order to follow the 3GPP established physical layer procedures for data channels for 5G-NR as taught on page 7 of 38.214.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Claim 10, 13 and 23 is rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Zhu further in view of CSI Feedback With Model-Driven Deep Learning of Massive MIMO Systems, IEEE COMMUNICATIONS LETTERS, VOL. 26, NO. 3, MARCH 2022, by Jianhua Guo et al. (hereinafter Guo)
Regarding claim 10, Prasad teaches The method of claim 1, as stated.
Prasad does NOT teach reporting, to the network node and within a timer from receiving the configuration, an output of a channel state information (CSI) encoder based on training the ML model;
receiving, from the network node, a global gradient transmitted from the network node; and updating the ML model based on the global gradient.
In the same field of endeavor, Zhu teaches reporting, to the network node and within a timer from receiving the configuration, an output of a [[channel state information (CSI) encoder]] based on training the ML model; (Zhu teaches in para. [0119] – [0120] that terminals participating in federated learning complete the model training within the training duration T at the reporting moment, the terminals may report the gradients in the current round of model training to the base station.)
receiving, from the network node, a global gradient transmitted from the network node; (Zhu teaches receiving an average gradient mapped to a global gradient based on federated learning as shown in Fig. 7C:
PNG
media_image5.png
696
1167
media_image5.png
Greyscale
and
updating the ML model based on the global gradient. (Zhu teaches as shown in Step 706 in Fig. 7C receiving the average gradient which is applied to the ML model at the UE. )
Neither Prasad nor Zhu teach that the configuration is the output of a “channel state information (CSI) encoder”.
In the same field of endeavor, Guo teaches providing the output of “a channel state information (CSI) encoder”. (Guo teaches “model-driven deep learning” for CSI feedback as shown in Fig. 1, and encoder and a decoder provide CSI feedback mapped as channel state information encoder output:
PNG
media_image9.png
417
793
media_image9.png
Greyscale
Guo, page 547, column 2, third paragraph teaches that the compress feedback information is fed back to the base station to create a CSI matrix.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Guo to teach the providing output of a CSI encoder. Each of Guo and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Guo with Prasad in order to take advantage of the powerful learning ability and address the difficulty in designing a deep learning model that applies to the physical layer of wireless communications including CSI as taught in Guo, page 547, second column, first paragraph.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 13, Prasad teaches The method of claim 1 as stated.
Prasad does NOT teach receiving, from the network node, an output of a channel state information (CSI)- reference signal (RS) transmitter based on training a network-side ML model at the network node; updating the ML model based on a loss computed from the output of the CSI-RS transmitter; and reporting, to the network node, a local gradient of the ML model based on training the ML model.
In the same field of endeavor, Zhu teaches receiving, from the network node, an [[output of a channel state information (CSI)- reference signal (RS) transmitter]] based on training a network-side ML model at the network node; (Zhu Fig. 4 and para. [0087] teach in step 401 a terminal receiving a configuration for training based on an AI model for federated learning)
updating the ML model based on a loss computed from the output of the CSI-RS transmitter; (Zhu Fig. 4 teaches a terminal training an AI model to obtain a gradient in step 402. Zhu para. [0073] teaches that the training process includes a “loss function” wherein the adjusting gradient is associated with a value of the loss function being less than a threshold value.)
and
reporting, to the network node, a local gradient of the ML model based on training the ML model. (Zhu Fig. 4 step 403 and para. [0089] teach reporting a local gradient of the model to the base station).
Neither Prasad nor Zhu teach “an output of a channel state information (CSI)- reference signal (RS) transmitter”.
In the same field of endeavor, Guo teaches “an output of a channel state information (CSI)- reference signal (RS) transmitter”. (Guo, page 547, first column, lines 1-5 teach that CSI estimates are sent from a transmitter and “fed back to the transmitter” and proposes a reduction in feedback overhead to the transmitter using a deep learning model.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Guo to teach the providing output of a CSI encoder. Each of Guo and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Guo with Prasad in order to take advantage of the powerful learning ability and address the difficulty in designing a deep learning model that applies to the physical layer of wireless communications including CSI as taught in Guo, page 547, second column, first paragraph.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 23, Prasad teaches The method of claim 16 as stated. Prasad does NOT teach, further comprising: transmitting, to the UE, a global gradient of an output of a channel state information (CSI)-reference signal (RS) transmitter based on training the ML model; and receiving, within a timer from transmitting the configuration or within a timer from transmitting a global gradient, a local gradient of an output of a channel estimation based on training a UE-side ML model at the UE; and updating, based on the local gradient and other received local gradients, the ML model.
In the same field of endeavor, Zhu teaches transmitting, to the UE, a global gradient of an output of a channel state information [[(CSI)-reference signal (RS) transmitter]] based on training the ML model; (Zhu Fig. 4 and para. [0087] teach in step 401 a terminal receiving a configuration for training based on an AI model for federated learning)
receiving, within a timer from transmitting the configuration or within a timer from transmitting a global gradient, a local gradient of an output of a channel estimation based on training a UE-side ML model at the UE; (Zhu teaches in Fig. 7B, above, receiving within a timer a training gradient associated with the ML model at the UE “within the training duration”:
PNG
media_image4.png
576
733
media_image4.png
Greyscale
and
updating, based on the local gradient and other received local gradients, the ML model. ( Zhu Fig. 7C teaches updating the “average gradient” in block 706 based on other received local gradients for the ML model:
PNG
media_image5.png
696
1167
media_image5.png
Greyscale
In the same field of endeavor, Guo teaches “an output of a channel state information (CSI)- reference signal (RS) transmitter”. (Guo, page 547, first column, lines 1-5 teach that CSI estimates are sent from a transmitter and “fed back to the transmitter” and proposes a reduction in feedback overhead to the transmitter using a deep learning model.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Guo to teach the providing output of a CSI encoder. Each of Guo and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Guo with Prasad in order to take advantage of the powerful learning ability and address the difficulty in designing a deep learning model that applies to the physical layer of wireless communications including CSI as taught in Guo, page 547, second column, first paragraph.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Claims 11, 12 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Zhu further in view of Guo and further in view of 3GPP TS 38.214 V16.8.0 (2021-12) (hereinafter 38.214)
Regarding claim 11, Prasad in view of Zhu and Guo teach The method of claim 10 as stated.
Prasad does NOT teach receiving, from the network node, downlink control information indicating resources for reporting the output of the CSI encoder, wherein training the ML model is based at least in part on a minimum timing gap between receiving the downlink control information and the resources for reporting the output of the CSI encoder.
In the same field of endeavor, Zhu teaches receiving, from the network node, [[downlink control information]] indicating resources for reporting the [[output of the CSI encoder]], wherein training the ML model [[is based at least in part on a minimum timing gap between receiving the downlink control information and the resources]] for reporting the [[output of the CSI encoder]] . (Zhu para. [0087] and Fig. 4 teaches that in Step 401: A base station sends first configuration information to a terminal participating in federated learning, where the first configuration information is used to configure at least one of the following: training duration, a time-frequency resource, or a reporting moment. Correspondingly, the terminal receives the first configuration information from the base station.)
Neither Zhu nor Prasad teach “a DCI … a minimum timing gap between receiving the downlink control information and the resources”.
In the analogous field of 3GPP 5G Wireless Standards, 38.214 teaches “a minimum timing gap between receiving the downlink control information and the resources”. (38.214 teaches in Section 5.2.1 page 50, that CSI reporting is triggered by DCI. Further, Section 5.2.1.5.1 “Triggering/activation of CSI Reports and CSI-RS teaches on page 59 teaches a “scheduling offset” mapped to the timing gap which is “between the last symbol of the PDCCH carrying the triggering DCI and the first symbol of the aperiodic CSI-RS resources”. When smaller than the UE reported threshold beamSwitchTiming… or smaller than 48, then the UE provides a value.)
Neither Prasad nor Zhu teach “reporting the output of the CSI encoder”
In the same field of endeavor, Guo teaches “reporting the output of the (CSI) encoder. (Guo teaches “model-driven deep learning” for CSI feedback as shown in Fig. 1, and encoder and a decoder provide CSI feedback mapped as channel state information encoder output:
PNG
media_image9.png
417
793
media_image9.png
Greyscale
Guo, page 547, column 2, third paragraph teaches that the compress feedback information is fed back to the base station to create a CSI matrix.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Guo to teach the providing output of a CSI encoder. Each of Guo and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Guo with Prasad in order to take advantage of the powerful learning ability and address the difficulty in designing a deep learning model that applies to the physical layer of wireless communications including CSI as taught in Guo, page 547, second column, first paragraph.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
It would have been obvious to one of ordinary skill in the art to combine Prasad and 38.214 to teach a minimum timing gap between receiving the downlink control information and the resources as applied to a machine learning model. Each of Prasad and 38.214 teach CSI reporting. One of ordinary skill would have been motivated to combine Prasad and 38.214 in order to follow the 3GPP established physical layer procedures for data channels for 5G-NR as taught on page 7 of 38.214.
Regarding claim 12, Prasad in view of Zhu and Guo teaches The method of claim 10 as stated. Prasad does NOT teach receiving a first signaling to trigger a semi-persistent reporting of an output of a channel state information (CSI) encoder, wherein the semi-persistent reporting comprises a plurality of reporting occasions; receiving, prior to each reporting occasion a second signaling conveying global gradients; where a timing gap between the second signaling and a reporting occasion is greater than a threshold: updating the ML model based on the global gradients; and
where the timing gap between the second signaling and a corresponding reporting occasion of the plurality of reporting occasions is not greater than a threshold: refraining from updating the ML model.
In the analogous field of 3GPP wireless standards, 38.214 teaches receiving a first signaling to trigger a semi-persistent reporting of [[an output of a channel state information (CSI) encoder]], wherein the semi-persistent reporting comprises a plurality of reporting occasions; (38.214 teaches triggering a semi-persistent CSI reporting in Section 5.2.1.5.2 page 63 that UE CSI report is semi-persistent after a triggering command wherein “When the UE would transmit a PUCCH with HARQ-ACK information in slot n corresponding to the PDSCH carrying the activation command, the indicated semi-persistent Reporting Setting should be applied starting from the first slot that is after slot 𝑛+3𝑁subframe, μ/slot where μ is the SCS configuration for the PUCCH.”)
Prasad nor 38.314 teaches receiving, prior to each reporting occasion a second signaling conveying global gradients.
However, Zhu teaches “receiving, prior to each reporting occasion a second signaling conveying global gradients.” (Zhu Fig. 7C, a second signaling conveying an average gradient of all UE local gradients:
PNG
media_image10.png
278
1267
media_image10.png
Greyscale
.)
and
where the timing gap between the second signaling and a corresponding reporting occasion of the plurality of reporting occasions is not greater than a threshold: refraining from updating the ML model. (Zhu teaches in Fig. 7B block 704 that the training is sent for updating if “within the training duration” mapped to not greater than a threshold:
PNG
media_image4.png
576
733
media_image4.png
Greyscale
Prasad does not teach “an output of a channel state information (CSI) encoder”
In the same field of endeavor, Guo teaches “an output of a channel state information (CSI) encoder”. (Guo teaches “model-driven deep learning” for CSI feedback as shown in Fig. 1, and encoder and a decoder provide CSI feedback mapped as channel state information encoder output:
PNG
media_image9.png
417
793
media_image9.png
Greyscale
Guo, page 547, column 2, third paragraph teaches that the compress feedback information is fed back to the base station to create a CSI matrix.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Guo to teach the providing output of a CSI encoder. Each of Guo and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Guo with Prasad in order to take advantage of the powerful learning ability and address the difficulty in designing a deep learning model that applies to the physical layer of wireless communications including CSI as taught in Guo, page 547, second column, first paragraph.
It would have been obvious to one of ordinary skill in the art to combine Prasad and 38.214 to teach a minimum timing gap between receiving the downlink control information and the resources as applied to a machine learning model. Each of Prasad and 38.214 teach CSI reporting. One of ordinary skill would have been motivated to combine Prasad and 38.214 in order to follow the 3GPP established physical layer procedures for data channels for 5G-NR as taught on page 7 of 38.214.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
Regarding claim 15, Prasad in view Zhu and Guo teaches The method of claim 13 as stated. Prasad does NOT teach receiving a first signaling to trigger a semi-persistent reporting of a local gradient of the ML model based on training the ML model, wherein the semi-persistent reporting comprises a plurality of reporting occasions; receiving, prior to each reporting occasion, a second signaling conveying an output of a channel state information (CSI)-reference signal (RS) transmitter based on training a network-side ML model at the network node; where a timing gap between the second signaling and a corresponding reporting occasion is greater than a threshold: updating the ML model based on the output of the CSI-RS transmitter; and where the timing gap between the second signaling and the corresponding reporting occasion is not greater than a threshold:
refraining from updating the ML model.
In the analogous field of 3GPP 5G wireless communications, 38.214 teaches receiving a first signaling to trigger a semi-persistent reporting of [[a local gradient of the ML model based on training the ML model]], wherein the semi-persistent reporting comprises a plurality of reporting occasions; (38.214 teaches triggering a semi-persistent CSI reporting in Section 5.2.1.5.2 page 63 that UE CSI report is semi-persistent after a triggering command wherein “When the UE would transmit a PUCCH with HARQ-ACK information in slot n corresponding to the PDSCH carrying the activation command, the indicated semi-persistent Reporting Setting should be applied starting from the first slot that is after slot 𝑛+3𝑁subframe, μ/slot where μ is the SCS configuration for the PUCCH.”)
Prasad nor 38.314 teaches that the reporting is a local gradient of the ML model based on training the ML model.
However, Zhu teaches “reporting a local gradient of the ML model based on training the ML model.” (Zhu Fig. 7C, above illustrates reporting a gradient of the ML model based on training the ML model as shown in block 706, above with regard to claim 10.)
Zhu further teaches receiving, prior to each reporting occasion, a second signaling conveying an [[output of a channel state information (CSI)-reference signal (RS) transmitter]] based on training a network-side ML model at the network node; (Zhu Fig. 7, step 706 teaches “deliver the average gradient” after the network side updates a model at a network node)
Neither Prasad nor Zhu teaches that “output of a channel state information (CSI)-reference signal (RS) transmitter”.
However Guo teaches In the same field of endeavor, Guo teaches “an output of a channel state information (CSI)- reference signal (RS) transmitter”. (Guo, page 547, first column, lines 1-5 teach that CSI estimates are sent from a transmitter and “fed back to the transmitter” and proposes a reduction in feedback overhead to the transmitter using a deep learning model.)
Zhu also teaches “where a timing gap between the second signaling and a corresponding reporting occasion is greater than a threshold: updating the ML model based on the output of the CSI-RS transmitter; and where the timing gap between the second signaling and the corresponding reporting occasion is not greater than a threshold: refraining from updating the ML model.” (Zhu para. [0123] teaches that “After the average gradient in the current round of model training is delivered in step 507, the terminal may update the parameter and the gradient of the AI model based on the average gradient in the current round of model training,” the reporting of gradients of the AI model to the base station is done “if the model training is completed within the training duration T.” Thus, if it is not with the threshold T, it is not sent, and if it is within the training duration T, it is sent.)
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Guo to teach the providing output of a CSI encoder. Each of Guo and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Guo with Prasad in order to take advantage of the powerful learning ability and address the difficulty in designing a deep learning model that applies to the physical layer of wireless communications including CSI as taught in Guo, page 547, second column, first paragraph.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
It would have been obvious to one of ordinary skill in the art to combine Prasad and 38.214 to teach a minimum timing gap between receiving the downlink control information and the resources as applied to a machine learning model. Each of Prasad and 38.214 teach CSI reporting. One of ordinary skill would have been motivated to combine Prasad and 38.214 in order to follow the 3GPP established physical layer procedures for data channels for 5G-NR as taught on page 7 of 38.214.
Claim 14 is rejected is rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Zhu and Guo further in view of US Pat. Pub. 20240388383 to Hao Tang et al. (hereinafter Tang)
Regarding claim 14, Prasad in view of Zhu and Guo teaches The method of claim 13 as stated.
Prasad does NOT teach receiving, from the network node, downlink control information indicating resources for reporting the local gradient, wherein training the ML model is based at least in part on a minimum timing gap between receiving the downlink control information and the resources for reporting the local gradient.
In the same field of endeavor, Tang teaches receiving, from the network node, downlink control information indicating resources for reporting the local gradient, wherein training the ML model is [[based at least in part on a minimum timing gap between receiving the downlink control information and the resources for reporting the local gradient]].(Tang teaches in para. [0141] that “some specific fields in the DCI could be set as predefined values to indicate training activation” including a resource indicator. Tang also teaches gradient calculation in para. [0099] describing AI neural network calculations.)
Neither Tang nor Prasad teach “based at least in part on a minimum timing gap between receiving the downlink control information and the resources for reporting the local gradient”.
However, 38.214 teaches “based at least in part on a minimum timing gap between receiving the downlink control information and the resources for reporting the local gradient”. (38.214 teaches in Section 5.2.1 page 50, that CSI reporting is triggered by DCI. Further, Section 5.2.1.5.1 “Triggering/activation of CSI Reports and CSI-RS teaches on page 59 teaches a “scheduling offset” mapped to the timing gap which is “between the last symbol of the PDCCH carrying the triggering DCI and the first symbol of the aperiodic CSI-RS resources”. When smaller than the UE reported threshold beamSwitchTiming… or smaller than 48, then the UE provides a value.)
It would have been obvious to one of ordinary skill in the art to combine Prasad and 38.214 to teach a minimum timing gap between receiving the downlink control information and the resources as applied to a machine learning model. Each of Prasad and 38.214 teach CSI reporting. One of ordinary skill would have been motivated to combine Prasad and 38.214 in order to follow the 3GPP established physical layer procedures for data channels for 5G-NR as taught on page 7 of 38.214.
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Tang to teach downlink control information indicating resources for reporting the local gradient. Each of Prasad and Tang are in the field of wireless communication and AI. One of ordinary skill in the art would have been motivated to combine Prasad and Tang in order to increase the reliability of adaptation for artificial intelligence training as taught in para. [0002] of Tang.
Claim 21 is rejected is rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Zhu further in view of Tang
Regarding claim 21, Prasad teaches The method of claim 16 as stated. Prasad does NOT teach, further comprising: transmitting downlink control information indicating resources for reporting an aperiodic local training gradient associated with training the ML model; and receiving the aperiodic local training gradient associated with training the ML model at the UE over the resources.
In the same field of endeavor, Tang teaches transmitting downlink control information indicating resources for reporting an aperiodic local training gradient associated with training the ML model. (Tang teaches in para. [0141] that “some specific fields in the DCI could be set as predefined values to indicate training activation” including a resource indicator. Tang also teaches gradient calculation in para. [0099] describing AI neural network calculations. Tang para. [0167] teaches that a UE may send information aperiodically.)
and
Prasad teaches receiving the aperiodic [[local training gradient]] associated with training the ML model at the UE over the resources (Prasad para. [0032] teaches “ when frequent feedback consumes significant radio resources, the gNB may, at 215, configure aperiodic CSI reports from the UE, which may be triggered by certain threshold conditions being satisfied.”)
However Prasad does not teach that the feedback is a “local training gradient”.
In the same field of endeavor, Zhu teaches a local training gradient as shown in Fig. 7B:
PNG
media_image4.png
576
733
media_image4.png
Greyscale
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Zhu with Prasad to teach a timing limit for ML. Each of Zhu and Prasad are in the field of wireless communications. One of ordinary skill in the art would have been motivated to combine Prasad and Zhu in order to address problems such as excessively high time-frequency resource usage and large delay caused by deployed federated learning as taught in Zhu para. [0004].
It would have been obvious to one of ordinary skill in the art prior to the effective date of the invention to have combined Prasad with Tang to teach downlink control information indicating resources for reporting the local gradient. Each of Prasad and Tang are in the field of wireless communication and AI. One of ordinary skill in the art would have been motivated to combine Prasad and Tang in order to increase the reliability of adaptation for artificial intelligence training as taught in para. [0002] of Tang.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARGARET MARIE ANDERSON whose telephone number is (703)756-1068. The examiner can normally be reached M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CHARLES JIANG can be reached at 571-270-7191. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARGARET MARIE ANDERSON/Examiner, Art Unit 2412 /CHARLES C JIANG/Supervisory Patent Examiner, Art Unit 2412