Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 16-21, 23, 25-31, 33 and 35 have been amended. Claims 1-15, 22 and 32 have been canceled. Claims 16-21, 23-31, and 33-35 are currently pending and have been considered by the Examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 19 and 29 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 19 is rendered indefinite for the following reasons. Lines 2-3 recite “splitting at least a portion of the at least the portion of the pre-trained second neural network”, and the second neural network could include only one portion. It is unclear if this limitation means breaking an existing portion of the second neural network into smaller pieces. It is unclear if this limitation means separating a portion from the rest of the portions of the second neural network. If so, it is unclear how the second neural network could be split where there is only one portion in the second neural network. The limitation in lines 4-5 is unclear for similar reason. Examiner treats claim 19 to mean separating a portion from the rest of the second neural network.
Claim 29 recites the same indefinite limitations as claim 19 and is therefore rejected for at least the same reasons.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 16-17, 20, 23, 26-27, 30, and 33 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Hall et al. (US 20220344049 A1, cited in the PTO-892 issued 05/18/2026).
Regarding claim 16, Hall teaches: A method performed by a target device, the method comprising: receiving, from a device on a network via an encoded signal, metadata associated with a pre-trained first neural network, wherein the metadata comprises predefined information of portions of a training process of the pre-trained first neural network performed at the device, wherein the device on the network is a different device than the target device; ([0005], lines 1-5; [0105], lines 1-3; [0122], lines 1-7; [0126] and [0132] discloses training a teacher model at a node by optimizing its weights, encrypting the weights, and sending encrypted weights to a central node. A “target device” is central node 40, a “device on a network” is a node 11, “metadata associated with a pre-trained first neural network” and “predefined information” are weights of a trained teacher model M1. Since the teacher’s weights may be encrypted before sending them to the central node, the signal containing the encrypted weights is an encoded signal.)
obtaining at least a portion of a pre-trained second neural network from memory at the target device; ([0062], [0122], lines 1-10, [0124] on page 11, col. 1, lines 29-34 (“sending a Teacher model to become a Student in another locality”), and Fig. 1A discloses applying a student model MS at the central node. The student model may be a previous-trained teacher model such as M2 which was trained at another locality, transferred to memory at the central node, and then obtained from the memory for distillation training. A “portion” as claimed includes any nodes or edges of M2 now acting as the student model.)
adapting the at least the portion of the pre-trained second neural network at the target device based on the predefined information of the portions of the training process of the pre-trained first neural network received from the device on the network at which the training of the first neural network was performed; and ([0018] and [0122], lines 7-10 discloses training the student model MS on the trained teacher model M1 using a distillation training technique. Adapting includes training the student model MS at the central node.)
performing an inference during deployment of the adapted at least the portion of the pre-trained second neural network. ([0122], lines 10-12)
Regarding claim 17, Hall teaches: The method of claim 16, wherein said metadata comprises at least one of: at least one batch size, at least one optimizer, at least one drop-out, at least one learning rate, a designation or a parameter of at least one loss function, at least one performance indicator related to an accuracy of training, at least one indicator related to an importance of at least one weight of at least one layer in the pre-trained first neural network or the at least the portion of the pre-trained second neural network, at least one type of information related to pre-processing performed on at least one element of a training set, or at least one type of information representative of at least one position inside the pre-trained first neural network where a prediction can be made. ([0037]-[0038] and [0115] in col. 2, lines 8-14 discloses a trained teacher model comprises layers. The final layer is used to make a prediction. Trained weights of a final layer correspond to “metadata” and “at least one type of information representative of at least one position inside the pre-trained first neural network where a prediction can be made.”)
Regarding claim 20, Hall teaches: The method of claim 16, wherein said adapting comprises pre-processing at least a part of a training data set based on the metadata, ([0125], lines 1-10 discloses sourcing a transfer dataset for training the student model. Since the data is permitted to be used and accessed by the trained teacher model (the metadata) during the distillation method, sourcing is based on the metadata.)
wherein the training data set is configured for use during a training process for fine-tuning the at least the portion of the pre-trained second neural network. ([0125], final 3 lines, where “fine-tuning” is training the student model.)
Regarding claim 23, Hall teaches: The method of claim 16, wherein the adapting comprises picking a sub-part of the at least the portion of the pre-trained second neural network or selecting a set of parameter settings for the at least the portion of the pre-trained second neural network based on the metadata. ([0005], lines 1-5 and [0122], lines 7-10 discloses selecting a set of parameter settings (student weight values) for the at least the portion of the pre-trained second neural network based on the metadata.)
Regarding claim 26, Hall teaches: A target device comprising: a transceiver; ([0122], lines 1-7 and [0136], lines 1-5 discloses a central node 40 can send and receive data. The central node is a target device comprising a transceiver.)
a memory; and a processor configured to: ([0282], lines 4-5 and 19-end disclose a memory and processors)
receive, via the transceiver from a device on a network via an encoded signal, metadata associated with a pre-trained first neural network, wherein the metadata comprises predefined information of portions of a training process of the pre-trained first neural network performed at the device, wherein the device on the network is a different device than the target device; ([0005], lines 1-5; [0105], lines 1-3; [0122], lines 1-7; [0126] and [0132] discloses training a teacher model at a node by optimizing its weights, encrypting the weights, and sending encrypted weights to a central node. A “target device” is central node 40, a “device on a network” is a node 11, “metadata associated with a pre-trained first neural network” and “predefined information” are weights of a trained teacher model M1. Since the teacher’s weights may be encrypted before sending them to the central node, the signal containing the encrypted weights is an encoded signal.)
obtain at least a portion of a pre-trained second neural network from the memory; ([0062], [0122], lines 1-10, [0124] on page 11, col. 1, lines 29-34 (“sending a Teacher model to become a Student in another locality”), and Fig. 1A discloses applying a student model MS at the central node. The student model may be a previously-trained teacher model such as M2 which was trained at another locality, transferred to memory at the central node, and then obtained from the memory for distillation training. A “portion” as claimed includes any nodes or edges of M2 now acting as the student model.)
adapt the at least the portion of the pre-trained second neural network at the target device based on the predefined information of the portions of the training process of the pre-trained first neural network received from the device on the network at which the training of the first neural network was performed; and ([0018] and [0122], lines 7-10 discloses training the student model MS on the trained teacher model M1 using a distillation training technique. Adapting includes training the student model MS at the central node.)
perform an inference during deployment of the adapted at least the portion of the pre-trained second neural network. ([0122], lines 10-12)
Claims 27, 30 and 33 each recites a product which implements the same features as the method of claims 17, 20 and 23, respectively, and are therefore rejected for at least the same reasons.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 18-19 and 28-29 are rejected under 35 U.S.C. 103 as being unpatentable over Hall et al. (US 20220344049 A1, cited in the PTO-892 issued 05/18/2026) in view of Mun et al. (US 20190088251 A1, cited in PTO-892 issued 01/28/2026).
Regarding claim 18, Hall teaches: The method of claim 16, wherein adapting the at least the portion of the pre-trained second neural network comprises [training]
However, Hall does not explicitly teach: compressing or pruning the at least the portion of the pre-trained second neural network
But Mun teaches: compressing or pruning the at least the portion of the server 520, and reducing (“pruning”) the combined model down to the personalization layer. The limitation of “the at least the portion of the… second neural network” is a combination of the personalization layer and the global model)
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have applied Mun’s pruning techniques to Hall. A motivation for the combination is to learn a model where some features are unique to a predetermined user and other features are shared by multiple users. (Mun, [0063], lines 1-7)
Regarding claim 19, Hall teaches: The method of claim 16, wherein the adapting comprises: [retraining] 2 now acting as a student model. The retraining is a distillation training technique based on the teacher model M1. The feature of “metadata” is the teacher model M1.)
However, Hall does not explicitly teach: splitting at least a portion of the at least the portion of the pre-trained second neural network
transmitting at least one split portion of the at least the portion of the pre-trained second neural network.
But Mun teaches: splitting at least a portion of the at least the portion of the rest of the second neural network. [0088], lines 1-11 and [0104], from line 9 to the end of the paragraph discloses a neural network comprising a personalization layer and a global model. An embodiment includes training a combination of a personalization layer and global model at a server 520, splitting the personalization layer from the global model, and transmitting the personalization layer back to device 510.)
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have applied Mun’s training, splitting, and transmitting techniques to Hall, where Mun’s neural network corresponds to Hall’s portions of the student model. A motivation for the combination is the same as the motivation given for claim 18.
Claims 28-29 each recites a product which implements the same features as the method of claims 18-19, respectively, and are therefore rejected for at least the same reasons.
Claims 21 and 31 are rejected under 35 U.S.C. 103 as being unpatentable over Hall et al. (US 20220344049 A1, cited in the PTO-892 issued 05/18/2026) in view of Bloom (US 20180336463 A1, cited in PTO-892 issued 09/18/2025).
Regarding claim 21, Hall teaches: The method of claim 16, wherein [0116], lines 1-6 discloses a trained teacher neural network can have multiple layers as depicted in Fig. 1E. Each convolution layer “conv” and fully-connected layer “FC” has weights.)
However, Hall does not explicitly teach: the metadata is received as compressed metadata that comprises compression of the metadata
wherein the target device comprises a decoder, and wherein the method comprises decoding the compressed metadata at the decoder.
But Bloom teaches: the metadata is received as compressed metadata that comprises compression of the metadata ([0054], lines 1-19 and Fig. 5. The claim limitation of “metadata” is input data 540, “compressed metadata” is encoded data 550, and receiving compressed metadata is the second ML model component 530 receiving encoded data 550.)
wherein the target device comprises a decoder, and wherein the method comprises decoding the compressed metadata at the decoder. ([0054], lines 1-19 discloses decoder 514 comprises the layers of second ML model component 530)
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated Bloom’s encoder portion into Hall’s node 11 for encrypting the trained teacher model, and to have incorporated Bloom’s decoder portion into Hall’s central node 40 for decrypting the encrypted, trained teacher model. Bloom’s input data 540 corresponds to Hall’s trained teacher model, and Bloom’s encoded data 550 corresponds to Hall’s encrypted teacher model. A motivation for the combination is to obscure data for transport and prevent an attacker from reconstructing the original data. (Bloom, [0022])
Claim 31 recites a product which implements the same features as the method of claim 21 and is therefore rejected for at least the same reasons.
Claims 24 and 34 are rejected under 35 U.S.C. 103 as being unpatentable over Hall et al. (US 20220344049 A1, cited in the PTO-892 issued 05/18/2026) in view of Sikka et al. (US 20210012212 A1, cited in PTO-892 issued 09/18/2025).
Regarding claim 24, Hall teaches: The method of claim 16, further comprising:
Hall teaches a user interface at [0267], final 5 lines. However, Hall does not explicitly teach: providing a portion of the metadata to a user on a user interface; and receiving an indication from the user in response to the portion of the metadata being provided, and wherein the adapting is performed based on the indication from the user.
But Sikka teaches: providing a portion of the metadata to a user on a user interface; and (Fig. 4, [0073], lines 1-6 and [0074], lines 5-end, where “a portion of the metadata” includes textual or graphical descriptions of a neural network.)
receiving an indication from the user in response to the portion of the metadata being provided, and ([0075])
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to allow a user to modify Hall’s teacher model using Sikka’s technique. In the combination of references, Hall’s student model would be trained based on the modified teacher model, and thus the combination teaches “the adapting is performed based on the indication from the user.” A motivation for the combination is that Sikka’s method gives a user finer control over how a neural network is modified. ([0075])
Claim 34 recites a product which implements the same features as the method of claim 24 and is therefore rejected for at least the same reasons.
Claims 25 and 35 are rejected under 35 U.S.C. 103 as being unpatentable over Hall et al. (US 20220344049 A1, cited in the PTO-892 issued 05/18/2026) in view of Alakuijala et al. (US 20190251444 A1, cited in PTO-892 issued 09/18/2025).
Regarding claim 25, Hall teaches: The method of claim 16, wherein the metadata indicates
However, Hall does not explicitly teach: an importance of one or more weights
But Alakuijala teaches: an importance of one or more weights (All of [0027] and [0067], lines 1-2)
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have estimated an edge utility of Hall’s teacher model as part of the metadata. A motivation for the combination is to measure an amount by which the edge contributes to correct predictions provided by the neural network. (Alakuijala, [0027])
Claim 35 recites a product which implements the same features as the method of claim 25 and is therefore rejected for at least the same reasons.
Response to Arguments
Below is the Examiner’s response to the Applicant’s arguments filed 08/18/2026.
Applicant’s Arguments Under 35 U.S.C. 102: On page 8 of the remarks, the Applicant argues that Claim 26 has been amended to clarify that the at least a portion of the second neural network obtained is pre-trained and that the at least the portion of the pre-trained second neural network is adapted. Applicant argues that the Office Action has not established that the art discloses pre-training at least a portion of a second neural network, much less adapting the at least the portion of the pre-trained second neural network at the target device based on the predefined information of the portions of the training process of the pre-trained first neural network received from the device on the network at which the training of the first neural network was performed.
Examiner’s Response: Applicant's arguments have been fully considered but they are not persuasive. The previous Office Action did not establish that the Hall reference teaches the second neural network being “pre-trained” because this feature was absent in the previously-considered claim 26. Hall at paragraphs [0062], [0122], lines 1-10, [0124] on page 11, col. 1, lines 29-34 (“sending a Teacher model to become a Student in another locality”), and Fig. 1A discloses applying a student model MS at the central node. The student model may be a previously-trained teacher model such as M2 which was trained at another locality, transferred to memory at the central node, and then obtained from the memory for distillation training. A “portion” as claimed includes any nodes or edges of M2 now acting as the student model.
Hall discloses the limitations of claim 26, lines 10-13 at [0018] and [0122], lines 7-10. Hall teaches training the student model MS on the trained teacher model M1 using a distillation training technique. Adapting includes training the student model MS at the central node.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Asher H. Jablon whose telephone number is (571)270-7648. The examiner can normally be reached Monday - Friday, 9:00 am - 6:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Al Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.H.J./Examiner, Art Unit 2127
/ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127