Prosecution Insights
Last updated: October 01, 2026
Application No. 18/745,564

APPARTUS AND METHOD FOR DISTILLING KNOWLEDGE FOR FLOW PREDICTION MODEL

Non-Final OA §101§103§112
Filed
Jun 17, 2024
Priority
Jul 03, 2023 — RE 10-2023-0086090
Examiner
ZENG, WENWEI
Art Unit
Tech Center
Assignee
Research & Business Foundation Sungkyunkwan University
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
23 currently pending
Career history
18
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on April 14, 2026, and June 17, 2024, were considered by the examiner. The submissions are compliant with the provisions of 37 CFR 1.97. Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. KR-10-2023-0086090, from the Republic of Korea, filed on July 3, 2023. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are from claim 1, “a student model former configured to…”, “a weight generator configured to generate…”, “a function generator configured to generate…”, and “and a knowledge distiller configured to distill…”, and corresponding structures are found in specification on page 8, lines 2-10. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1, 7, and 13 recite the limitation “the plurality of prediction results”, however it is unclear if these prediction results refer to “a plurality of hierarchical prediction results” or other type of results. There is insufficient antecedent basis for this limitation in the claims. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Claims 1-13 are rejected under 35 U.S.C. 101 since the claimed invention is directed to an abstract idea (a mental process or math concept) without significantly more. Claim 1: Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “1. An apparatus for distilling knowledge for a scene flow prediction model, the apparatus comprising: a controller comprising: a student model former configured to form a student model to have single bidirectional flow embedding and a flow predictor structure of a teacher model; a weight generator configured to generate a weight based on a plurality of hierarchical prediction results of the teacher model and predetermined ground truth data; a function generator configured to generate a loss function by using the weight and the plurality of prediction results; and a knowledge distiller configured to distill the knowledge of the teacher model to the student model by using the loss function,” and an apparatus or machine is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept or mental process but for recitation of generic computer …to generate a weight based on a plurality of hierarchical prediction results of the teacher model and predetermined ground truth data; (This recites a mathematical relationship, formula or equation, or mathematical calculation, see page 10, lines 6-9, from the specification state “The weight generator 120 obtains a differentiation between a hierarchical prediction result of the flow predictor (FP) of the teacher model and ground truth data (data close a correct answer), and applies the differentiation to the inversed Softmax function to generate the weight,” showing using an inversed softmax function, which recites math to generate weights, see MPEP 2106.04(a)(2), subsection I), …to generate a loss function by using the weight and the plurality of prediction results; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in page 9, lines 8- 11, from the specification stating “The function generator 130 multiplies each of a plurality of prediction results of the flow predictor (FP) for each layer of the teacher model by the weight, and adds multiplication results to each other to calculate the loss function. Here, the loss function may be defined as Attentive Mixed Loss (AML)” , showing math operations are performed to calculate a loss function, also see specification on page 11, lines 11-17 describe a loss function is a calculated value using math operations, PNG media_image1.png 496 1272 media_image1.png Greyscale , see MPEP 2106.04(a)(2), subsection I), to distill the knowledge of the teacher model to the student model by using the loss function, (This recites a mathematical relationship, formula or equation, or mathematical calculation, see specification on page 11, lines 11-17 from above, describe a loss function is a calculated value using math operations, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: An apparatus for distilling knowledge for a scene flow prediction model, the apparatus comprising: a controller comprising: … (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)), a student model former configured to form a student model to have single bidirectional flow embedding and a flow predictor structure of a teacher model; (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), a weight generator configured…, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), a function generator configured… (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), and a knowledge distiller configured… (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements v, vi, vii, and viii recite mere instructions to apply the judicial exception using generic computer components and additional element iv recites using generic computer components as a tool, which are not indicative of significantly more. Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claim 2: Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Claim 2 recites the following abstract idea: ...calculates a differentiation between the plurality of hierarchical prediction results of the teacher model and the ground truth data, and inputs the calculated differentiation into a predetermined inversed Softmax function to generate the weight, (This recites a mathematical relationship, formula or equation, or mathematical calculation, see page 9, lines 3-7 in specification stating "The weight generator 120 calculates a differentiation between a plurality of prediction results of a flow predictor (FP) for each hierarchy of a teacher model and ground truth data, and inputs the calculated differentiation into a predetermined inversed Softmax function to generate a weight. Here, the differentiation may be a value closer to 1 as the difference between the prediction result and the ground truth data is smaller," showing a calculation performed to obtain a differentiation value which is a number, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 2 recites the following additional element: 2. The apparatus of claim 1, wherein the weight generator…, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 3: Regarding claim 3, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis applied to claim 2. Further, claim 3 recites the following abstract idea: ... multiplies the plurality of respective hierarchical prediction results of the teacher model by the weight, and adds multiplication results to each other to calculate the loss function, (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in page 9, lines 8 -11 from the specification stating “The function generator 130 multiplies each of a plurality of prediction results of the flow predictor (FP) for each layer of the teacher model by the weight, and adds multiplication results to each other to calculate the loss function. Here, the loss function may be defined as Attentive Mixed Loss (AML)” , showing math operations are performed to calculate a loss function, also see specification on page 11, lines 11-17 describe a loss function is a calculated value using math operations, PNG media_image1.png 496 1272 media_image1.png Greyscale , see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 3 recites the following additional element: 3. The apparatus of claim 2, wherein the function generator ..., (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 4: Regarding claim 4, it is dependent upon claim 3 and thereby incorporates the limitations of, and corresponding analysis applied to claim 3. Further, claim 4 recites the following abstract idea: … combines the loss function with a predetermined multi scale loss function to generate a training loss function for the student model, (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in page 9, lines 12-16, from specification, stating “The function generator 130 combines the loss function with a predetermined multi scale loss (MSL) function to finally generate the loss function. The multi scale loss (MSL) function may be acquired by a plurality of layers of a conventional model having different scales. Hereinafter, a final loss function may be defined as a training loss function.” Combining involves math operations since the multi scale loss function involve different scales or values, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 4 recites the following additional element: 4. The apparatus of claim 3, wherein the function generator ... (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 5: Regarding claim 5, it is dependent upon claim 4 and thereby incorporates the limitations of, and corresponding analysis applied to claim 4. Further, claim 5 recites the following abstract idea: ... to distill the knowledge of the teacher model to the student model by using the training loss function, (This recites a mathematical relationship, formula or equation, or mathematical calculation, see specification on page 11, lines 11-17 PNG media_image1.png 496 1272 media_image1.png Greyscale , describe a loss function is a calculated value using math operations, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 5 recites the following additional element: 5. The apparatus of claim 4, wherein the knowledge distiller... (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 6: Regarding claim 6, it is dependent on claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 6 recites the following additional element: The apparatus of claim 1, wherein the teacher model has bidirectional flow embedding and a flow structure of an iteration performing structure, (In step 2A, prong 2, this is considered mere instructions to apply an exception using a generic computer – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using a generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 7: Regarding claim 7, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites a “7. A method for distilling knowledge for a scene flow prediction model, the method comprising: …; ”, and a method is one of the four statutory categories of invention. Further, since claim 7 recites similar limitations as corresponding independent claim 1 listed above, it is rejected for similar reasons under 35 U.S.C. 101. Claims 8-12: Since claims 8-12 recite similar limitations as corresponding claims 2-6 listed above, and are rejected for similar reasons under 35 U.S.C. 101. Claim 13: Regarding claim 13, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “13. A non-transitory computer readable medium containing program instructions executed by a processor, the computer readable medium comprising: program instructions that … ”, and a non-transitory computer readable medium or machine is one of the four statutory categories of invention. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: “13. A non-transitory computer readable medium containing program instructions executed by a processor, the computer readable medium comprising: program instructions that … ” (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element i recites mere instructions to apply the judicial exception using generic computer components, which is not indicative of significantly more. Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Further, since claim 13 recites similar limitations as corresponding independent claim 1 listed above, it is rejected for similar reasons under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 6, 7, 12, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Li, Z., et al., in "Deep Learning for Scene Flow Estimation on Point Clouds: A Survey and Prospective Trends," published on April 3, 2023, available at: https://onlinelibrary.wiley.com/doi/full/10.1111/cgf.14795 , (hereafter, Li), in view of Yu, M. et al., in PG Pub. No. CN115953643A, published on April 11, 2023, (hereafter, Yu), further in view of Liu, Y., et al., in "Adaptive multi-teacher multi-level knowledge distillation," published on July 28, 2020, available at https://www.sciencedirect.com/science/article/pii/S0925231220311565 , (hereafter, Liu). Claim 1: Regarding claim 1, Li teaches “a student model former configured to form a student model to have single bidirectional flow embedding and a flow predictor structure of a teacher model; See Li in page 8 , section Bi-PointFlowNet note “Built upon successful bidirectional learning in time series-based tasks and 2D optical flow estimation, Bi-PointFlowNet [CK22] develops the first bidirectional model for 3D scene flow estimation. Bi-PointFlowNet targets at estimating the optimal non-rigid transformation that represents the best alignment from the source to the target frame. Previous standard procedure (i.e. grouping -> concatenation -> MLP -> max-pooling) usually leads to redundant computations. To address this issue, Bi-PointFlowNet decomposes the MLP weights in bidirectional flow embedding layer into three sub-weights. In this way, the local coordinates, the propogated feature, and the replicated feature of two point clouds can be transformed to produce a new fused feature vector. The following upsampling and warping layer are the same as PointPWC-Net. Compared to PointPWC-Net [WWL*20], Bi-PointFlowNet reduces the total operation by 44% and accelerates the inference by 33%.” Here, Li mentions using the concept of knowledge distillation of teacher model and student model for a bidirectional flow embedding layer. Also, see Li describe in page 17, section 7.5 Knowledge distillation note “Collecting large-scale dynamic scene data requires complex calibration. In addition, the cost of transforming the original data into a trainable format is expensive. As a consequence, a labelled dataset for scene flow estimation is very rare. Therefore, applying the knowledge distillation model for training small scene flow estimation networks would be a possible solution for data-hungry networks. In machine learning, knowledge distillation (KD) is the process of compressing the knowledge in a large model into a smaller one. As shown in Figure 7, the traditional knowledge distillation model consists of a teacher model and a student model. In many proposed deep learning models, there are often heavy parameters. Although it's commonly accepted that integrating multiple models and introducing more parameters improves the accuracy of a model, we have to bear high computational costs in the meanwhile [MFL*20]. KD allows training smaller models with minimal loss in performance. The main innovation of KD is that the student network is trained not only via the information provided by true labels but also by observing how the teacher network works with the data. To our best knowledge, DCA-SRSFE [JLA*22] is the only method that applies the KD model to point-based scene flow estimation so far.” Here, Li shows using knowledge distillation which brings a teacher network model to later distill the data that the teacher model has learned onto the student model. The student model learns the single bidirectional flow embedding and how to make predictions from the teacher The data used in this case is scene flow estimation, which is using bidirectional flow embeddings according to page 8 section Bi-PointFlowNet, describing using a model called “Bi-PointFlowNet decomposes the MLP weights in bidirectional flow embedding layer.” Further, see Li in page 2, Difference in point density, mention "A LiDAR system identifies the position of the light energy returns from a target to the LiDAR sensor." Here, Li shows using a LIDAR system, which includes a computer to perform operations such as forming student models from knowledge distillation. Examiner construes a student model former to be any part of a computer system that builds models for knowledge distillation, including teacher models and student models. Further, Li teaches “ a weight generator configured to generate a weight based on a plurality of hierarchical prediction results of the teacher model and predetermined ground truth data;” See Li in page 2, Difference in point density, mention "A LiDAR system identifies the position of the light energy returns from a target to the LiDAR sensor." Here, Li shows using a LIDAR system, which includes a computer to perform operations such as calculating weights for the teacher models. Further, see Li in page 3, section 3.2 PointConv feature pyramid note “PointConv layer is proposed to learn point features hierarchically. The PointConv method involves inputting the positions of point clouds and training an MLP to estimate a weight function. The method also involves applying an inverse density scale to the learned weights…” Here, Li shows weights are generated. Li also explicitly mentions learning each point of the scene flow hierarchically and relates to hierarchical results. Further, see Li in page 7, section 5. Methodology, sub section PointPWC-Net, note "PointConvFormer [WSF22] modifies the feature learning mechanism via transformers. It explores the computation of convolutional weights, leveraging the difference in features between points to recalculate the convolutional weights. Additionally, PointConvFormer uses a sigmoid activation for the attention weights that outperformed the use of softmax." Here, Li shows that the method calculates weights, and the PointConvFormer is considered a weight generator. Within page 7 and section 5, Li also notes the method "PointPWC-Net [WWL*20] that predicts scene flow via constructing the cost volume at each feature pyramid level." Here, Li further elaborates that the PointPWC-Net can make scene flow predictions at each feature pyramid level (which is interpreted as hierarchical level here). Since PointConvFormer and PointPWC-Net are part of the same system that Li describes, these all work together to generate weight using hierarchical prediction results of a model and data. Also, see Li describe in page 12, in section DCA-SRSFE note that “Jin et al. [JLA*22] proposed a mean-teacher framework for unsupervised domain adaptation from synthetic data to real data. DCA-SRSFE [JLA*22] consists of a student model that uses ground-truth scene flow labels for supervision and a teacher model updated as the Exponential Moving Average (EMA) of the student model weights. A deformation regularization module and a correspondence refinement module are introduced to produce high-quality pseudo labels. In the deformation regularization module, a rigid motion between the first point cloud and the warped point cloud is predicted via Kabsch algorithm [Kab76]. This module encourages shape distortion awareness in the student model and promotes adaptive deformations for the target domain. The flow vector is later improved with surface correspondence by refining local geometry. DCA-SRSFE is supervised by ground truth flow labels in the source domain and trained with a consistency loss over the target domain. The proposed synthetic dataset GTA-SF is a large-scale dataset with real-world labels. According to the experiments, DCA-SRSFE has narrowed down the performance gap between synthetic datasets and real-world scenarios.” Here, Li shows using a teacher model and provide knowledge on how to make inferences onto the student model. When Li mentions the student model uses ground-truth scene flow labels for supervision, Li shows that the weights are generated using prediction results from the updated teacher model, and the ground-truth scene flow labels (i.e. relate to ground truth data). Li also describes that the updated teacher model trains a student model to learn in predicting results. When Li mentions ‘a rigid motion between the first point cloud and the warped point cloud is predicted’, the hierarchical results in this case are also considered predicted results. See Li in page 8 in FH-Net note “FH-Net [DDX*22] deals with multi-scale flows from different layers with a much faster speed. To this end, FH-Net extracts keypoint features via hierarchical Trans-flow layer. The computed sparse flow is then used to obtain hierarchical flows at different resolutions through an inverse Trans-up layer.” Here, Li mentions in page 8 that the system handles image data at different resolutions, where different resolutions relate to hierarchical prediction since hierarchical means to process images at multiple levels of resolution from specification on page 1, lines 20-22 state “The bi-point flow net is constituted by four hierarchical structures in order to predict the scene flow. The four hierarchical structures can include hierarchical feature extraction, bidirectional flow embedding, upsampling, and a flow predictor.” Further, Li teaches “a function generator configured to generate a loss function by using … the plurality of prediction results;” See Li describe in page 8, first paragraph, from section 5. Methodology, sub section SCTN, “Additionally, SCTN proposes a feature-aware spatial consistency loss to improve its ability to distinguish different motion fields.” Here, Li mentions producing a loss value to better differentiate motion scenes of image data, where consistency loss relates to being part of a loss function, and for one of the skill in the art, the consistency loss is added directly into a total loss function. Also, see Li mention in pages 7-8, section 5. Methodology, sub section SCTN “SCTN [LZGG22] introduces a voxel based convolution to produce consistent flows in 3D space. SCTN uses a combination of sparse convolution for feature extraction and a transformer module for accurate scene flow prediction. It is the first work to incorporate the transformer with sparse convolution, which allows it to learn relation-based contextual information on point clouds.” Here, Li further solidified using the prediction results produced from the SCTN method onto creating a consistency loss or loss function. SCTN here is considered a function generator. However, Li did not teach “1. An apparatus for distilling knowledge for a scene flow prediction model, the apparatus comprising: a controller comprising:…” or “a function generator configured to generate a loss function by using the weight …;” or “and a knowledge distiller configured to distill the knowledge of the teacher model to the student model by using the loss function.” In an analogous art, Yu teaches “1. An apparatus for distilling knowledge for a scene flow prediction model, the apparatus comprising: a controller comprising:…” See Yu in paragraph [n0161] describe “Another embodiment of this application also provides an electronic device. As shown in Figure 10, the electronic device 1000 includes a memory 1001 and a processor 1002; the memory 1001 and the processor 1002 are coupled; the memory 1001 is used to store computer program code, which includes computer instructions. When the processor 1002 executes computer instructions, it causes the electronic device 900 to perform each step of the method.” Note the examiner construes a controller to be a part of a general computer system, which includes a memory and processor as part of the system. Here, Yu shows that the electronic device 1000 corresponds to a controller. Further, Yu teaches “and a knowledge distiller configured to distill the knowledge of the teacher model to the student model by using the loss function,” See Yu in paragraph [n0054] describes a “server 230 extracts output information from terminal 210 and output information from terminal 220, matching the student model's output information with the teacher model's output information and calculating the output distillation loss; server 230 improves the student model's loss function based on the calculated feature distillation loss function and output distillation loss function, and guides the training of the student model based on the improved loss function, so that the trained student model has object detection capabilities similar to the teacher model.” Here, Yu shows the system builds a loss function then use this loss to train the student model to match the information provided by the teacher model. This relates to using a server 230 (i.e. knowledge distiller) to use a teacher model to train and distill knowledge with a loss function to a student model. Yu describes a system that updates the student model so its behavior aligns with that of the teacher model. Examiner construes knowledge distiller to be part of a computing system that can run knowledge distillation operations using teacher and student models. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Li and incorporate into the teachings of Yu because both references teach knowledge distillation in scene flow prediction. One of ordinary skill in the art would be motivated to do so because “this example implementation provides a complete heterogeneous distillation scheme, which allows the use of a higher-performance two-stage detection network as a teacher model during the training of the target detection model. This improves the performance ceiling of distillation, expands the range of teacher selection, and broadens the application scope of the distillation algorithm,” (see Yu in [n0144] ). However, Li in view of Yu, did not teach “a function generator configured to generate a loss function by using the weight …;” Further, Liu teaches “a function generator configured to generate a loss function by using the weight …;” See Liu in page 106, Introduction mention “associate a latent representation with each teacher to indicate its characteristics and adaptively determine the importance weights of different teachers with respect to a specific instance based on both the teacher representation and instance representation gotten from a student model. These learned weights are leveraged for the weighted combination of soft-targets from corresponding teachers.” Here, Liu mentions using weights of various teacher models and incorporate these by importance weighting in predictions. Also, see Liu in page 108 section 3.1 Motivation and overview, note "The adapter is responsible for adaptively learning instance-level teacher importance weights for integrating soft-targets, which are further utilized to derive the standard knowledge distillation loss Lkd and angle based loss ℒangle. ℒ kd and ℒangle help to learn the weighted dark knowledge and structural relation between examples [respectively]. Moreover, a multi-group hint based loss ℒHT is for transferring intermediate-level knowledge from multiple teacher layers. The overall optimization target of the proposed AMTML-KD is given as: ... PNG media_image2.png 67 866 media_image2.png Greyscale .” Here, Liu shows creating a loss function of Loss ℒHT, from both ℒkd and ℒangle by using weights and the soft targets (i.e. relate to prediction results) from each teacher model. Further, see Liu in page 112, from section 5.3 Ablation study, last paragraph of section mention "We can see that even if some teachers’ predictions are wrong, the adverse effects of these mistakes can be shielded by weight allocation, such as shown in Fig. 4, Fig. 4(b)," Here, Liu shows generating and then assigning the weights for each teacher model to evaluate if the teacher model's predictions have the correct predictions or are close to ground-truth results. Liu shows that the predictions are used in this method. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Li and Yu with the teachings of Liu by using the teachings by Li and Yu, and incorporate with Liu’s teaching of generating weights and loss function for the teacher model. One of ordinary skill in the art would be motivated to do so because by integrating Liu’s framework into the methods of Li and Yu, one with ordinary skill in the art would achieve the goal of providing “since the teacher network is only involved in the training procedure and does not take up any resource for deployment and use, it is effective and efficient for applying the learned student model into practice. It has aroused some successful applications in computer vision,” (see Liu in page 106, 1. Introduction ). Claim 6: Regarding claim 6, Li in view of Yu, and further in view of Liu, teach the limitations of claim 1. Further, Li teaches “6. The apparatus of claim 1, wherein the teacher model has bidirectional flow embedding…” See Li in page 8 , section Bi-PointFlowNet note “Built upon successful bidirectional learning in time series-based tasks and 2D optical flow estimation, Bi-PointFlowNet [CK22] develops the first bidirectional model for 3D scene flow estimation. Bi-PointFlowNet targets at estimating the optimal non-rigid transformation that represents the best alignment from the source to the target frame. Previous standard procedure (i.e. grouping -> concatenation -> MLP -> max-pooling) usually leads to redundant computations. To address this issue, Bi-PointFlowNet decomposes the MLP weights in bidirectional flow embedding layer into three sub-weights. In this way, the local coordinates, the propogated feature, and the replicated feature of two point clouds can be transformed to produce a new fused feature vector. The following upsampling and warping layer are the same as PointPWC-Net. Compared to PointPWC-Net [WWL*20], Bi-PointFlowNet reduces the total operation by 44% and accelerates the inference by 33%.” Here, Li mentions using the concept of knowledge distillation of teacher model and student model for a bidirectional flow embedding layer. Also, see Li describe in page 17, section 7.5 Knowledge distillation note “Collecting large-scale dynamic scene data requires complex calibration. In addition, the cost of transforming the original data into a trainable format is expensive. As a consequence, a labelled dataset for scene flow estimation is very rare. Therefore, applying the knowledge distillation model for training small scene flow estimation networks would be a possible solution for data-hungry networks. In machine learning, knowledge distillation (KD) is the process of compressing the knowledge in a large model into a smaller one. As shown in Figure 7, the traditional knowledge distillation model consists of a teacher model and a student model. In many proposed deep learning models, there are often heavy parameters. Although it's commonly accepted that integrating multiple models and introducing more parameters improves the accuracy of a model, we have to bear high computational costs in the meanwhile [MFL*20]. KD allows training smaller models with minimal loss in performance. The main innovation of KD is that the student network is trained not only via the information provided by true labels but also by observing how the teacher network works with the data. To our best knowledge, DCA-SRSFE [JLA*22] is the only method that applies the KD model to point-based scene flow estimation so far.” Here, Li shows using knowledge distillation which brings a teacher network model to later distill the data that the teacher model has learned onto the student model. The student model learns the single bidirectional flow embedding and how to make predictions from the teacher The data used in this case is scene flow estimation, which is using bidirectional flow embeddings according to page 8 section Bi-PointFlowNet, describing using a model called “Bi-PointFlowNet decomposes the MLP weights in bidirectional flow embedding layer.” Further, see Li describe in page 3 PNG media_image3.png 608 806 media_image3.png Greyscale Here, Li describes a calculation of a flow embedding. However, Li in view of Yu, did not teach “the teacher model has … a flow predictor structure of an iteration performing structure,” Further, Liu teaches “the teacher model has … a flow predictor structure of an iteration performing structure,” See Liu in page 3 in section Notation mention “For our teacher model, we denote I1, I2 ∈ RH× W× 3 for two consecutive RGB images, where H and W are height and width respectively. Our goal is to estimate a forward optical flow wf ∈ RH× W× 2 from I1 to I2. After obtaining wf, we can warp I2 towards I1 to get a warped image Iw 2 . Here, we also estimate a backward optical flow wb from I2 to I1 and a backward warp image Iw 1 . Since there are many cases where one pixel is only visible in one image but not visible in the other image, namely occlusion, we denote Of, Ob ∈ RH× W× 1 as the forward and backward occlusion map respectively.” Here, Liu shows bidirectional flow by using a forward flow and a backward flow in optical images. Also, see Liu in page 6 in Training procedure note “After that, we initialize the student model with the weights from our teacher model, and train both the teacher model (with Lp) and the student model (with Lp + Lo) together for 300k iterations.” Here, Liu shows the method of using bidirectional flow from page 3 to predict optical flow data, which is repeated 300k times in this method. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Li and Yu with the teachings of Liu by using the teachings by Li and Yu, and incorporate with Liu’s teaching of a flow predictor structure of an iteration performing structure. One of ordinary skill in the art would be motivated to do so because by integrating Liu’s framework into the methods of Li and Yu, one with ordinary skill in the art would achieve the goal of providing “since the teacher network is only involved in the training procedure and does not take up any resource for deployment and use, it is effective and efficient for applying the learned student model into practice. It has aroused some successful applications in computer vision,” (see Liu in page 106, 1. Introduction ). Claim 7: Regarding claim 7, it comprises of similar additional limitations as independent claim 1, and is rejected under the same rationale under 35 U.S.C. 103. Claim 12: Regarding claim 12, it comprises of similar additional limitations as corresponding claim 6, and is rejected under the same rationale under 35 U.S.C. 103. Claim 13: Regarding claim 13, regarding the limitation “A non-transitory computer readable medium containing program instructions executed by a processor,…” See Yu in [n0028] note “Fourthly, this application provides a computer-readable storage medium storing computer instructions that, when executed on an electronic device, cause the electronic device to perform the knowledge distillation-based model training method as described in the first aspect and any possible design thereof.” Here, Yu mentions a computer -readable storage medium that helps store instructions that runs a physical device, and relates to a non-transitory computer readable medium. Further, see Yu in [n0003] note “Therefore, this application proposes a model training method, apparatus, electronic device and storage medium based on knowledge distillation to solve the series of heterogeneous distillation problems mentioned above.” Yu shows using the device for knowledge distillation. Also, see Yu in paragraph [n0161] describe "Another embodiment of this application also provides an electronic device. As shown in Figure 10, the electronic device 1000 includes a memory 1001 and a processor 1002; the memory 1001 and the processor 1002 are coupled; the memory 1001 is used to store computer program code, which includes computer instructions. When the processor 1002 executes computer instructions, it causes the electronic device 900 to perform each step of the method." Note the examiner construes a controller to be a part of a general computer system, which includes a physical memory and processor as part of the system. Here, Yu shows that the electronic device 1000 corresponds to a controller. Referring to claim 13, the claim comprises of similar additional limitations as independent claim 1, and is rejected under the same rationale under 35 U.S.C. 103. Claims 2, 3, 4, 5, and 8, 9, 10, 11 are rejected under 35 U.S.C. 103 as being unpatentable over Li, in view of Yu, further in view of Liu, further in view of Passban P. et al., in US PG Pub. No. US20220076136A1, published on March 10, 2022, (hereafter, Passban), and further in view of Peng, B. et al., in "Hierarchical Dense Correlation Distillation for Few-Shot Segmentation," published from June 17-24, 2023 for a conference, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10204860&tag=1 , (hereafter, Peng). Claim 2: Regarding claim 2, Li in view of Yu, further in view of Liu, teach the limitations of claim 1. However, Li in view of Yu, further in view of Liu, did not teach “2. The apparatus of claim 1, wherein the weight generator calculates a differentiation between the plurality of hierarchical prediction results of the teacher model and the ground truth data, and inputs the calculated differentiation into a predetermined inversed Softmax function to generate the weight.” In an analogous art, Passban teaches teach “2. The apparatus of claim 1, wherein the weight generator calculates a differentiation between the plurality of hierarchical prediction results of the teacher model and the ground truth data, and inputs the calculated differentiation into a predetermined … Softmax function to generate the weight,” See Passban note in [0088] “ Where the output inferred (i.e. predicted) by the neural network model are penalized with its own standard loss against the ground truth but also a loss against the output generated by hidden layers of the teacher model given by q(y=ν|x; θT), which may be referred to as the KD loss. In Equation (2), the first component of the loss function or the KD loss, namely q((y=ν|x; θT)), is often referred to as the soft loss as it is compared to the inferences (i.e. predictions) (also known as soft labels) made by a softmax function of the teacher model. The remaining loss components, such as the standard loss, are known as the hard loss. Accordingly, in typical KD approaches, the overall loss function includes at least two loss terms, namely the standard loss and the KD loss, accordingly to Equation (3): PNG media_image4.png 56 333 media_image4.png Greyscale ” Here, Passban shows comparing the results of the predictions of the teacher model, weighing this by a weight factor alpha and 1-alpha. The predicted results are compared to ground truth, which then produces a loss value called the KD loss to account for the difference, which shows a differentiation between the plurality of hierarchical prediction results of the teacher model and the ground truth data. Further, See Passban describe in [0042-0043] shows generating the weights: PNG media_image5.png 314 863 media_image5.png Greyscale Here, Passban shows how attention weights are produced from viewing the difference of predictions vs. actual outputs of the teacher model and student model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Li, Yu, and Liu, with the teachings of Passban by using the teachings by Li, Yu, and Liu,, and incorporate with Passban’s teaching of a differentiation between the plurality of hierarchical prediction results of the teacher model and the ground truth data. One of ordinary skill in the art would be motivated to do so because by integrating Passban’s framework into the methods of Li, Yu, and Liu, one with ordinary skill in the art would achieve a method like “the CKD* method may improve upon the performance for BERT and neural machine translation models”, (see Passban in [0132]). However, Li in view of Yu, further in view of Liu, and further in view of Passban, did not teach “ a... inversed Softmax function” In an analogous art, Peng teaches “ a... inversed Softmax function”, See Peng in page 23645, section 3.3 note “where Wq, Wk∈Rc×c denote the learnable parameters, ∥⋅∥ indicates L2 norm, and t is a hyperparameter to control the distribution range, empirically set to 0.1 in all experiments. Inspired by [28], [33], we propose the inverse softmax layer that normalizes the correlation matrix along the query axis... PNG media_image6.png 965 1064 media_image6.png Greyscale ” see Figure 2 for details. Here, Peng shows using an inverse SoftMax on the correlation step, which is connected to adapting a classifier's weights or producing weights. Later, see Peng in page 23642, in Section 2, Related work, Transformer and Knowledge distillation sections describe "In [23], the classifier weight transformer adapts the classifier's weights to address the intra-class variation issue. CyCTR [49] is a cycle-consistent transformer by generating query and key sequences from the query and support set, respectively. Transformer architecture helps FSS transcend the limitation of semantic-level prototypes and leverage pixel-wise alignment. Previous transformer-based methods are still difficult to handle noise interference and over-fitting. We, in this paper, propose a new transformer structure decoupling the downsampling and matching processes and design the matching module constructed on correlation mechanism and distillation...KD attempts to transfer learned knowledge from the large model (a.k.a. the teacher) to another light model (a.k.a. the student) with tolerable loss in performance." Here, Peng shows using a transformer architecture in adapting a classifier's weights or is interpreted to producing the weights, which are constructed based on correlation and knowledge distillation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Li, Yu, Liu, and Passban with the teachings of Peng by using the teachings by Li, Yu, and Liu,, and incorporate with Peng’s teaching of using an inverse softmax. One of ordinary skill in the art would be motivated to do so because by integrating Peng’s framework into the methods of Li, Yu, Liu, and Passban, one with ordinary skill in the art would achieve a goal that “processes a self-distillation framework to distill knowledge within itself to improve model accuracy,” (see Peng in page 23642, Knowledge distillation section). Claim 3: Regarding claim 3, Li in view of Yu, further in view of Liu, further in view of Passban, and further in view of Peng, teach the limitations of claim 2. Further, Passban teaches “3. The apparatus of claim 2, wherein the function generator multiplies the plurality of respective hierarchical prediction results of the teacher model by the weight, and adds multiplication results to each other to calculate the loss function,” See Passban in [0047] note "In some examples, the energy function Φ(hi S,hj T) is computed as the dot product of the output of the ith hidden layer of the student (hi S) and the weighted value of the output generated by the jth hidden layer of the teacher (hj T) by: PNG media_image7.png 60 349 media_image7.png Greyscale ” Here, Passban shows the weighted values (i.e. weights) of the prediction outputs of the teacher model calculated by dot product, which is construed to mean one form of multiplication operation. Further, see Passban note in [0088] “ Where the output inferred (i.e. predicted) by the neural network model are penalized with its own standard loss against the ground truth but also a loss against the output generated by hidden layers of the teacher model given by q(y=ν|x; θT), which may be referred to as the KD loss. In Equation (2), the first component of the loss function or the KD loss, namely q((y=ν|x; θT)), is often referred to as the soft loss as it is compared to the inferences (i.e. predictions) (also known as soft labels) made by a softmax function of the teacher model. The remaining loss components, such as the standard loss, are known as the hard loss. Accordingly, in typical KD approaches, the overall loss function includes at least two loss terms, namely the standard loss and the KD loss, accordingly to Equation (3): PNG media_image4.png 56 333 media_image4.png Greyscale ” Here, Passban shows comparing the results of the predictions of the teacher model, weighing this by a weight factor alpha and 1-alpha. The predicted results are compared to ground truth, which then produces a loss value called the KD loss to account for the difference, which shows a differentiation between the plurality of hierarchical prediction results of the teacher model and the ground truth data. Later, see Passban in [0134] describe “ PNG media_image8.png 187 839 media_image8.png Greyscale ” where Passban describes adding the multiplication results to calculate the loss function Lqkd It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Li, Yu, and Liu, with the teachings of Passban by using the teachings by Li, Yu, and Liu,, and incorporate with Passban’s teaching of a system that multiplies the respective prediction results of the teacher model by the weight, and adds multiplication results to each other to calculate the loss function. One of ordinary skill in the art would be motivated to do so because by integrating Passban’s framework into the methods of Li, Yu, and Liu, one with ordinary skill in the art would achieve a method like “the CKD* method may improve upon the performance for BERT and neural machine translation models”, (see Passban in [0132]). Claim 4: Regarding claim 4, Li in view of Yu, further in view of Liu, further in view of Passban, and further in view of Peng, teach the limitations of claim 3. Further, Passban teaches “4. The apparatus of claim 3, wherein the function generator combines the loss function with a predetermined multi scale loss function to generate a training loss function for the student model,” See Passban also note in [0149] “Referring back to FIG. 10, at step 1008, the weight KD losses L q KD for all K teachers 204 are summed up as the total KD loss of the teachers 204 at each time step. It is to be understood that although embodiments described herein includes LKD loss, in some other embodiments, depending on the design of each problem, training loss functions other than L KD may also be applicable.” Further, see Passban mention in abstract that this overall loss is “An agnostic combinatorial knowledge distillation (CKD) method for transferring trained knowledge of neural model from a complex model (teacher) to a less complex model (student) is described. In addition to training the student to generate a final output that approximates both the teacher's final output and a ground truth of a training input, the method further maximizes knowledge transfer by training hidden layers of the student to generate outputs that approximate a representation of a subset of teacher hidden layers are mapped to each of the student hidden layers for a given training input.” Passban shows overall loss is from a standard loss and the KD loss, which relates to being part of a multi-scale loss. Then, Passban mentions that this loss is used in training a teacher model to a student model from abstract (i.e. relates to generating a training loss function) for the student model. By transferring trained knowledge, Passban refers to the teacher model teaching a student model. Examiner construes multi-scale loss to mean a loss value combined from multiple sources such as using the standard loss and the KD loss. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Li, Yu, and Liu, with the teachings of Passban by using the teachings by Li, Yu, and Liu,, and incorporate with Passban’s teaching of a system that generates a training loss function for the student model. One of ordinary skill in the art would be motivated to do so because by integrating Passban’s framework into the methods of Li, Yu, and Liu, one with ordinary skill in the art would achieve a method like “the CKD* method may improve upon the performance for BERT and neural machine translation models”, (see Passban in [0132]). Claim 5: Regarding claim 5, Li in view of Yu, further in view of Liu, further in view of Passban, and further in view of Peng, teach the limitations of claim 4. Further, Passban teaches “5. The apparatus of claim 4, wherein the knowledge distiller distills the knowledge of the teacher model to the student model by using the training loss function,” See Passban also note in [0149] “Referring back to FIG. 10, at step 1008, the weight KD losses L KD q for all K teachers 204 are summed up as the total KD loss of the teachers 204 at each time step. It is to be understood that although embodiments described herein includes LKD loss, in some other embodiments, depending on the design of each problem, training loss functions other than L KD may also be applicable.” Further, see Passban mention in abstract that this overall loss is “An agnostic combinatorial knowledge distillation (CKD) method for transferring trained knowledge of neural model from a complex model (teacher) to a less complex model (student) is described. In addition to training the student to generate a final output that approximates both the teacher's final output and a ground truth of a training input, the method further maximizes knowledge transfer by training hidden layers of the student to generate outputs that approximate a representation of a subset of teacher hidden layers are mapped to each of the student hidden layers for a given training input.” Here, Passban emphasizes that knowledge is distilled, or the teacher model teaches the student model to produce a final result using a teacher model's knowledge so the student also predicts results the same the teacher model does using the loss function described in paragraph [0149]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Li, Yu, and Liu, with the teachings of Passban by using the teachings by Li, Yu, and Liu,, and incorporate with Passban’s teaching of a system that distills knowledge from a teacher model to the student model. One of ordinary skill in the art would be motivated to do so because by integrating Passban’s framework into the methods of Li, Yu, and Liu, one with ordinary skill in the art would achieve a method like “the CKD* method may improve upon the performance for BERT and neural machine translation models”, (see Passban in [0132]). Claim 8: Regarding claim 8, it comprises of similar additional limitations as claim 2, and is rejected under the same rationale under 35 U.S.C. 103. Claim 9: Regarding claim 9, it comprises of similar additional limitations as claim 3, and is rejected under the same rationale under 35 U.S.C. 103. Claim 10: Regarding claim 10, it comprises of similar additional limitations as claim 4, and is rejected under the same rationale under 35 U.S.C. 103. Claim 11: Regarding claim 11, it comprises of similar additional limitations as claim 5, and is rejected under the same rationale under 35 U.S.C. 103. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WenWei Zeng/Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Jun 17, 2024
Application Filed
Sep 21, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month