Prosecution Insights
Last updated: October 02, 2026
Application No. 18/954,233

SYSTEMS AND METHODS FOR FEATURE DROPOUT KNOWLEDGE DISTILLATION

Non-Final OA §103
Filed
Nov 20, 2024
Priority
Nov 21, 2023 — provisional 63/601,508
Examiner
KOLB JR, BRETT DAVID
Art Unit
Tech Center
Assignee
Datum Point Labs Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 20240256831 A1; Hereinafter referred to as Li) and further in view of Helmrich et al. (US 20210409708 A1, hereinafter referred to as Helmrich) and Sinha et al. (US 20230094954 A1, hereinafter referred to as Sinha). Regarding Claim 1, Li teaches A method of knowledge distillation (“The knowledge from the first model can be distilled or passed”; See Paragraph 70) from a teacher model (First Model; See Paragraph 63; The first model is a generative model teaching the student model.) to a student model (Second Model; See Paragraph 31; “The second model can be referred to as a student”), comprising: Encoding (“The knowledge from the first model can be distilled or passed to the encoder”; See Paragraph 70), via a first encoder (Encoder; See Paragraph 5; “The second model can include at least one of an encoder or a decoder.”; The second model can include multiple encoder and decoders.), a feature (Second Features; See Paragraph 71; “The second features can include extracted representations or tensors”; The Second features coming from the second model (Student model).) of the student model (Second Model) to provide first principal components (extracted representations or tensors; See Paragraph 71; “The second features can include extracted representations or tensors”); Encoding (“While the at least one image is generated or while the at least one real image is being encoded”; See Paragraph 64), via a second encoder (“the generative model is a combination of encoder and generator, such as Variational Autoencoder (VAE)”; See paragraph 63; Both The teacher model (Generative model) and the Student model (Second Model) includes an encoder.), a feature (First features) of the teacher model (First Model; See Paragraph 63-64; “the first features, from the generative model”; The first model is a generative model teaching the student model. ) to provide second principal components (extracted representations or tensors; See paragraph 64; “The First features can include extracted representations or tensors”); Li does not teach decoding, via a first decoder, the first principal components to provide first decoded components. Li also does not teach decoding, via a second decoder, the second principal components to provide second decoded components. Helmrich teaches decoding, via a first decoder ("FIGS. 1 to 3 have been presented as an example where the inventive concept described further below may be implemented in order to form specific examples for encoders and decoders"; See para 49; Helmrich teaches using multiple decoders in variations of their inventive concept ("encoders may comprise a functionality corresponding to claimed decoders"; See Paragraph 112) and by saying you can combine embodiments (“ In addition, features of the different embodiments described hereinafter may be combined with each other, unless specifically noted otherwise.”; See Paragraph 26.)), the first principal components (" However, this fixed approach was found to yield relatively uneven distribution of said coding gain across the two processed component signals. To compensate for this issue, a more general rotation-based approach, realized using a size-two KLT also known as principal component analysis (PCA), may be pursued."; See paragraph 90; Implies two Principal components because they are doing principal component analysis on two processed component signals (The two Principal components).) to provide first decoded components (decoded first and second components; See table 7); Helmrich also teaches decoding, via a second decoder ("FIGS. 1 to 3 have been presented as an example where the inventive concept described further below may be implemented in order to form specific examples for encoders and decoders"; See para 49; Helmrich teaches using multiple decoders in variations of their inventive concept ("encoders may comprise a functionality corresponding to claimed decoders"; See Paragraph 112) and by saying you can combine embodiments (“ In addition, features of the different embodiments described hereinafter may be combined with each other, unless specifically noted otherwise.”; See Paragraph 26.)), the second principal components ("However, this fixed approach was found to yield relatively uneven distribution of said coding gain across the two processed component signals. To compensate for this issue, a more general rotation-based approach, realized using a size-two KLT also known as principal component analysis (PCA), may be pursued."; See paragraph 90; Implies two Principal components because they are doing principal component analysis on two processed component signals (The two Principal components).) to provide second decoded components (decoded first and second components; See table 7); It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified Li to incorporate the teaching of Helmrich’s method of using two decoders with Li’s already preestablished two encoders to provide primary and secondary decoded components to achieve computation of logits for loss calculation. Li further teaches Computing first logits (“P.sub.τ.sup.r is the logit determined by the second model”; See Paragraph 80; This is a variable in a formula or a piece of code which would be ran with all the features through it to make multiple logits.) via the student model (Second Model); Computing second logits from the teacher model (First Model) (“the interpreter receives the first features (e.g., the multi-level features) outputted from the generator as input and feeds the first features into a series of Feature Fusion Layers (FFLs) to lower the feature dimension and fuse with the next-level features, to output per-pixel logits.”; See Paragraph 78); The combination of Li, Helmrich, and Sinha teach this limitation of the claim as this limitation makes this claim a Markush claim. Li teaches computing a loss (“The training system can determine the loss (associated with the second features) with respect to the first features. The loss can include one or more of attention loss, feature regression loss, knowledge distillation loss, softmax loss, and so on. In some examples, the overall or total loss can be calculated to be the sum or combination of one or more of the types of loss. For example, the overall loss for a feature can be determined using the following expression…”; See Paragraph 75; They teach using many types of loss and using different features and attributes to determine it.) based on at least one of: Li does not teach but Helmrich teaches a comparison of the first principal components to the second principal components ("the variance of one of the resulting components becomes small compared to the variance of the other component"; See paragraph 68; They teach comparison of first and second principal components through various methods of analysis.); a comparison of the first decoded components to the feature of the student model; a comparison of the second decoded components to the feature of the teacher model; or and Li does not teach but Sinha teaches a comparison of the first logits to the second logits ("In some instances, a discriminator component of the trained machine-learning model generates a first logit value representing the set of predicted socio-demographic attributes and a second logit value representing the set of training socio-demographic attributes. Then, the discriminator determines a loss (e.g., a classification loss) between the two logit values."; See paragraph 105;); Li teaches and updating parameters of the student model based on the loss (“The training system can update the second model using the loss”; See paragraph 82). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified Li to incorporate the teachings of Helmrich and Sinha. The motivation to incorporate Helmrich and Sinha into Li is to use their comparisons for loss calculations into Li’s loss calculations not only to cover many different types of calculating data loss but to also update the parameters of the student model based on that loss to improve the calculations of the student model. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRETT DAVID KOLB JR whose telephone number is (571)270-0751. The examiner can normally be reached Monday-Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /B.D.K./Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Nov 20, 2024
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month