Prosecution Insights
Last updated: October 02, 2026
Application No. 18/559,003

NEURAL NETWORK LEARNING APPARATUS, NEURAL NETWORK LEARNING METHOD, AND PROGRAM

Non-Final OA §103
Filed
Nov 03, 2023
Priority
May 17, 2021 — nonprovisional of PCTJP2021018588
Examiner
TRAN, TAN H
Art Unit
Tech Center
Assignee
Nippon Telegraph and Telephone Corporation
OA Round
1 (Non-Final)
61%
Grant Probability
Moderate
1-2
OA Rounds
7m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
195 granted / 320 resolved
+0.9% vs TC avg
Strong +33% interview lift
Without
With
+32.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 6m
Avg Prosecution
46 currently pending
Career history
374
Total Applications
across all art units

Statute-Specific Performance

§101
13.4%
-26.6% vs TC avg
§103
59.8%
+19.8% vs TC avg
§102
16.5%
-23.5% vs TC avg
§112
6.3%
-33.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 320 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 2. This action is in response to the original filing on 11/03/2023. Claims 4, 7, and 11 are pending and have been considered below. Election/Restrictions 3. Claims 1-3, 5-6, and 8-10 are withdraw from further consideration pursuant to 37 CFR 1.142(b) as being drawn to nonelected group I. Election was made without traverse in the reply filed on 6/23/2016. Information Disclosure Statement 4. The information disclosure statement (IDS(s)) submitted on 11/03/2023 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections – 35 USC § 103 5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 6. Claims 4, 7, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Tran (Learning to Make Predictions on Graphs with Autoencoders, arXiv, published 2018, pages 1-9) in view of Saini et al. (U.S. Patent Application Pub. No. US 20190333400 A1). Claim 4: Tran teaches a neural network learning device that performs learning of a neural network (i.e. We implement the autoencoder architecture using Keras [4] on top of the GPU-enabled TensorFlow; Section IV.B, page 5) including an encoder that converts an input vector into a latent variable vector having a latent variable as an element and a decoder that converts the latent variable vector into an output vector such that the input vector and the output vector are substantially identical to each other (i.e. Let ai∈RN be an adjacency vector of A that contains the local neighborhood of the ith node. Our proposed autoencoder architecture comprises a set of non-linear transformations on ai summarized in two component parts: encoder g(ai):RN→RD, and decoder f(zi):RD→RN. We stack two layers of the encoder part to derive D-dimensional latent feature representation of the ith node zi∈RD, and then stack two layers of the decoder part to obtain an approximate reconstruction output a^i∈RN, resulting in a four-layer autoencoder architecture; Section II, pages 2-3), the neural network learning device comprising a learning circuitry configured to perform learning by repeating parameter update processing of updating parameters included in the neural network (i.e. During the forward pass, or inference, the model takes as input an adjacency vector ai and computes its reconstructed output a^i=h(ai) for link prediction. The parameters θ are learned via backpropagation. During the backward pass, we estimate θ by minimizing the Masked Balanced Cross-Entropy (MBCE) loss, which only allows for the contributions of those parameters associated with observed edges … For reconstruction and link prediction, we train for 50 epochs using mini-batch size of 8 samples … We employ the Adam algorithm [13] for gradient descent optimization with a fixed learning rate of 0.001; Section II, IV, pages 3-5), wherein the encoder, when each of pieces of input information included in a predetermined input information group corresponds to one of three ways of positive information, negative information, and no information (i.e. For a partially observed graph, A∈{1,0,UNK}N×N, where 1 denotes a known present positive edge, 0 denotes a known absent negative edge, and UNK denotes an unknown status (missing or unobserved) edge; Section II, page 2), inputs an input vector that represents each of the pieces of input information by a positive information bit that is 1 in a case where the input information corresponds to positive information, and is 0 in a case where there is no information or in a case where the input information corresponds to negative information (i.e. For a partially observed graph, A∈{1,0,UNK}N×N, where 1 denotes a known present positive edge, 0 denotes a known absent negative edge, and UNK denotes an unknown status (missing or unobserved) edge … We impute missing or UNK elements in the adjacency matrix with 0; Section II, IV, pages 2, 5), and the encoder includes a plurality of layers (i.e. Let ai∈RN be an adjacency vector of A that contains the local neighborhood of the ith node. Our proposed autoencoder architecture comprises a set of non-linear transformations on ai summarized in two component parts: encoder g(ai):RN→RD, and decoder f(zi):RD→RN. We stack two layers of the encoder part to derive D-dimensional latent feature representation of the ith node zi∈RD, and then stack two layers of the decoder part to obtain an approximate reconstruction output a^i∈RN, resulting in a four-layer autoencoder architecture; Section II, pages 2-3), a layer having the input vector as an input obtains a plurality of output values from the input vector (i.e. Let ai∈RN be an adjacency vector of A that contains the local neighborhood of the ith node. Our proposed autoencoder architecture comprises a set of non-linear transformations on ai summarized in two component parts: encoder g(ai):RN→RD, and decoder f(zi):RD→RN. We stack two layers of the encoder part to derive D-dimensional latent feature representation of the ith node zi∈RD, and then stack two layers of the decoder part to obtain an approximate reconstruction output a^i∈RN, resulting in a four-layer autoencoder architecture zi=g(ai)=ReLU(W⋅ReLU(Vai+b(1))+b(2)); Section II, page 2), and the parameter update processing is performed such that a value of a loss function is smaller (i.e. During the forward pass, or inference, the model takes as input an adjacency vector ai and computes its reconstructed output a^i=h(ai) for link prediction. The parameters θ are learned via backpropagation. During the backward pass, we estimate θ by minimizing the Masked Balanced Cross-Entropy (MBCE) loss, which only allows for the contributions of those parameters associated with observed edges … For reconstruction and link prediction, we train for 50 epochs using mini-batch size of 8 samples … We employ the Adam algorithm [13] for gradient descent optimization with a fixed learning rate of 0.001; Section II, IV, pages 3-5), the loss function including a sum for all pieces of input information of an input information group for learning of loss that has (i.e. we compute the MBCE loss as follows: LBCE=−ailog(σ(a^i))⋅ζ−(1−ai)log(1−σ(a^i)),LMBCE=mi⊙LBCE∑mi; Section II, page 3), aggregates the masked per entry losses to form L, in a case where the input information corresponds to positive information, a larger value as a probability that input information obtained by the decoder corresponds to positive information is smaller (i.e. we compute the MBCE loss as follows: LBCE=−ailog(σ(a^i))⋅ζ−(1−ai)log(1−σ(a^i)),LMBCE=mi⊙LBCE∑mi; Section II, page 3), this loss increases as the decoder’s predicted probability of a positive edge decreases, and in a case where the input information corresponds to negative information, a larger value as a probability that input information obtained by the decoder corresponds to negative information is smaller (i.e. we compute the MBCE loss as follows: LBCE=−ailog(σ(a^i))⋅ζ−(1−ai)log(1−σ(a^i)),LMBCE=mi⊙LBCE∑mi; Section II, page 3), the loss increases as that negative state probability decreases, and in a case where the input information does not exist, a value of substantially zero (i.e. we compute the MBCE loss as follows: LBCE=−ailog(σ(a^i))⋅ζ−(1−ai)log(1−σ(a^i)),LMBCE=mi⊙LBCE∑mi. Here, LBCE is the balanced cross-entropy loss with weighting factor ζ=1−# present links# absent links,σ(⋅) is the sigmoid function, ⊙ is the Hadamard (element-wise) product, and mi is the boolean function: mi=1 if ai≠UNK, else mi=0 … note that imputed edges are not observed elements in the adjacency matrix and hence do not contribute to the masked loss computations during training; Section II, IV, pages 3-5), assigns a mask value of zero when an adjacency entry is unknown, because the entry’s BCE loss is multiplied by that zero valued mask, missing input information contributes zero to the MBCE training loss. Tran does not explicitly teach a negative information bit that is 1 in a case where the input information corresponds to negative information, and is 0 in a case where there is no information or in a case where the input information corresponds to positive information, each of the output values is obtained by adding together all of values of positive information bits included in the input vector to which weight parameters are respectively given and values of negative information bits included in the input vector to which weight parameters are respectively given. However, Saini teaches when each of pieces of input information included in a predetermined input information group corresponds to one of three ways of positive information, negative information, and no information (i.e. Once a learner attempts a question (or takes a hint), the user's response can be encoded and provided to co-learning model 200 to update the learner's knowledge state. For example, three outcomes for a particular interaction with a question (correct response, incorrect response, hint taken) can be encoded in various ways, as will be understood by those of ordinary skill in the art. In some embodiments, one-hot encoding is used to encode (qt, rt, ht) into a vector based on a set of distinct questions Q; para. [0034] and Table 1), inputs an input vector that represents each of the pieces of input information by a positive information bit that is 1 in a case where the input information corresponds to positive information, and is 0 in a case where there is no information or in a case where the input information corresponds to negative information (i.e. one-hot encoding is used to encode (qt, rt, ht) into a vector based on a set of distinct questions Q. In some embodiments, the first |Q| dimensions of the encoding are a one-hot vector representing a correct attempt on the question. For example, in the case of a correct attempt, the vector can have 1 at the index for the question and 0 everywhere else. Similarly, the next |Q| dimensions can encode an incorrect attempt. In some embodiments, hint-taking (or use of some other learning aid) may be similarly one-hot encoded. However, the accuracy of co-learning model 200 may be improved by encoding hint-taking as a binary value representing hint-taking across different questions. This indicates that it is important to know how many times a learner has taken a hint, but the identity of the questions on which the hints are taken does not impact the two prediction tasks. As such, a particular interaction is advantageously encoded into a vector of length 2|Q|+1. Table 1 provides an example of such input encoding where there are a total of two questions, Q1 and Q2; para. [0034] and Table 1), and a negative information bit that is 1 in a case where the input information corresponds to negative information, and is 0 in a case where there is no information or in a case where the input information corresponds to positive information (i.e. one-hot encoding is used to encode (qt, rt, ht) into a vector based on a set of distinct questions Q. In some embodiments, the first |Q| dimensions of the encoding are a one-hot vector representing a correct attempt on the question. For example, in the case of a correct attempt, the vector can have 1 at the index for the question and 0 everywhere else. Similarly, the next |Q| dimensions can encode an incorrect attempt. In some embodiments, hint-taking (or use of some other learning aid) may be similarly one-hot encoded. However, the accuracy of co-learning model 200 may be improved by encoding hint-taking as a binary value representing hint-taking across different questions. This indicates that it is important to know how many times a learner has taken a hint, but the identity of the questions on which the hints are taken does not impact the two prediction tasks. As such, a particular interaction is advantageously encoded into a vector of length 2|Q|+1. Table 1 provides an example of such input encoding where there are a total of two questions, Q1 and Q2; para. [0034] and Table 1), each of the output values is obtained by adding together all of values of positive information bits included in the input vector to which weight parameters are respectively given and values of negative information bits included in the input vector to which weight parameters are respectively given (i.e. Generally, the learner's most recent interaction 255 xt=(qt, rt, ht) is encoded, and provided as an input. This encoding can be converted into response embedding 257 vt using response embedding matrix B: v t =B*x t; para. [0032, 0034, 0040]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Tran to include the feature of Saini. One would have been motivated to make this modification because it would preserve the distinction between negative and missing information during encoder processing and permit the neural network to learn separate weights for positive and negative information. Claims 7 and 11 are similar in scope to Claim 4 and are rejected under a similar rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Macready et al. (Pub. No. US 20210089884 A1), Ratings may comprise binary ratings (e.g., thumbs-up/thumbs-down), categorical ratings with more than two possible values (e.g., a rating out of five stars), and/or continuous-valued ratings (e.g., a percentage value). The ratings may be provided via a graphical user interface either explicitly (e.g., by selecting a rating from a group of possible ratings) and/or implicitly (e.g., by tracking which items a user interacts with, such as by clicking on them). It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)). Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TAN H TRAN/Primary Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Nov 03, 2023
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748960
Analog Hardware Realization of Neural Networks
5y 7m to grant Granted Sep 29, 2026
Patent 12718079
Systems and Methods for Generating Libraries for Hardware Realization of Neural Networks
5y 5m to grant Granted Aug 25, 2026
Patent 12718088
DESIGNING LADDER AND LAGUERRE ORTHOGONAL RECURRENT NEURAL NETWORK ARCHITECTURES INSPIRED BY DISCRETE-TIME DYNAMICAL SYSTEMS
4y 8m to grant Granted Aug 25, 2026
Patent 12688413
METHODS FOR RELIABLE OVER-THE-AIR COMPUTATION WITH PULSES FOR DISTRIBUTED LEARNING AND WITH FEDERATED EDGE LEARNING WITHOUT CHANNEL STATE INFORMATION
4y 1m to grant Granted Jul 21, 2026
Patent 12682274
MODEL INTEGRATION APPARATUS, MODEL INTEGRATION METHOD, COMPUTER-READABLE STORAGE MEDIUM STORING A MODEL INTEGRATION PROGRAM, INFERENCE SYSTEM, INSPECTION SYSTEM, AND CONTROL SYSTEM
5y 0m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
61%
Grant Probability
94%
With Interview (+32.6%)
3y 6m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 320 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month