Prosecution Insights
Last updated: October 02, 2026
Application No. 18/558,458

LEARNING APPARATUS, PREDICTION APPARATUS, LEARNING METHOD, PREDICTION METHOD AND PROGRAM

Final Rejection §103
Filed
Nov 01, 2023
Priority
May 07, 2021 — nonprovisional of PCTJP2021017568
Examiner
LE, HUNG D
Art Unit
Tech Center
Assignee
Nippon Telegraph and Telephone Corporation
OA Round
2 (Final)
90%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
990 granted / 1099 resolved
+30.1% vs TC avg
Moderate +6% lift
Without
With
+6.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
20 currently pending
Career history
1120
Total Applications
across all art units

Statute-Specific Performance

§101
14.1%
-25.9% vs TC avg
§103
41.3%
+1.3% vs TC avg
§102
20.5%
-19.5% vs TC avg
§112
8.2%
-31.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1099 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 1. This Office Action is in response to the amendment filed on 07/21/2026. Claims 1, 4 and 6 have been amended. Claim 7 has been canceled. Claims 10-12 have been added. Claims 1-6 and 8-12 are pending. Response to Arguments 2. Applicant's arguments with respect to claims 1-6 and 8-12 have been considered but are moot in view of the new ground(s) of rejection. Examiner's Note 3. A latent vector (According to Google): "A latent vector is a set of numerical values (a vector) that represents compressed, abstract features of data in a lower-dimensional space. The term "latent" means "hidden" or "not directly observable," indicating that these vectors capture underlying, essential characteristics of the data rather than raw input features (like individual pixels)." A recurrent neural network (According to Google): “A recurrent neural network (RNN) is a type of artificial neural network designed to process sequential data by keeping a short-term memory of previous inputs. How It Works: (1) Loops and Memory: Unlike standard networks, an RNN has a feedback loop that passes the output of a hidden state from one step to the input of the next step. (2) : This design lets the network remember past information to understand the current input, much like remembering previous words in a sentence to grasp the meaning of the current word. Common Uses: (1) Language Translation: Translating a sequence of words from one language to another. (2) Speech Recognition: Turning spoken audio into written text. (3) Time-Series Forecasting: Predicting stock prices or weather based on past numbers. Limitations: (1) Vanishing Gradients: Basic RNNs struggle to remember long-term context over very long sequences because the signals get too weak during training. (2) Advanced Variants: Specialized models like LSTMs (Long Short-Term Memory) and GRUs (Gated Recurrent Units) were created to fix this memory loss problem.” Osborn et al, US 20230072423, [Osborn: Paragraphs 27 and 87 ("generating a statistical model for predicting musculoskeletal position information using signals recorded from sensors" and "a system for predicting body state information") [Osborn: Paragraphs 266-267 ("Signals sensed by wearable sensors placed at locations on a user's body may be provided as input to an inference model trained to generate spatial information for rigid segments of a multi-segment articulated rigid-body model of a human body. The spatial information may include, for example, position information of one or more segments, orientation information of one or more segments, joint angles between segments, and the like ")] [Osborn: Paragraph 272 ("to predict information about a position and/or a movement of a portion of a user's arm and/or the user's hand, which may be represented as a multi-segment articulated rigid-body system with joints connecting the multiple segments of the rigid-body system")] [Osborn: Paragraph 274 ("information of a segment relative to other segments in the model are predicted using a trained inference model")] [Osborn: Paragraphs 295 and 565("the neural network may include an output layer that is a softmax layer, such that outputs of the inference model add up to one and may be interpreted as probabilities. For instance, outputs of the softmax layer may be a set of values corresponding to a respective set of control signals, with each value indicating a probability that the user wants to perform a particular control action. As one non-limiting example, the outputs of the softmax layer may be a set of three probabilities (e.g., 0.92, 0.05, and 0.03) indicating the respective probabilities that a detected pattern of activity is one of three known patterns") [Osborn: Paragraph 1315 ("For example, a user can perform an index pinch and the system can properly classify the pinch and associate and plot a corresponding first latent vector that can be presented to the user. The user can instruct the system that it is going to perform the same gesture again. When the user performs the gesture again, they can do SO with a slight modification (e.g., different wrist angle or degree of rotation). Based on the processed EMG data for that second gesture, the system can associate and plot a corresponding second latent vector that can be presented to the user. The system can quantify the distance between the first and second latent vectors and use that calculation to improve its ability to detect that specific gesture classification")]. Hu et al, US 20220086401, [Hu: Paragraphs 8-9 ("extracting video features of a first portion of the video stream, the first portion of the video stream including a first plurality of video frames and corresponding to a first time period; forming, based on the video features, a first vector in a latent space; computing a first index of a first match of the first vector in the latent space, wherein a first similarity score is associated with the first index and the first match corresponds to a second vector in the latent space; determining if the first similarity score exceeds a similarity threshold; and when the similarity threshold is exceeded, transmitting first data related to the first index; wherein the second vector is pre-determined by: extracting text features of a natural language query, and forming, based on the text features, the second vector in the latent space. ")]. Shu et al, US 20180285348, [Shu: Abstract and paragraphs 6 and 15 ("converting each term in a Kth round of a query sentence into a first word vector, and calculating a positive latent vector and a negative latent vector of each term according to the first word vector, where K denotes a positive integer greater than or equal to 2. A content topic of the Kth round of the query sentence is obtained, and converted into a second word vector. An initial latent vector output for the Kth round of the query sentence is determined according to the second word vector, the positive latent vector of a last term in the Kth round of the query sentence, a latent vector of a last term in a (K- 1)th round of a reply sentence output for a (K- 1)th round of the query sentence, and an initial latent vector of the (K 1)th round of the reply sentence output for the (K- 1)th round of the query sentence ")] [Shu: Paragraph 33 ("A second latent vector output for the K.sup.th round of the query sentence is determined according to the initial latent vector output for the K.sup.th round of the query sentence and a word vector of a preset identification character, and the first reply term to be output for the K.sup.t round of the query sentence is determined according to the second latent vector; contribution of each term in the K.sup.t round of the query sentence to generation of the second reply term is calculated according to the second latent vector and the latent vector of each term in the K.sup.th round of the query sentence; a third latent vector is calculated according to the contribution of each term in the K.sup.t round of the query sentence to the generation of the second reply term, the second latent vector, and a word vector of the first reply term; and the second reply term for the K.sup.th round of the query sentence is generated according to the third latent vector, and the process is repeated to generate the reply sentence for the K.sup.th round of the query sentence ")] [Shu: Paragraph 39 ("a weight of each term in the K.sup.th round of the query sentence for the generation of the second reply term is calculated according to the second latent vector and the latent vector of each term in the K.sup.t round of the query sentence; a weighted sum of the latent vector of each term in the K.sup. round of the query sentence is calculated according to the weight of each term in the K.sup. round of the query sentence for the generation of the second reply term, and the weighted sum is used as the contribution of each term in the K.sup. th round of the query sentence to the generation of the second reply term", i.e., "weighted sum" is 'an intensity function') ')] [Shu: Paragraph 46 ("The intention layer 22 outputs the initial latent vector, inputs the initial latent vector and a word vector whose first character is an identification character " EOS " to the decoding layer 23, and updates the latent vector by using a neural network to obtain a second latent vector. The second latent vector generates a probability distribution of ten thousand terms by using a softmax regression algorithm. A term "wo" has a highest probability, and therefore a reply term "wo" is output. The second latent vector and a word vector of the reply term "wo" are used as an input, to generate a third latent vector. ", i.e., "softmax regress algorithm" metrics is also considered as 'an intensity function') [Shu: Paragraph 22 ("Softmax: Promotion of a logistic regression model in a multi-categorization problem" and "the initial latent vector h.sup.(in,k) in an interval of [0, 1], to improve a nonlinear representation capability of the mode ")] [Shu: Paragraph 75 ("A second latent vector output for the K.sup.th round of the query sentence is determined according to the initial latent vector output")]. Margolin, US 20210406218, [Margolin: Abstract and paragraph 44 ("The recommendation module 120 generates a likelihood scalar value indicating a likelihood of the query is answered by a candidate user in a set of users using a combination of the vector of first latent features (e.g., 218) and a vector of second latent features different from the vector of first latent features (e.g., 220). The recommendation module 120 also can extract a plurality of second features of the candidate user from a user profile associated with the candidate user to generate a set of second visible features for the candidate user. In some aspects, the user profile can be obtained from the user data repository 128. The recommendation module 120 can generate the vector of second latent features from the set of second visible features using the machine learning-trained classifier 216. In some aspects, the vector of second latent features includes latent representations of the set of second visible features. In some aspects, the vector of second latent features indicates a latent representation of the set of second visible features in a same feature space as that of the set of second visible features. In some aspects, the vector of second latent features includes a plurality of second latent feature fields, in which each of the plurality of second latent feature fields includes a different linear combination of the plurality of second features in the set of second visible features. ")] [Margolin: Paragraph 75 ("In generating the vector of. first latent features, the recommendation module 120 can generate a first embedding representation of the set of first visible features with a pre-trained language model and apply the first embedding representation to the machine learning-trained classifier to generate the vector of first latent features. In generating the vector of second latent features, the recommendation module 120 can generate a second embedding representation of the set of second visible features with the pre-trained language model and apply the second embedding representation to the machine learning-trained classifier to generate the vector of second latent features. ")]. Gregory et al, US 12,457,110, [Gregory: Column 3, lines 14-27 ("The machine learning model may be a temporal self-attention network. The operations may further include: generaling a first regional graph based at least on the first publicly verifiable exposure data; embedding the first regional graph as a first set of points in a first latent vector space; generating a second regional graph based at least on the second publicly available exposure data; embedding the second regional graph as a second set of points in a second latent vector space; and generating the temporal self-attention network based at least on the first latent vector space and the second latent vector space. The first regional graph may include multiple nodes representing respective locations in a geospatial region and multiple edges representing traffic flow between the respective locations. ")]. Oezer, US 20240104385, [Oezer: Paragaph 48 ("It shall also be understood that throughout this discussion that components may be described as separate functional units, which may comprise sub-units, but those skilled in the art will recognize that various components, or portions thereof, may be divided into separate components or may be integrated together, including integrated within a single system or component. ")] [Oezer: Paragraph 63 ("a neural model may be a neural network or may be a neural field. In one or more embodiments, a neural field is a fully connected neural network with neurons that have two distinct vector parameters, where a first parameter is used for input space partitioning and a second parameter is used for approximation. In one or more embodiments, the first vector (partitioning) may be learned via clustering and the second vector (weight vector for approximation) may be learned via a linear least squares method. In one or more embodiments, a reinforcement reward may be represented by one or more world state properties (e.g., property of an object). In one or more embodiments, for plan generation, the neural models may be inverted via piecewise linear programming and piecewise linear programming may apply linear programming to each partition of the neural model. It should be noted that, in one or more embodiments, a neural field partition may be explicitly represented by the first (partition) vector")] [Oezer: Paragraph 923 ("Forecasting was tested with simulated data that simulates any typical forecasting task. The simulated data had the following components: (1) a long-term trend; (2) periodicity; (3) dependence on current features; (4) effect of special events; and (5) noise. The long-term trend was generated with a logistic function (i.e., The trend may be thought to simulate a long-term growth. Periodicity was generated via three sine waves with different frequencies and phases. This simulated periodic low, medium, and high frequency fluctuations. The three sine waves were scaled and added to the trend. A weighted value of random features was added. Features may have a positive or negative impact on the function value. Events represent the occurrence of special situations or anomalies. Events may be Boolean features, which may be set to 1 if the event was occurring. Events may also be generated with a fixed event probability. If the event occurred, it may increase or decrease the function value at the time of the event. The input feature vector was extended with the recent time series data (i.e., prediction results for times t-1, t-2, t-3, ). Random noise was added to all input features. ")]. Chen et al, US 20210287116, [Chen: Paragraph 4 ("Machine learning defines models that can be used to predict occurrence of an event, for example, from sensor data or signal data, or recognize/classify an object, for example, in an image, in text, in a web page, in voice data, in sensor data, etc. Machine learning algorithms can be classified into three categories: unsupervised learning, supervised learning, and semi-supervised learning. ")] [Chen: Paragraph 88 ("subject to (V.sub.m.sup.r).sup.TV.sub.m.sup.r=I and Z.sub.i=l.sup.ns.sub.i.sup.T1=1 where (2.sub.m.sup.r(V.sub.m.sup.r).sup.T).sub.i represents an i.sup.th column of the estimated matrix Z, e ER.sup.cx1 is a first auxiliary vector, b ER.sup.Nx1 is a second auxiliary vector, ")]. Chen et al, US 10,929,762, [Chen: Column 1, lines 34-62 ("Machine learning defines models that can be used to predict occurrence of an event, for example, from sensor data or signal data, or recognize/classify an object, for example, in an image, in text, in a web page, in voice data, in sensor data, etc. ")] [Chen: Column 2, lines 15-67 ("(A) A next classified observation vector is selected from the plurality of classified observation vectors. (B) A distance value is computed between the selected next classified observation vector and each cluster center of the defined cluster centers. (C) When the target variable value of the selected next classified observation vector is not the unique class determined for a cluster center having a minimum computed distance value, a first distance value is selected as the minimum computed distance value, a second distance value is selected as the computed distance value to the cluster center having the unique class of the target variable value of the selected next classified observation vector, a ratio value is computed between the selected second distance value and the selected first distance value, and the target variable value of the selected next classified observation vector is changed to the unique class determined for the cluster center having the minimum computed distance value when the computed ratio value satisfies a predefined label correction threshold. (D) (A) through (C) are repeated until each observation vector of the plurality of classified observation vectors is selected in (A). (E) A classification matrix is defined using the plurality of observation vectors. (F) The target variable value is determined for each observation vector of the plurality of unclassified observation vectors based on the defined classification matrix. (G) The target variable value is output for each observation vector of the plurality of observation vectors, wherein the target variable value selected for each observation vector of the plurality of observation vectors is defined to represent the label for a respective observation vector", i.e., 'output a second latent vector based on each of the output first latent vectors ')] [Chen: Column 23, lines 58-67 through column 24, lines 1-7 ("Session manager device 800 may or may not include input classified data 124 and a portion of input unclassified data 126 divided into input unclassified data subset 814. For example, session manager device 800 may coordinate the distribution of input unclassified data 126 with or without storing a portion of input unclassified data", i.e., "a portion of input unclassified data" is considered as 'a support set extracted from a set of previous data' that is divided into "unclassified data subset" (dividing 'support set' into 'a plurality of sections ')] [Chen: Column 13, lines 9-33 ("In operation 228, a ratio value r.sub.v is computed between a second distance value and a first distance value, where the first distance value is the minimum computed distance value and the second distance value is the computed distance value to the cluster center having the unique class of the target variable value of the selected next classified observation vector. For example, the ratio value may be computed using where d.sub.1 is the first distance value, and d.sub.2 is the second distance value. r.sub. is the probability that the selected next classified observation vector is actually a member of the cluster associated with the minimum computed distance value instead of the cluster associated with the cluster center having the unique class of the target variable value of the selected next classified observation vector. This is the result because the smaller the first distance value is, the larger the probability that the selected next classified observation vector belongs to a specific cluster. " i.e., 'output a second latent vector based on each of the output first latent vectors')]. Bernhardsson, US 9,110,955, [Bernhardsson: Abstract (“A server computes n-dimensional latent vectors for each user and for each item. The server iteratively optimizes the user vectors and item vectors based on the data points. Each iteration includes a first phase in which the item vectors are held constant, and a second phase in which the user vectors are held constant. In the first phase, the server computes first phase parameters based on data points, the user vectors, and the item vectors, and updates the user vectors. In the second phase, the server similarly computes second phase parameters for the item vectors and updates the item vectors. The server receives a request from a user for an item recommendation, and selects an item vector based on proximity in n-dimensional space. The server then recommends the selected item to the user.”)] [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 11, lines 1-30 (“a modeling module 422, which utilizes the historical data in the access log 346 to predict which items 324 a user 112 will like.”, i.e., ‘a sequence of previous events’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 13, lines 64-67 through column 14, lines 1-27 (“estimating the latent user vectors 348 and latent item vectors 350. The goal is to find a set of latent vectors that best corresponds to the occurrence matrix 800. The process is iterative, and in each iteration the process uses the current user and item vector estimates to compute a new set of estimates.” And “Because P(u,i) is a probability distribution, .SIGMA..sub.u,iP(u,i)=1, so Z=.SIGMA..sub.u,i exp ({right arrow over (u)}{right arrow over (i)}), as illustrated by equation 904 in FIG. 9. In FIG. 9, the matrix 900 is a probability distribution over both users and items, and the user and item vectors 348 and 350 are computed so that the probability distribution corresponds to the occurrence matrix 800 as closely as possible”, i.e., ‘predicting an occurrence of an event’)] [Bernhardsson: Column 18, lines 12-34 and column 18, lines 47-58 (“The latent user vectors 348 and latent item vectors 350 are stored for later use, such as in the database 118. Later, the latent user and item vectors 348 and 350 are used to make item recommendations. A request for an item recommendation is received (1064) from a user 112. In some implementations, the user 112 corresponds to a latent user vector u.sub.0 348. In some cases, the user 112 does not correspond to stored latent user vector 348 (e.g., a new user). … user might like” and “full steam of log events”, i.e., ‘a sequence of previous events extracted from a set of previous data’ and ‘predicting an occurrence of an event’)]. Claim Rejections - 35 USC § 103 4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 6. Claims 1-6 and 8-12 are rejected under 35 U.S.C. 103 as being unpatentable over Bernhardsson (US 9,110,955), in view of Shu et al (US 20180285348). Claim 1: Bernhardsson suggests a learning device for predicting an occurrence of an event [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)], comprising: a memory; and a processor configured to: a dividing unit which divide a support set comprising a sequence of previous events extracted from a set of previous data for learning into a plurality of sections [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 11, lines 1-30 (“a modeling module 422, which utilizes the historical data in the access log 346 to predict which items 324 a user 112 will like.”, i.e., ‘a sequence of previous events’)] [Bernhardsson: Column 18, lines 12-34 and column 18, lines 47-58 (“The latent user vectors 348 and latent item vectors 350 are stored for later use, such as in the database 118. Later, the latent user and item vectors 348 and 350 are used to make item recommendations. A request for an item recommendation is received (1064) from a user 112. In some implementations, the user 112 corresponds to a latent user vector u.sub.0 348. In some cases, the user 112 does not correspond to stored latent user vector 348 (e.g., a new user). … user might like” and “full steam of log events”, i.e., ‘a sequence of previous events extracted from a set of previous data’ and ‘predicting an occurrence of an event’)]. Bernhardsson suggests outputting a first latent vector based on each of the plurality of divided sections and outputting a second latent vector based on each of the output first latent vectors [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)]. Bernhardsson suggests outputting an intensity function indicating a likelihood of the event occurring at a time during a prediction period based on the second latent vector and the time [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)]. Shu suggests a learning device for predicting an occurrence of an event [Shu: Paragraph 17 (“Recurrent neural network (RNN): A recurrent neural network may be used to model a time sequence behavior”). Paragraph 18 (“predict an important event with a quite long interval and delay in a time sequence”)]. Both references (Bernhardsson and Shu) taught features that were directed to analogous art and they were directed to the same field of endeavor, such as data processing. It would have been obvious to one of ordinary skill in the art at the time the invention was made, having the teachings of Bernhardsson and Shu before him/her, to modify the system of Bernhardsson with the teaching of Shu in order to implement machine learning models in processing or predicting data [Shu: Paragraph 17 (“Recurrent neural network (RNN): A recurrent neural network may be used to model a time sequence behavior”). Paragraph 18 (“predict an important event with a quite long interval and delay in a time sequence”)]. Claim 2: The combined teachings of Bernhardsson and Shu suggests updating any parameter of a first model for outputting the first latent vector, a second model for outputting the second latent vector, and a third model for outputting the intensity function based on the intensity function [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)]. Claim 3: The combined teachings of Bernhardsson and Shu suggests wherein the processor outputs the first latent vector based on each of the plurality of divided sections by parallel distributed processing [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 14, lines 1-10 (“Because the process involves a substantial amount of data, most of the operations are distributed across multiple computers operating in parallel (e.g, using cluster servers 124).”)]. Claim 4: Claim 4 is essentially the same as claim 1 and rejected under the same reasons as applied above. Claim 5: The combined teachings of Bernhardsson and Shu suggests predicting a situation of occurrences of events of an event in a prediction period using the intensity function [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)]. Claim 6: Claim 6 is essentially the same as claim 1 except that it sets forth the claimed invention as a method rather than a device and rejected under the same reasons as applied above. Claim 8: Claim 8 is essentially the same as claim 1 except that it sets forth the claimed invention as a program product rather than a device and rejected under the same reasons as applied above. Claim 9: Claim 9 is essentially the same as claim 4 except that it sets forth the claimed invention as a program product rather than a device and rejected under the same reasons as applied above. Claim 10: The combined teachings of Bernhardsson and Shu suggests wherein the processor is configured to divide the support set into the plurality of sections based on specified time intervals or by equalizing an expected value of a number of events included in each of the plurality of sections [Bernhardsson: Column 4, lines 19-28 (“the goal of a content provider is to increase the number of active users. In general, it is difficult or impossible to directly optimize the number of users. Instead, implementations typically focus on other measureable metrics, such as the number of "skip-forward" actions by users while listening to a stream of audio tracks; the length of time users interact with the provided content items;”)] [Bernhardsson: Column 5, lines 36-50 (“In other implementations, one or more items that the user likes are known (e.g., by explicit feedback, or interacting with the item multiple times)..”)]. Claim 11: The combined teachings of Bernhardsson and Shu suggests wherein additional information is added to the sequence of previous events, and the processor is configured to obtain a third latent vector based on the second latent vector and the additional information, and output the intensity function based on the third latent vector and the time [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)]. Claim 12: The combined teachings of Bernhardsson and Shu suggests wherein the processor is further configured to: extract a query set from the sequence of previous events, wherein the query set comprises events occurring within a time period after a time period of the support set; calculate a negative logarithmic likelihood from the intensity function and the query set; and update parameters of a first model for outputting the first latent vector, a second model for outputting the second latent vector, and a third model for outputting the intensity function, based on the negative logarithmic likelihood [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 9, lines 18-67 through column 10, lines 1-35 (“The database 118 also includes a list of users 336, who are typically registered users. This allows the content server to track the likes and dislikes of the users, and thus present users with content items 324 that better match a user's likes.”)]. Conclusion 7. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. 8. Any inquiry concerning this communication or earlier communications from the examiner should be directed to [Hung D. Le], whose telephone number is [571-270-1404]. The examiner can normally be communicated on [Monday to Friday: 9:00 A.M. to 5:00 P.M.]. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Apu Mofiz can be reached on [571-272-4080]. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, contact [800-786-9199 (IN USA OR CANADA) or 571-272-1000]. Hung Le 09/09/2026 /HUNG D LE/Primary Examiner, Art Unit 2161
Read full office action

Prosecution Timeline

Nov 01, 2023
Application Filed
Apr 28, 2026
Non-Final Rejection mailed — §103
Jul 21, 2026
Response Filed
Sep 14, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743447
READING COMPRESSED DATA DIRECTLY INTO AN IN-MEMORY STORE
1y 8m to grant Granted Sep 22, 2026
Patent 12724780
RECOVERY OF ROW DATA STORED VIA DIFFERENT STORAGE MECHANISMS
2y 1m to grant Granted Sep 01, 2026
Patent 12705279
DISPLAY APPARATUS, BACKGROUND MUSIC PROVIDING METHOD THEREOF AND BACKGROUND MUSIC PROVIDING SYSTEM
2y 6m to grant Granted Aug 11, 2026
Patent 12694009
TRACKING EVALUATION OF WORKLOAD STABILITY THROUGH PERFORMANCE INDEXING
2y 1m to grant Granted Jul 28, 2026
Patent 12682286
GENERATING OPPORTUNITY PROFILE INSIGHTS
3y 1m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
90%
Grant Probability
96%
With Interview (+6.4%)
2y 4m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1099 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month