Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
1. This Office Action is in response to the amendment filed on 07/21/2026.
Claims 1, 4 and 6 have been amended.
Claim 7 has been canceled.
Claims 10-12 have been added.
Claims 1-6 and 8-12 are pending.
Response to Arguments
2. Applicant's arguments with respect to claims 1-6 and 8-12 have been considered but are moot in view of the new ground(s) of rejection.
Examiner's Note
3. A latent vector (According to Google): "A latent vector is a set of numerical values (a
vector) that represents compressed, abstract features of data in a lower-dimensional space. The
term "latent" means "hidden" or "not directly observable," indicating that these vectors capture
underlying, essential characteristics of the data rather than raw input features (like individual
pixels)."
A recurrent neural network (According to Google): “A recurrent neural network (RNN) is a type of artificial neural network designed to process sequential data by keeping a short-term memory of previous inputs. How It Works: (1) Loops and Memory: Unlike standard networks, an RNN has a feedback loop that passes the output of a hidden state from one step to the input of the next step. (2) : This design lets the network remember past information to understand the current input, much like remembering previous words in a sentence to grasp the meaning of the current word. Common Uses: (1) Language Translation: Translating a sequence of words from one language to another. (2) Speech Recognition: Turning spoken audio into written text. (3) Time-Series Forecasting: Predicting stock prices or weather based on past numbers. Limitations: (1) Vanishing Gradients: Basic RNNs struggle to remember long-term context over very long sequences because the signals get too weak during training. (2) Advanced Variants: Specialized models like LSTMs (Long Short-Term Memory) and GRUs (Gated Recurrent Units) were created to fix this memory loss problem.”
Osborn et al, US 20230072423, [Osborn: Paragraphs 27 and 87 ("generating a statistical model
for predicting musculoskeletal position information using signals recorded from sensors" and "a system for predicting body state information") [Osborn: Paragraphs 266-267 ("Signals sensed
by wearable sensors placed at locations on a user's body may be provided as input to an
inference model trained to generate spatial information for rigid segments of a multi-segment
articulated rigid-body model of a human body. The spatial information may include, for
example, position information of one or more segments, orientation information of one or more
segments, joint angles between segments, and the like ")] [Osborn: Paragraph 272 ("to predict
information about a position and/or a movement of a portion of a user's arm and/or the user's
hand, which may be represented as a multi-segment articulated rigid-body system with joints
connecting the multiple segments of the rigid-body system")] [Osborn: Paragraph 274
("information of a segment relative to other segments in the model are predicted using a trained
inference model")] [Osborn: Paragraphs 295 and 565("the neural network may include an
output layer that is a softmax layer, such that outputs of the inference model add up to one and
may be interpreted as probabilities. For instance, outputs of the softmax layer may be a set of
values corresponding to a respective set of control signals, with each value indicating a
probability that the user wants to perform a particular control action. As one non-limiting
example, the outputs of the softmax layer may be a set of three probabilities (e.g., 0.92, 0.05, and
0.03) indicating the respective probabilities that a detected pattern of activity is one of three
known patterns") [Osborn: Paragraph 1315 ("For example, a user can perform an index pinch
and the system can properly classify the pinch and associate and plot a corresponding first latent
vector that can be presented to the user. The user can instruct the system that it is going to
perform the same gesture again. When the user performs the gesture again, they can do SO with a slight modification (e.g., different wrist angle or degree of rotation). Based on the processed
EMG data for that second gesture, the system can associate and plot a corresponding second latent vector that can be presented to the user. The system can quantify the distance between the
first and second latent vectors and use that calculation to improve its ability to detect that
specific gesture classification")].
Hu et al, US 20220086401, [Hu: Paragraphs 8-9 ("extracting video features of a first portion of
the video stream, the first portion of the video stream including a first plurality of video frames
and corresponding to a first time period; forming, based on the video features, a first vector in a
latent space; computing a first index of a first match of the first vector in the latent space,
wherein a first similarity score is associated with the first index and the first match corresponds
to a second vector in the latent space; determining if the first similarity score exceeds a
similarity threshold; and when the similarity threshold is exceeded, transmitting first data
related to the first index; wherein the second vector is pre-determined by: extracting text features
of a natural language query, and forming, based on the text features, the second vector in the
latent space. ")].
Shu et al, US 20180285348, [Shu: Abstract and paragraphs 6 and 15 ("converting each term in
a Kth round of a query sentence into a first word vector, and calculating a positive latent vector
and a negative latent vector of each term according to the first word vector, where K denotes a
positive integer greater than or equal to 2. A content topic of the Kth round of the query sentence
is obtained, and converted into a second word vector. An initial latent vector output for the Kth
round of the query sentence is determined according to the second word vector, the positive
latent vector of a last term in the Kth round of the query sentence, a latent vector of a last term in
a (K- 1)th round of a reply sentence output for a (K- 1)th round of the query sentence, and an initial latent vector of the (K 1)th round of the reply sentence output for the (K- 1)th round of
the query sentence ")] [Shu: Paragraph 33 ("A second latent vector output for the K.sup.th
round of the query sentence is determined according to the initial latent vector output for the
K.sup.th round of the query sentence and a word vector of a preset identification character, and
the first reply term to be output for the K.sup.t round of the query sentence is determined
according to the second latent vector; contribution of each term in the K.sup.t round of the
query sentence to generation of the second reply term is calculated according to the second
latent vector and the latent vector of each term in the K.sup.th round of the query sentence; a
third latent vector is calculated according to the contribution of each term in the K.sup.t round
of the query sentence to the generation of the second reply term, the second latent vector, and a
word vector of the first reply term; and the second reply term for the K.sup.th round of the query
sentence is generated according to the third latent vector, and the process is repeated to
generate the reply sentence for the K.sup.th round of the query sentence ")] [Shu: Paragraph 39
("a weight of each term in the K.sup.th round of the query sentence for the generation of the
second reply term is calculated according to the second latent vector and the latent vector of
each term in the K.sup.t round of the query sentence; a weighted sum of the latent vector of
each term in the K.sup. round of the query sentence is calculated according to the weight of
each term in the K.sup. round of the query sentence for the generation of the second reply
term, and the weighted sum is used as the contribution of each term in the K.sup. th round of the
query sentence to the generation of the second reply term", i.e., "weighted sum" is 'an intensity
function') ')] [Shu: Paragraph 46 ("The intention layer 22 outputs the initial latent vector, inputs
the initial latent vector and a word vector whose first character is an identification character
" EOS " to the decoding layer 23, and updates the latent vector by using a neural network to obtain a second latent vector. The second latent vector generates a probability distribution of ten
thousand terms by using a softmax regression algorithm. A term "wo" has a highest probability,
and therefore a reply term "wo" is output. The second latent vector and a word vector of the
reply term "wo" are used as an input, to generate a third latent vector. ", i.e., "softmax regress
algorithm" metrics is also considered as 'an intensity function') [Shu: Paragraph 22
("Softmax: Promotion of a logistic regression model in a multi-categorization problem" and
"the initial latent vector h.sup.(in,k) in an interval of [0, 1], to improve a nonlinear
representation capability of the mode ")] [Shu: Paragraph 75 ("A second latent vector output for
the K.sup.th round of the query sentence is determined according to the initial latent vector
output")].
Margolin, US 20210406218, [Margolin: Abstract and paragraph 44 ("The recommendation
module 120 generates a likelihood scalar value indicating a likelihood of the query is answered
by a candidate user in a set of users using a combination of the vector of first latent features
(e.g., 218) and a vector of second latent features different from the vector of first latent features
(e.g., 220). The recommendation module 120 also can extract a plurality of second features of
the candidate user from a user profile associated with the candidate user to generate a set of
second visible features for the candidate user. In some aspects, the user profile can be obtained
from the user data repository 128. The recommendation module 120 can generate the vector of
second latent features from the set of second visible features using the machine learning-trained
classifier 216. In some aspects, the vector of second latent features includes latent
representations of the set of second visible features. In some aspects, the vector of second latent
features indicates a latent representation of the set of second visible features in a same feature space as that of the set of second visible features. In some aspects, the vector of second latent
features includes a plurality of second latent feature fields, in which each of the plurality of
second latent feature fields includes a different linear combination of the plurality of second
features in the set of second visible features. ")] [Margolin: Paragraph 75 ("In generating the
vector of. first latent features, the recommendation module 120 can generate a first embedding
representation of the set of first visible features with a pre-trained language model and apply the
first embedding representation to the machine learning-trained classifier to generate the vector
of first latent features. In generating the vector of second latent features, the recommendation
module 120 can generate a second embedding representation of the set of second visible features
with the pre-trained language model and apply the second embedding representation to the
machine learning-trained classifier to generate the vector of second latent features. ")].
Gregory et al, US 12,457,110, [Gregory: Column 3, lines 14-27 ("The machine learning model
may be a temporal self-attention network. The operations may further include: generaling a first
regional graph based at least on the first publicly verifiable exposure data; embedding the first
regional graph as a first set of points in a first latent vector space; generating a second regional
graph based at least on the second publicly available exposure data; embedding the second
regional graph as a second set of points in a second latent vector space; and generating the
temporal self-attention network based at least on the first latent vector space and the second
latent vector space. The first regional graph may include multiple nodes representing respective
locations in a geospatial region and multiple edges representing traffic flow between the
respective locations. ")].
Oezer, US 20240104385, [Oezer: Paragaph 48 ("It shall also be understood that throughout
this discussion that components may be described as separate functional units, which may
comprise sub-units, but those skilled in the art will recognize that various components, or
portions thereof, may be divided into separate components or may be integrated together,
including integrated within a single system or component. ")] [Oezer: Paragraph 63 ("a neural
model may be a neural network or may be a neural field. In one or more embodiments, a neural
field is a fully connected neural network with neurons that have two distinct vector parameters,
where a first parameter is used for input space partitioning and a second parameter is used for
approximation. In one or more embodiments, the first vector (partitioning) may be learned via
clustering and the second vector (weight vector for approximation) may be learned via a linear
least squares method. In one or more embodiments, a reinforcement reward may be represented
by one or more world state properties (e.g., property of an object). In one or more embodiments,
for plan generation, the neural models may be inverted via piecewise linear programming and
piecewise linear programming may apply linear programming to each partition of the neural
model. It should be noted that, in one or more embodiments, a neural field partition may be
explicitly represented by the first (partition) vector")] [Oezer: Paragraph 923 ("Forecasting
was tested with simulated data that simulates any typical forecasting task. The simulated data
had the following components: (1) a long-term trend; (2) periodicity; (3) dependence on current
features; (4) effect of special events; and (5) noise. The long-term trend was generated with a
logistic function (i.e.,
The trend may be thought to simulate a long-term growth. Periodicity was generated via three
sine waves with different frequencies and phases. This simulated periodic low, medium, and high frequency fluctuations. The three sine waves were scaled and added to the trend. A weighted
value of random features was added. Features may have a positive or negative impact on the
function value. Events represent the occurrence of special situations or anomalies. Events may
be Boolean features, which may be set to 1 if the event was occurring. Events may also be
generated with a fixed event probability. If the event occurred, it may increase or decrease the
function value at the time of the event. The input feature vector was extended with the recent time series data (i.e., prediction results for times t-1, t-2, t-3, ). Random noise was added to all
input features. ")].
Chen et al, US 20210287116, [Chen: Paragraph 4 ("Machine learning defines models that can
be used to predict occurrence of an event, for example, from sensor data or signal data, or
recognize/classify an object, for example, in an image, in text, in a web page, in voice data, in
sensor data, etc. Machine learning algorithms can be classified into three categories:
unsupervised learning, supervised learning, and semi-supervised learning. ")] [Chen: Paragraph 88 ("subject to (V.sub.m.sup.r).sup.TV.sub.m.sup.r=I and Z.sub.i=l.sup.ns.sub.i.sup.T1=1 where (2.sub.m.sup.r(V.sub.m.sup.r).sup.T).sub.i represents an i.sup.th column of the estimated matrix Z, e ER.sup.cx1 is a first auxiliary vector, b ER.sup.Nx1 is a second auxiliary vector, ")].
Chen et al, US 10,929,762, [Chen: Column 1, lines 34-62 ("Machine learning defines models
that can be used to predict occurrence of an event, for example, from sensor data or signal data,
or recognize/classify an object, for example, in an image, in text, in a web page, in voice data, in
sensor data, etc. ")] [Chen: Column 2, lines 15-67 ("(A) A next classified observation vector is
selected from the plurality of classified observation vectors. (B) A distance value is computed between the selected next classified observation vector and each cluster center of the defined
cluster centers. (C) When the target variable value of the selected next classified observation
vector is not the unique class determined for a cluster center having a minimum computed
distance value, a first distance value is selected as the minimum computed distance value, a
second distance value is selected as the computed distance value to the cluster center having the
unique class of the target variable value of the selected next classified observation vector, a ratio
value is computed between the selected second distance value and the selected first distance
value, and the target variable value of the selected next classified observation vector is changed
to the unique class determined for the cluster center having the minimum computed distance
value when the computed ratio value satisfies a predefined label correction threshold. (D) (A)
through (C) are repeated until each observation vector of the plurality of classified observation
vectors is selected in (A). (E) A classification matrix is defined using the plurality of observation
vectors. (F) The target variable value is determined for each observation vector of the plurality
of unclassified observation vectors based on the defined classification matrix. (G) The target
variable value is output for each observation vector of the plurality of observation vectors,
wherein the target variable value selected for each observation vector of the plurality of
observation vectors is defined to represent the label for a respective observation vector", i.e.,
'output a second latent vector based on each of the output first latent vectors ')] [Chen: Column
23, lines 58-67 through column 24, lines 1-7 ("Session manager device 800 may or may not
include input classified data 124 and a portion of input unclassified data 126 divided into input
unclassified data subset 814. For example, session manager device 800 may coordinate the
distribution of input unclassified data 126 with or without storing a portion of input unclassified
data", i.e., "a portion of input unclassified data" is considered as 'a support set extracted from a set of previous data' that is divided into "unclassified data subset" (dividing 'support set' into
'a plurality of sections ')] [Chen: Column 13, lines 9-33 ("In operation 228, a ratio value r.sub.v
is computed between a second distance value and a first distance value, where the first distance
value is the minimum computed distance value and the second distance value is the computed
distance value to the cluster center having the unique class of the target variable value of the
selected next classified observation vector. For example, the ratio value may be computed using
where d.sub.1 is the first distance value, and d.sub.2 is the second distance value. r.sub. is the
probability that the selected next classified observation vector is actually a member of the cluster
associated with the minimum computed distance value instead of the cluster associated with the
cluster center having the unique class of the target variable value of the selected next classified
observation vector. This is the result because the smaller the first distance value is, the larger the
probability that the selected next classified observation vector belongs to a specific cluster. "
i.e., 'output a second latent vector based on each of the output first latent vectors')].
Bernhardsson, US 9,110,955, [Bernhardsson: Abstract (“A server computes n-dimensional latent vectors for each user and for each item. The server iteratively optimizes the user vectors and item vectors based on the data points. Each iteration includes a first phase in which the item vectors are held constant, and a second phase in which the user vectors are held constant. In the first phase, the server computes first phase parameters based on data points, the user vectors, and the item vectors, and updates the user vectors. In the second phase, the server similarly computes second phase parameters for the item vectors and updates the item vectors. The server receives a request from a user for an item recommendation, and selects an item vector based on proximity in n-dimensional space. The server then recommends the selected item to the user.”)] [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 11, lines 1-30 (“a modeling module 422, which utilizes the historical data in the access log 346 to predict which items 324 a user 112 will like.”, i.e., ‘a sequence of previous events’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 13, lines 64-67 through column 14, lines 1-27 (“estimating the latent user vectors 348 and latent item vectors 350. The goal is to find a set of latent vectors that best corresponds to the occurrence matrix 800. The process is iterative, and in each iteration the process uses the current user and item vector estimates to compute a new set of estimates.” And “Because P(u,i) is a probability distribution, .SIGMA..sub.u,iP(u,i)=1, so Z=.SIGMA..sub.u,i exp ({right arrow over (u)}{right arrow over (i)}), as illustrated by equation 904 in FIG. 9. In FIG. 9, the matrix 900 is a probability distribution over both users and items, and the user and item vectors 348 and 350 are computed so that the probability distribution corresponds to the occurrence matrix 800 as closely as possible”, i.e., ‘predicting an occurrence of an event’)] [Bernhardsson: Column 18, lines 12-34 and column 18, lines 47-58 (“The latent user vectors 348 and latent item vectors 350 are stored for later use, such as in the database 118. Later, the latent user and item vectors 348 and 350 are used to make item recommendations. A request for an item recommendation is received (1064) from a user 112. In some implementations, the user 112 corresponds to a latent user vector u.sub.0 348. In some cases, the user 112 does not correspond to stored latent user vector 348 (e.g., a new user). … user might like” and “full steam of log events”, i.e., ‘a sequence of previous events extracted from a set of previous data’ and ‘predicting an occurrence of an event’)].
Claim Rejections - 35 USC § 103
4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. Claims 1-6 and 8-12 are rejected under 35 U.S.C. 103 as being unpatentable over Bernhardsson (US 9,110,955), in view of Shu et al (US 20180285348).
Claim 1:
Bernhardsson suggests a learning device for predicting an occurrence of an event [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)], comprising: a memory; and a processor configured to: a dividing unit which divide a support set comprising a sequence of previous events extracted from a set of previous data for learning into a plurality of sections [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 11, lines 1-30 (“a modeling module 422, which utilizes the historical data in the access log 346 to predict which items 324 a user 112 will like.”, i.e., ‘a sequence of previous events’)] [Bernhardsson: Column 18, lines 12-34 and column 18, lines 47-58 (“The latent user vectors 348 and latent item vectors 350 are stored for later use, such as in the database 118. Later, the latent user and item vectors 348 and 350 are used to make item recommendations. A request for an item recommendation is received (1064) from a user 112. In some implementations, the user 112 corresponds to a latent user vector u.sub.0 348. In some cases, the user 112 does not correspond to stored latent user vector 348 (e.g., a new user). … user might like” and “full steam of log events”, i.e., ‘a sequence of previous events extracted from a set of previous data’ and ‘predicting an occurrence of an event’)].
Bernhardsson suggests outputting a first latent vector based on each of the plurality of divided sections and outputting a second latent vector based on each of the output first latent vectors [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)].
Bernhardsson suggests outputting an intensity function indicating a likelihood of the event occurring at a time during a prediction period based on the second latent vector and the time [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)].
Shu suggests a learning device for predicting an occurrence of an event [Shu: Paragraph 17 (“Recurrent neural network (RNN): A recurrent neural network may be used to model a time sequence behavior”). Paragraph 18 (“predict an important event with a quite long interval and delay in a time sequence”)].
Both references (Bernhardsson and Shu) taught features that were directed to analogous art and they were directed to the same field of endeavor, such as data processing. It would have been obvious to one of ordinary skill in the art at the time the invention was made, having the teachings of Bernhardsson and Shu before him/her, to modify the system of Bernhardsson with the teaching of Shu in order to implement machine learning models in processing or predicting data [Shu: Paragraph 17 (“Recurrent neural network (RNN): A recurrent neural network may be used to model a time sequence behavior”). Paragraph 18 (“predict an important event with a quite long interval and delay in a time sequence”)].
Claim 2:
The combined teachings of Bernhardsson and Shu suggests updating any parameter of a first model for outputting the first latent vector, a second model for outputting the second latent vector, and a third model for outputting the intensity function based on the intensity function [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)].
Claim 3:
The combined teachings of Bernhardsson and Shu suggests wherein the processor outputs the first latent vector based on each of the plurality of divided sections by parallel distributed processing [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 14, lines 1-10 (“Because the process involves a substantial amount of data, most of the operations are distributed across multiple computers operating in parallel (e.g, using cluster servers 124).”)].
Claim 4:
Claim 4 is essentially the same as claim 1 and rejected under the same reasons as applied above.
Claim 5:
The combined teachings of Bernhardsson and Shu suggests predicting a situation of occurrences of events of an event in a prediction period using the intensity function [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)].
Claim 6:
Claim 6 is essentially the same as claim 1 except that it sets forth the claimed invention as a method rather than a device and rejected under the same reasons as applied above.
Claim 8:
Claim 8 is essentially the same as claim 1 except that it sets forth the claimed invention as a program product rather than a device and rejected under the same reasons as applied above.
Claim 9:
Claim 9 is essentially the same as claim 4 except that it sets forth the claimed invention as a program product rather than a device and rejected under the same reasons as applied above.
Claim 10:
The combined teachings of Bernhardsson and Shu suggests wherein the processor is configured to divide the support set into the plurality of sections based on specified time intervals or by equalizing an expected value of a number of events included in each of the
plurality of sections [Bernhardsson: Column 4, lines 19-28 (“the goal of a content provider is to increase the number of active users. In general, it is difficult or impossible to directly optimize the number of users. Instead, implementations typically focus on other measureable metrics, such as the number of "skip-forward" actions by users while listening to a stream of audio tracks; the length of time users interact with the provided content items;”)] [Bernhardsson: Column 5, lines 36-50 (“In other implementations, one or more items that the user likes are known (e.g., by explicit feedback, or interacting with the item multiple times)..”)].
Claim 11:
The combined teachings of Bernhardsson and Shu suggests wherein additional information is added to the sequence of previous events, and the processor is configured to obtain a third
latent vector based on the second latent vector and the additional information, and output the intensity function based on the third latent vector and the time [Bernhardsson: Column 2, lines 45-60 (“It would be useful to express each entry in this matrix as a function of a user vector u and an item vector i (these are latent vectors). Although this cannot be done exactly, user and item vectors can be chosen so that the vector products approximate the entries in the usage matrix, up to a multiplicative constant”, i.e., i.e., “vector products” = ‘output a second latent vector based on each of the output first latent vectors’)] [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 5, lines 36-50 (“The method selects another item that is close to one of the known desirable items by computing inner products of latent item vectors. The "close" item is then recommended to the user”, i.e., “close” = ‘intensity score’)] [Bernhardsson: Column 12, lines 61-67 through column 13, lines 1-19 (“In this example, the vectors are shown with only three components, but a typical implementation would construct latent vectors with 30 or 40 components. The first vector 702 has components 704, 706, and 708, the second vector 710 has components 712, 714, and 716. When the vectors are expressed algebraically, the inner product of two vectors is the sum of the products of the individual components, and is commonly referred to as the dot product. For example, the inner product 720 i.sub.4i.sub.5=(0.227*0.943)+(0.793*0.106)+(0.566*0.318), which is approximately 0.478. Similarly, the inner product 722 i.sub.5i.sub.6=(0.943*0.333)+(0.106*0.667)+(0.318*0.667), which is approximately 0.597. Finally, the inner product 724 i.sub.4i.sub.6=(0.227*0.333)+(0.793*0.667)+(0.566*0.667), which is approximately 0.982. In this example, the lengths of the vectors have been normalized to 1 (e.g., (0.227).sup.2+(0.793).sup.2+(0.566).sup.2.apprxeq.1). In this way, the inner products between pairs of vectors correspond to the angles. Therefore, it is apparent that vectors 702 and 718 are the closest pair of vectors. FIG. 7 also illustrates how easy it is computationally to compute the inner product of two vectors when their components are known. Even if the vectors have 40 or 50 components, an ordinary computer (e.g., content server 106 or cluster server 124) can compute an inner product almost instantaneously”, i.e., “product of two vectors” = ‘output a second latent vector based on each of the output first latent vectors’)].
Claim 12:
The combined teachings of Bernhardsson and Shu suggests wherein the processor is further
configured to: extract a query set from the sequence of previous events, wherein the query set comprises events occurring within a time period after a time period of the support set; calculate a negative logarithmic likelihood from the intensity function and the query set; and update parameters of a first model for outputting the first latent vector, a second model for outputting the second latent vector, and a third model for outputting the intensity function, based on the negative logarithmic likelihood [Bernhardsson: Column 4, lines 29-35 (“An issue with some collaborative filtering algorithms for implicit data is that they focus on predicting future user events … There are also fewer popular items to choose from, meaning that the user is more likely to choose one of those specific items”, i.e., ‘predicting an occurrence of an event’ using ‘ a sequence of previous events’)] [Bernhardsson: Column 9, lines 18-67 through column 10, lines 1-35 (“The database 118 also includes a list of users 336, who are typically registered users. This allows the content server to track the likes and dislikes of the users, and thus present users with content items 324 that better match a user's likes.”)].
Conclusion
7. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
8. Any inquiry concerning this communication or earlier communications from the examiner should be directed to [Hung D. Le], whose telephone number is [571-270-1404]. The examiner can normally be communicated on [Monday to Friday: 9:00 A.M. to 5:00 P.M.].
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Apu Mofiz can be reached on [571-272-4080]. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, contact [800-786-9199 (IN USA OR CANADA) or 571-272-1000].
Hung Le
09/09/2026
/HUNG D LE/Primary Examiner, Art Unit 2161