Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the application and claims filed 02 May 2024. Claims 1-7 are pending and have been examined. Claims 1-7 are rejected.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 02 May 2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a first calculation unit” in claim 1, “second calculation unit” in claim 2, “update unit” in claim 3, “initialization unit” in claim 5.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-7 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract
idea without significantly more.
Step 1: Claim 1 is a machine type claim. Claim 6 is a process type claim. Claim 7 is a manufacture type claim. Therefore, claims 1-7 are directed to either a process, machine, manufacture or composition of matter.
As per claim 1,
2A Prong 1:
"Calculate a cumulative intensity function based on an output from the … and a product of a parameter and time" This is a mathematical calculation, in which the cumulative intensity function is calculated as the sum of a function output and the product of a parameter and time, per Formula (1) at paragraph [0066] of the specification.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
"An information processing apparatus comprising…", "a first calculation unit configured to… " (mere instructions to apply the exception using a generic computer component);
"a monotonic neural network", "the monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
"An information processing apparatus comprising…", "a first calculation unit configured to… " (mere instructions to apply the exception using a generic computer component);
"a monotonic neural network", "the monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
As per claim 2,
2A Prong 1:
"Calculate an intensity function related to a point process based on the calculated cumulative intensity function" This is a mathematical calculation, in which the intensity function of a point process, itself a probability model, is calculated by differentiating the cumulative intensity function (specification paragraphs [0002] and [0068]).
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
"The information processing apparatus", "a second calculation unit" (mere instructions to apply the exception using a generic computer component);
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
"The information processing apparatus", "a second calculation unit" (mere instructions to apply the exception using a generic computer component);
As per claim 3,
2A Prong 1:
"Update the parameter based on the calculated intensity function" This is a mathematical calculation, in which the parameter is updated by optimizing an evaluation function, such as a negative log likelihood, calculated from the intensity function, using error backpropagation (specification paragraphs [0071]-[0072]).
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
"The information processing apparatus", "an update unit" (mere instructions to apply the exception using a generic computer component);
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
"The information processing apparatus", "an update unit" (mere instructions to apply the exception using a generic computer component);
As per claim 4,
2A Prong 1:
"Update the parameter by using all events included in a sequence including a plurality of events discretely arranged on a continuous time or a number of the plurality of events included in the sequence as an input" This is a mathematical calculation, in which the parameter is calculated as the output of a function that takes the events of a sequence, or the count of those events, as its input (specification paragraph [0114]).
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
"The information processing apparatus" (mere instructions to apply the exception using a generic computer component);
"a neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The neural network here is a generic, unnamed neural network with no detail beyond taking a sequence or a count of events as an input and outputting a parameter. This could be any off the shelf neural network).
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
"The information processing apparatus" (mere instructions to apply the exception using a generic computer component);
"a neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The neural network here is a generic, unnamed neural network with no detail beyond taking a sequence or a count of events as an input and outputting a parameter. This could be any off the shelf neural network).
As per claim 5,
2A Prong 1:
"Initialize a plurality of weights applied to the … based on a distribution with a positive average" This is a mathematical calculation, in which initial weight values are generated as random numbers drawn according to a statistical distribution with a positive average, such as a normal or uniform distribution (specification paragraphs [0234]-[0237]).
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
"an initialization unit" (mere instructions to apply the exception using a generic computer component);
"the monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
"an initialization unit" (mere instructions to apply the exception using a generic computer component);
"the monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
As per claim 6,
2A Prong 1:
"Outputting a monotonically increasing function from a …" A monotonically increasing function is a mathematical function, and this limitation recites the mathematical operation of computing the output of that function.
"calculating a cumulative intensity function based on the output monotonically increasing function and a product of a parameter and time" This is a mathematical calculation, in which the cumulative intensity function is calculated as the sum of the function output and the product of a parameter and time, per Formula (1) at paragraphs [0066]-[0067] of the specification.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
"a monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
"a monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
As per claim 7,
2A Prong 1:
"Outputting a monotonically increasing function from a …" A monotonically increasing function is a mathematical function, and this limitation recites the mathematical operation of computing the output of that function.
"calculating a cumulative intensity function based on the output monotonically increasing function and a product of a parameter and time" This is a mathematical calculation, in which the cumulative intensity function is calculated as the sum of the function output and the product of a parameter and time, per Formula (1) at paragraphs [0066]-[0067] of the specification.
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
"A non-transitory storage medium storing a program for causing a computer to execute" (mere instructions to apply the exception using a generic computer component);
"a monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
"A non-transitory storage medium storing a program for causing a computer to execute" (mere instructions to apply the exception using a generic computer component);
"a monotonic neural network" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: The monotonic neural network here is generic, off the shelf machine learning. The claim recites no architecture, activation function, or other implementation detail, and the specification's background admits that monotonic neural networks were a known technique (paragraph [0003]-[0004], citing Non Patent Literature 1). This could be any off the shelf monotonic neural network).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Examiner’s Note: Some rejections will include an Examiner’s Note (labeled ‘EN’) to provide additional context or rationale explaining the basis for the rejection.
Claims 1-3, 6, and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Omi et al. ("Fully Neural Network based Model for General Temporal Point Processes,"), hereinafter "Omi" in view of Aigner et al. ("A competing risks interpretation of Hawkes processes,"), hereinafter "Aigner".
Claim 1
Omi teaches,
An information processing apparatus comprising: (Abstract, p. 1, "We herein propose a novel RNN based model in which the time course of the intensity function is represented in a general manner. In our approach, we first model the integral of the intensity function using a feedforward neural network and then obtain the intensity function as its derivative." Section 4, p. 6, "We performed these computations under a GPU environment provided by Google Colaboratory." - EN: This denotes Omi’s neural-network model executed on GPU-equipped computing hardware, which corresponds to the claimed information processing apparatus.")
a monotonic neural network; and (Section 2.4, p. 4, "In the present study, we model the cumulative hazard function using a feedforward neural network (a cumulative hazard function network; Fig. 1) for flexible modeling. The cumulative hazard function is a monotonically increasing function of τ and is positive-valued. The cumulative hazard function network is designed to reproduce these properties." Section 2.4, p. 4, "To summarize, the weights of the particular network connections are constrained to be positive (Fig. 1)." Section 2.4, p. 4, footnote 3, "In order to enforce the weights to be positive, if a weight is updated to be a negative value during training, we replace it with its absolute value." - EN: This denotes Omi’s cumulative hazard function network having an output that is monotonically increasing with respect to elapsed time, which corresponds to the claimed monotonic neural network.")
a first calculation unit configured to calculate a cumulative intensity function based on an output from the monotonic neural network (...) (Section 2.4, p. 4, eq. (12), "The cumulative hazard function Φ(τ|h_i) and the hazard function φ(τ|h_i) are now formulated as follows based on the output Z_i(τ) of the cumulative hazard function network: Φ(τ|h_i) = Z_i(τ)" Section 2.4, p. 3, eq. (9), "Rather than directly modeling the hazard function, we herein propose to model the cumulative hazard function Φ(τ|h_i), defined as follows: Φ(τ|h_i) = ∫_0^τ φ(s|h_i)ds." - EN: This denotes calculating Omi’s cumulative hazard function Φ(τ∣h i ), which integrates the point-process intensity and is obtained from the network output Z i (τ), which corresponds to calculating a cumulative intensity function based on an output from the monotonic neural network.)
Omi does not explicitly teach:
and a product of a parameter and time.
However, Aigner teaches:
and a product of a parameter and time. (Section 2.4, p. 5, "The Hawkes process compensator has the form Λ(t) = ηt + Σ_{i=1}^{N(t)} ∫{t_i}^t h(s - t_i)ds" Section 2.4, p. 5, "Finally, consider the task of simulating a sample (T_i){i=1}^n from a point process with stochastic intensity λ(t), or more generally with a compensator Λ(t) which when absolutely continuous is represented as Λ(t) = ∫_0^t λ(s)ds." Section 2.4, p. 5, "If we choose for h(t) a hazard function, however, and define the cumulative hazard H(t) = ∫0^t h(s)ds, then the compensator of N may be written as Λ(t) = ηt + Σ{i=1}^{N(t)} H(t - t_i)" - EN: This denotes Aigner’s cumulative intensity function Λ(t) including the term ηt, which corresponds to calculating the cumulative intensity function based on a product of the parameter η and time t.")
Before the effective filing date, it would have been obvious to modify Omi’s calculation of the cumulative intensity function to include the linear baseline term ηt taught by Aigner. The motivation for doing so would be to provide a linear component representing the cumulative background-event rate while allowing the neural-network output to represent the contribution associated with previous events. As Aigner explains in Section 1, p. 1, “The baseline intensity parameter η>0 represents the rate of background events, while the excitation kernel h(t) … models the influence of previous events on the conditional intensity.”
Claim 2
Omi further teaches:
The information processing apparatus according to claim 1, further comprising
a second calculation unit configured to calculate an intensity function related to a point process based on the calculated cumulative intensity function. (Section 2.4, p. 4, eq. (10), "The hazard function itself can be then obtained by differentiating the cumulative hazard function with respect to τ as follows: φ(τ|h_i) = ∂/∂τ Φ(τ|h_i)." Section 2.4, p. 5, "The differentiation term ∂Z_i(τ)/∂τ, which is the derivative of the network output with respect to the network input, is computed using automatic differentiation" Section 2.2, p. 3, eq. (5), "The conditional intensity function is then formulated as a function of the elapsed time from the most recent event and the hidden state of the RNN, given as follows: λ(t|H_t) = φ(t - t_i|h_i), where φ is a non-negative function referred to as a hazard function." - EN: the hazard function φ is the conditional intensity function of the point process per eq. (5), so obtaining φ(τ|h_i) by differentiating the calculated cumulative hazard function reads on calculating "an intensity function related to a point process based on the calculated cumulative intensity function." Under the broadest reasonable interpretation, the automatic differentiation process constitutes the "second calculation unit.")
Claim 3
Omi further teaches:
The information processing apparatus according to claim 2, further comprising an update unit configured to update (...) based on the calculated intensity function. (Section 2.2, p. 3, "The parameter values of the model are estimated by maximizing the log-likelihood function. For this purpose, the backpropagation through time (BPTT) is employed to obtain the gradient of the log-likelihood function." Section 2.4, p. 5, "The log-likelihood function in eq. (11) of our model is then given based on the network output via the eqs. (12) and (13). The gradient of the log-likelihood function with respect to the parameters is obtained using backpropagation" - EN: Omi's log-likelihood function, eq. (11), is composed of the hazard function (the intensity function) and the cumulative hazard function, so updating the parameter values by gradient-based maximization of the log-likelihood is an update based on the calculated intensity function. Under the broadest reasonable interpretation, the optimization process that performs this update constitutes the "update unit.")
Omi does not explicitly teach:
the parameter
However, Aigner teaches:
the parameter (Section 2.2, p. 4, "For data observed in the interval [0, T], say, the complete data log-likelihood of some parameter vector θ is then [Lewis and Mohler, 2011]" Section 2.2, p. 4, "The first two terms are the likelihood of a homogeneous Poisson process with intensity η_θ, applied to the background events." Section 2.2, p. 4, "The complete-data log-likelihood is essential to an EM-type estimation for Hawkes processes [Mohler et al., 2011]; optimising the expected complete-data log-likelihood makes up the M-step of such an algorithm." - EN: Aigner's baseline parameter η_θ is a component of the parameter vector θ, and the log-likelihood containing the terms z_0i log η_θ and −η_θT is optimized to estimate that vector. The parameter η of the term ηt relied upon at claim 1 is therefore a parameter of the model of the cumulative intensity function that is updated by likelihood optimization, which reads on "the parameter.")
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the gradient-based likelihood maximization of Omi with the baseline parameter η of the term ηt taught by Aigner, such that the update unit updates the parameter η together with the weights of the cumulative hazard function network during training. The motivation for doing so would be to fit the linear background-rate term to the observed event sequence rather than requiring its value to be fixed in advance. As Aigner elaborates regarding the benefit of this likelihood-based estimation methodology in Section 2.2, p. 4, "optimising the expected complete-data log-likelihood makes up the M-step of such an algorithm."
Claim 6
Omi teaches,
An information processing method comprising: (Abstract, p. 1, "In our approach, we first model the integral of the intensity function using a feedforward neural network and then obtain the intensity function as its derivative.")
outputting a monotonically increasing function from a monotonic neural network; and (Section 2.4, p. 4, "In this setting, the network output is monotonically increasing with respect to the elapsed time τ and takes only a positive value, which mimics the cumulative hazard function." Section 2.4, p. 4, "The positivity of the network output can be ensured using an output unit, in which activation function is positive-valued.")
The remaining limitation of claim 6 is substantially the same as the first calculation unit limitation of claim 1, therefore claim 6 is rejected under the same rationale as claim 1.
Claim 7
Omi teaches,
A non-transitory storage medium storing a program for causing a computer to execute: (Section 2.4, p. 3, footnote 1, "A source code is available online. https://github.com/omitakahiro/NeuralNetworkPointProcess" Section 2.4, p. 5, "Automatic differentiation is a method to calculate the derivative of an arbitrary function, and it can be easily carried out using neural network libraries such as TensorFlow and PyTorch." - EN: Omi's model is a program, distributed as source code and implemented with neural network libraries, that is stored and executed by a computer; under the broadest reasonable interpretation, the storage holding that program constitutes "a non-transitory storage medium storing a program for causing a computer to execute" the recited operations.)
The remaining limitations of claim 7 are substantially the same as method claim 6, therefore claim 7 is rejected under the same rationale as claim 6.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Omi in view of Aigner, and further in view of Zhang et al., (“Self-Attentive Hawkes Process”), hereinafter “Zhang”
Claim 4
Zhang teaches:
The information processing apparatus according to claim 1, further comprising a neural network configured to update the parameter by using all events included in a sequence including a plurality of events discretely arranged on a continuous time or a number of the plurality of events included in the sequence as an input. (Section 3.1, p. 2, “A temporal point process (TPP) is a stochastic process whose realization is a list of discrete events at time t∈R + .” Section 2.1, p. 2, “We indicate with S={(v i ,t i )} i=1 L an event sequence, where the tuple (v i ,t i ) is the i-th event of the sequence S, v i ∈U is the event type, and t i is the timestamp of the i-th event.” Section 4, p. 4, “Given a series of historical events until t i , to compute the intensity of the type-u at the timestamp t, we need to consider the influence of all types of events before it.” Section 4, p. 4, “This generates a hidden vector that summarizes the influence of all previous events.” Section 4, p. 4, “Since the intensity function of Hawkes processes is history-dependent, we compute three parameters of the intensity function based on the history hidden vector h u,i+1 ” via equations (12)–(14).) – EN: Zhang’s self-attention neural network receives the historical event sequence, including the event types and continuous-time timestamps, and generates a hidden vector h u,i+1 summarizing all previous events. Zhang then calculates the intensity-function parameters μ u,i+1 , η u,i+1 , and γ u,i+1 from that hidden vector. Because these parameters are recalculated as the historical event sequence changes, Zhang’s network updates an intensity-function parameter using all events included in the historical sequence as input. In the combination, Zhang’s sequence-conditioned parameter calculation is applied to the parameter of Aigner’s linear term ηt. Since the claim recites “using all events … or a number of … events …” satisfying one or the other suffices.
Before the effective filing date, it would have been obvious to combine Omi’s calculation of a cumulative intensity function based on the output of a monotonic neural network and Aigner’s linear term ηt with Zhang’s self-attention neural network that summarizes all previous events and calculates intensity-function parameters from the resulting history representation. The motivation for doing so would have been to allow the parameter of the linear term to adapt to the observed event sequence, thereby improving the expressive power of the linear component and the accuracy with which the cumulative intensity function represents sequence-specific, long-term event dynamics. As Zhang explains in Section 1, p. 1, “The vanilla Hawkes processes specify a fixed and static intensity function, which limits the capability of capturing complicated dynamics.”
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Omi in view of Aigner, and further in view of You et al. ("Deep Lattice Networks and Partial Monotonic Functions”), hereinafter "You".
Claim 5
You teaches:
The information processing apparatus according to claim 1, further comprising
an initialization unit configured to initialize a plurality of weights applied to the monotonic neural network based on a distribution with a positive average. (Section 5, p. 6, "For linear embedding layers, we initialize each component in the linear embedding matrix with IID Gaussian noise N(2, 1). The initial mean of 2 is to bias the initial parameters to be positive so that they are not clipped to zero by the first monotonicity projection." Section 2, p. 2, eq. (1), "To preserve monotonicity on the embedded vector W_t^m x_t^m, we impose the following linear inequality constraints: W_t^m[i, j] ≥ 0 for all (i, j)." - EN: You's Gaussian distribution N(2, 1) is a distribution with an average of 2, a positive average, and the initialized components are the weights of a deep network whose monotonicity is enforced by the non-negativity constraint of eq. (1). This reads on initializing "a plurality of weights applied to the monotonic neural network based on a distribution with a positive average." In the combination stated at claim 1, the initialization is applied to the weights of Omi's cumulative hazard function network, which are constrained to be positive.)
Before the effective filing date, it would have been obvious to initialize the positively constrained weights of Omi’s monotonic neural network using the positive-mean distribution taught by You. The motivation for doing so would be to bias the initialized weights toward positive values, thereby reducing the extent to which Omi’s positivity-enforcement operation would need to modify the weights. As You explains in Section 5, p. 6, “The initial mean of 2 is to bias the initial parameters to be positive so that they are not clipped to zero by the first monotonicity projection.”
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAYMUR RAHMAN ALI whose telephone number is (571)272-0007. The examiner can normally be reached Mon-Fri. 9:30-6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAYMUR RAHMAN ALI/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123