Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on May 7, 2024, January 27, 2025, and May 21, 2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
“an initialization unit configured to initialize a plurality of weights based on a distribution with a positive average” in claim 1 and claims 2-6 by dependency,
“a first calculation unit configured to calculate a cumulative intensity function based on an output from the monotonic neural network” in claim 1 and claims 2-6 by dependency,
“a second calculation unit configured to calculate an intensity function related to a point process based on the calculated cumulative intensity function” in claim 2,
“the initialization unit is configured to initialize each of the plurality of weights with a positive fixed value” in claim 3, and
“the initialization unit is configured to generate a random number based on the distribution, and initialize each of the plurality of weights with the random number” in claim 4 for the random-number generation and initialization limitation, respectively.
The modifiers “initialization”, “first calculation”, and “second calculation” describe the respective functions assigned to the claimed “units”, but do not impart sufficiently definite structure to the generic placeholder ‘unit’ for performing the recited functions. The corresponding structures and algorithms are identified below.
Regarding Initialization Unit
For claims 1-6, the claimed function is to “initialize a plurality of weights based on a distribution with a positive average”. Claim 3 further requires initializing each weight with a positive fixed value. Claim 4 further requires generating a random number based on the distribution and initializing each weight with the generated random number.
The corresponding structure is control circuit 10 programmed to perform the Rule Y initialization algorithm through initialization unit 42, and equivalents thereof. The algorithm:
Identifies the weights among parameters p2 to be applied to monotonic neural network 44-1;
Initialize the weights according to Rule Y by:
assigning positive fixed values;
generating values according to a normal distribution having a positive mean and applying them to the weights; or
generating values according to a uniform distribution having a nonnegative minimum and positive maximum and applying them to the weights; and
Provides the initialized parameters p2 for application to monotonic neural network 44-1. (See ¶¶ [0234]-[0238], [0262], [0280], and Figure 20)
Paragraphs [0234]-[0238] expressly disclose and link initialization unit 42 to this algorithm. Paragraph [0236] discloses positive fixed values, including 0.01 and
2.0
∙
10
-
3
. Paragraph [0237] discloses a normal distribution having mean/average
α
1
and standard deviation
α
2
n
, where
α
1
and
α
2
are positive. Paragraph [0238] discloses a uniform distribution having minimum
α
3
, where
α
3
≥
0, and a maximum
α
4
, where
α
4
> 0. Paragraph [0262] and Figure 20, step S71, provide corresponding operational disclosure. See also ¶[0280].
Accordingly, the corresponding structure is the disclosed programmed computer performing the Rule Y initialization algorithm, rather than any processor merely capable of initializing weights.
Regarding First Calculation Unit
For claims 1-6, the claimed function is to “calculate a cumulative intensity function based on an output from the monotonic neural network”.
The corresponding structure is control circuit 10 programmed to perform either of the following disclosed cumulative-intensity algorithms, and equivalents thereof.
In the first embodiment, cumulative intensity function calculation unit 24-2:
receives or accesses
f
(
z
,
t
)
and
f
(
z
,
0
)
from monotonic neural network 24-1 and accesses parameter
β
;
calculates:
Δ
t
=
∫
0
t
λ
u
d
u
=
f
z
,
t
+
β
t
-
f
(
z
,
0
)
; and
provides
λ
(
t
)
to automatic differentiation unit 24-3. (See ¶¶ [0065]-[0068] and Eq. (1)).
In the second embodiment, cumulative intensity function calculation unit 44-2:
receives or accesses
f
(
z
,
t
)
and
f
(
z
,
0
)
from monotonic neural network 44-1;
calculates:
Δ
t
=
∫
0
t
λ
u
d
u
=
f
z
,
t
-
f
(
z
,
0
)
; and
provides
λ
(
t
)
to automatic differentiation unit 44-3. (See ¶¶ [0265]-[0267], Eq. (3), and Figure 20, steps S74-S75).
Paragraph [0366] expressly discloses combining the first-embodiment cumulative-intensity calculation with the second-embodiment Rule Y initialization. Both Formula (1) and Formula (3) are therefore clearly linked corresponding algorithms for the claimed first calculation function.
The operation that generates
f
(
z
,
t
)
and
f
(
z
,
0
)
is performed by the separately claimed monotonic neural network. The first calculation unit receives those outputs and performs the disclosed cumulative-intensity calculation.
Regarding Second Calculation Unit
For claim 2, the claimed function is to “calculate an intensity function related to a point process based on the calculated cumulative intensity function”.
The corresponding structure is control circuit 10 programmed to perform the following algorithm, and equivalents thereof:
receive or access the calculated cumulative intensity function
λ
(
t
)
;
automatically differentiate
λ
(
t
)
with respect to time (t); and
output the resulting intensity function
λ
(
t
)
.
The algorithm is performed by automatic differentiation units 24-3 and 44-3. (See ¶¶ [0068], [0244]). Paragraph [0267] and Figure 20, step S76, provide corresponding operational disclosure. Paragraphs [0042] and [0063] establish that the resulting intensity function is related to the modeled point process.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 5 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 5 recites that “the second value is a positive square root of a value obtained by dividing a positive real number by a number of nodes of the monotonic neural network”. It is unclear which node population is represented by “a number of nodes”. For a multilayer monotonic neural network, the phrase may refer to the total number of network nodes, the number of nodes in the layer whose weights are initialized, the number of nodes in a preceding layer, or another layer-specific node count. These interpretations produce different second values.
Paragraph [0237] states that
n
is “the number of nodes of the layer”, without identifying the layer, while paragraph [0060] describes initialization using the number of nodes in a previous layer. Accordingly, the specification does not resolve the ambiguity. Applicant may overcome the rejection by expressly identifying the layer or node population used to calculate the second value.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 1 is provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 5 of copending Application No. 18/706,833 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other for the reasons stated below.
For purposes of this rejection, the “initialization unit” and “first calculation unit” limitations of present claim 1 are interpreted under 35 U.S.C. § 112(f), consistent with the claim interpretation set forth above. The corresponding structure includes control circuit 10 programmed to perform the disclosed initialization and cumulative-intensity-function calculation algorithms, and equivalents thereof.
Reference claim 5 depends from reference claim 1 and therefore incorporates all the limitations of reference claim 1. Collectively, reference claim 1 and 5 recite:
an information processing apparatus;
a monotonic neural network;
a first calculation unit configured to calculate a cumulative function based on an output from the monotonic neural network and a product of a parameter and time; and
an initialization unit configured to initialize a plurality of weights applied to the monotonic neural network based on a distribution with a positive average.
Regarding the limitation of present claim 1 reciting:
“an initialization unit configured to initialize a plurality of weights based on a distribution with a positive average;”
Reference claim 5 expressly recites:
“an initialization unit configured to initialize a plurality of weights applied to the monotonic neural network based on a distribution with a positive average.”
The initialization function recited by reference claim 5 is therefore the same as the initialization function recited by present claim 1 and additionally specifics that the initialized weights are applied to the monotonic neural network, as separately required by present claim 1.
Under 35 U.S.C. § 112(f), the initialization unit of reference claim 5 is construed in view of the corresponding structure and algorithm disclosed in the reference application. The reference application discloses a programmed control circuit implementing Rule Y, which initializes weights of the monotonic neural network based on a distribution having a positive average. This is the same initialization algorithm identified as corresponding structure for the initialization unit of present claim 1.
Regarding the limitation of present claim 1 reciting:
“a monotonic neural network to which the plurality of weights is applied”
Reference claim 1, from which reference claim 5 depends, recites:
“a first calculation unit configured to calculate a cumulative intensity function based on an output from the monotonic neural network and a product of a parameter and time.”
Under 35 U.S.C. § 112(f), the corresponding structure for the first calculation unit of present claim 1 includes control circuit 10 programmed to calculate the cumulative intensity function according to:
Λ
t
=
f
z
,
t
+
β
t
-
f
z
,
0
and:
Λ
t
=
f
z
,
t
-
f
(
z
,
0
)
and equivalents thereof.
The first-calculation function of reference claim 5 corresponds to the first disclosed alternative, in which the cumulative intensity function is calculated according to:
Λ
t
=
f
z
,
t
+
β
t
-
f
z
,
0
The “product of a parameter and time” recited by reference claim 1 corresponds to the term
β
t
. Reference claim 5 therefore claims the Formula (1) species expressly encompassed by the corresponding structure of present claim 1.
The additional requirement in reference claim 5 that the cumulative intensity function also be based on a product of a parameter and time narrows reference claim 5 relative to present claim 1. It does not patentably distinguish the broader subject matter of present claim 1 from the narrower subject matter of reference claim 5.
Accordingly, reference claim 5 includes every limitation of present claim 1 and further limits the first calculation unit to the Formula (1) species encompassed by present claim 1. The entire scope of reference claim 5 therefore fails within the scope of present claim 1. For purposes of Nonstatutory double patenting, reference claim 5 anticipates the genus of present claim 1. Issuance of present claim 1 would improperly extend the right to exclude that may be granted by issuance of reference claim 5.
Claim 2 is provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 5 of copending Application No. 18/706,833 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other.
Claim 2 incorporates all limitations of claim 1 and further recites:
“a second calculation unit configured to calculate an intensity function related to a point process based on the calculated cumulative intensity function.”
Reference claim 5, also depending from reference claim 1, supplies the positive-average initialization limitation discussed above.
Thus, reference claims 2 and 5 respectively claim, on the same reference claim 1 apparatus, the second-calculation limitation and the positive-average initialization limitation required by present claim 2.
It would have been an obvious variation of the subject matter of reference claims 2 and 5 to employ both expressly claimed features in the same apparatus. The initialization unit operates upstream to initialize the weights of the monotonic neural network, whereas the second calculation unit operates downstream to calculate the intensity function from the resulting cumulative intensity function. The functions are complementary and do not interfere with one another, and each would perform its expressly claimed function in the combined apparatus.
Accordingly, present claim 2 is not patentably distinct from reference claims 2 and 5.
Claim 3 is provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 5 of copending Application No. 18/706,833 in view of Keller et al. (US20220129755A1). Although the claims at issue are not identical, they are not patentably distinct from each other.
Reference claim 5 recites, in pertinent part, an apparatus including:
“an initialization unit configured to initialize a plurality of weights applied to the monotonic neural network based on a distribution with a positive average.”
Thus, reference claim 5 expressly claims initialization of the weights of the monotonic neural network based on a distribution having a positive average.
Present claim 3 depends from claim 1 and further recites:
“the initialization unit is configured to initialize each of the plurality of weights with a positive fixed value.”
Reference claim 5 does not expressly require that each of the plurality of weights be initialized with a positive fixed value.
Keller, however, teaches this limitation.
Keller teaches initializing neural-network weights prior to training using a positive fixed value. In particular, Keller teaches initializing weights using a small positive constant and provides an exemplary initialization value of 0.01:
“const float InitialWeight = 0.01f;” (Keller, pg. 18, Table 2)
Thus, Keller teaches initializing neural-network weights with a positive fixed value, as required by present claim 3.
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the weight initialization of the apparatus defined by reference claim 5 by using Keller’s known positive-fixed-value initialization technique.
Reference claim 5 already requires initialization of the weights applied to the monotonic neural network. Keller teaches a known manner of initializing neural-network weights using a predetermined positive constant. One of ordinary skill in the art would have recognized Keller’s positive-fixed-value initialization as a predictable alternative for providing initial positive values to the neural-network weights of the reference-claim-5 apparatus. Such a modification would merely substitute a known neural-network weight-initialization technique for the initialization technique of the claimed apparatus and would have predictably resulted in each of the plurality of weights being initialized with a positive fixed value.
One of ordinary skill in the art would have had a reasonable expectation of success because Keller’s technique is directed to the same general operation of initializing neural-network weights before training and would require only the use of predetermined positive initial values for the weights already required to be initialized by reference claim 5.
Accordingly, the subject matter of present claim 3 would have been an obvious variation of the invention defined by reference claim 5 in view of Keller and is therefore not patentably distinct from reference claim 5.
This is a provisional nonstatutory double patenting rejection.
Claims 4-6 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 5 of copending Application No. 18/706,833 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other.
Reference claim 5 recites the initialization unit configured to initialize the plurality of weights applied to the monotonic neural network based on a distribution with a positive average. Under the §112(f) construction of that limitation, the corresponding Rule Y algorithm identifies implementations for initializing weights based on a distribution having a positive average, including the random-number, normal-distribution, and uniform-distribution implementations recited in present claims 4-6.
Regarding claim 4
Claim 4 further recites that the initialization unit is configured to:
“generate a random number based on the distribution, and initialize each of the plurality of weights with the random number.”
Rule Y expressly includes applying to a weight “a random number generated according to a distribution with a positive average”. Thus, the random-number implementation of claim 4 is expressly within the Rule Y algorithm corresponding to the initialization unit of reference claim 5. Claim 4 is therefore not patentably distinct from reference claim 5.
Regarding claim 5
Claim 5 further requires a normal distribution having a positive first value as its average and a standard deviation equal to a positive square root of a positive real number divided by a number of nodes of the monotonic neural network.
Rule Y expressly identifies a normal-distribution initialization having an average
α
1
and standard deviation:
α
2
n
where
α
1
and
α
2
may each be positive and
n
is the number of nodes of the layer.
Thus, the initialization recited by claim 5 is an expressly identified Rule Y species within the scope of reference claim 5. One of ordinary skill would have at once envisaged and found it obvious to select that expressly identified normal-distribution implementation. Claim 5 is therefore not patentably distinct from reference claim 5.
Regarding claim 6
Claim 6 further requires a uniform distribution having a minimum value that is zero or greater and a maximum value that is a positive real number.
Rule Y expressly identifies initialization according to a uniform distribution having a minimum value
α
3
and a maximum value
α
4
, wherein any real number of zero or more may be used for
α
3
and any positive real number may be used for
α
4
.
Thus, claim 6 recites another expressly identified Rule Y species within the scope of reference claim 5. One of ordinary skill would have envisaged and found it obvious to select that expressly identified uniform-distribution implementation. Claim 6 is therefore not patentably distinct from reference claim 5.
Claim 7 is provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 5 and 6 of copending Application No. 18/706,833 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other.
Present claim 7 recites:
“initializing a plurality of weights based on of a distribution with a positive average;”
Reference claim 5 expressly claims an initialization unit configured to initialize a plurality of weights applied to the monotonic neural network based on a distribution with a positive average.
Present claim 7 further recites:
“calculating an output according to a monotonically increasing function from a monotonic neural network to which the plurality of weights is applied;”
Reference claim 6 expressly recites:
“outputting a monotonically increasing function from a monotonic neural network.”
Present claim 7 further recites:
“calculating a cumulative intensity function based on the calculated output.”
Reference claim 6 expressly recites calculating a cumulative intensity function based on the output monotonically increasing function and additionally requires that the cumulative intensity function also be based on a product of a parameter and time.
Thus, reference claim 6 claims the same method operations while further narrowing the cumulative-intensity calculation by requiring the
β
t
term, and reference claim 5 claims initialization of the weights applied to that monotonic neural network using a positive-average distribution.
It would have been an obvious variation to perform, before the method operations of reference claim 6, the weight-initialization operation expressly claimed in reference claim 5. The initialization merely supplies initial weights to the same monotonic neural network before the network produces its monotonically increasing output and does not alter the subsequent method operations. The resulting method performs each limitation of present claim 7, while retaining the additional Formula (1) limitation of reference claim 6.
Accordingly, claim 7 is not patentably distinct from reference claims 5 and 6.
Claim 8 is provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 5 and 7 of copending Application No. 18/706,833 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other.
Present claim 8 recites a storage medium storing a program for causing a computer to execute the same initialization, monotonic-network-output, and cumulative-intensity-calculation operations recited by claim 7.
Reference claim 7 expressly claims:
“A non-transitory storage medium storing a program for causing a computer to execute.”
and further recites outputting a monotonically increasing function from a monotonic neural network and calculating a cumulative intensity function based on that output and a product of a parameter and time.
Reference claim 5 expressly claims the positive-average initialization of the weights applied to the monotonic neural network.
It would have been an obvious variation to include in the program of reference claim 7 insgtructions for performing the positive-average weight initialization expressly claimed by reference claim 5 before executing the monotonic neural network and cumulative-intensity operations. Such instructions merely cause the computer to perform the initialization required for the same monotonic neural network and would predictably perform their claimed function without altering the subsequent stored-program operations.
Further, reference claim 7’s requirement of a non-transitory storage medium is narrower than the unrestricted “storage medium” recited by present claim 8 and therefore does not patentably distinguish present claim 8.
Accordingly, claim 8 is not patentably distinct from reference claims 5 and 7.
These are provisional nonstatutory double patenting rejections because the patentably indistinct claims have not in fact been patented.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-8 are rejected under 35 U.S.C. 101 because the claims recite an abstract idea without significantly more.
Regarding claim 1
Claim 1 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is to a machine.
Claim 1 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“an initialization unit configured to initialize a plurality of weights based on a distribution with a positive average;” – this limitation recites a mathematical concept under MPEP § 2106.04(a)(2)(I). A distribution and its average or mean describe mathematical/statistical relationships among numerical values, and initializing numerical model weights based upon those relationship constitute a mathematical calculation.
“and a first calculation unit configured to calculate a cumulative intensity function based on an output from the monotonic neural network.” – this limitation recites a mathematical concept under MPEP § 2106.04(a)(2)(I) because it requires calculation of a mathematical function form a numerical function output and therefore recites both a mathematical relationship and a mathematical calculation.
Claim 1 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements:
“an initialization unit” and “a first calculation unit” – these units are implemented by the programmed control circuit executing the mathematical initialization and cumulative-intensity algorithms identified above. The units therefore amount to instructions to implement the mathematical concepts using computer components and do not impose a meaningful limitation beyond using the computer as a tool to perform the exception. See MPEP § 2106.05(f).
“a monotonic neural network to which the plurality of weights is applied” – this limitation generally links the mathematical initialization and cumulative-intensity calculations to a neural-network technological environment. The claim does not require a particular architecture, activation function, training sequence, hardware implementation, or other specific technological configuration that is integral to the performance of the mathematical concepts. Thus, the limitation amounts to generally linking the judicial exception to a particular technological environment. See MPEP § 2106.05(h).
Claim 1 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are:
“an initialization unit” and “a first calculation unit” – the specification describes control circuit 10 as including a CPU, RAM, ROM, and the like, and explains that the CPU loads a program stored in the ROM or memory into the RAM and executes the program to cause the event prediction device to function as the disclosed processing units. Accordingly, the recited initialization unit and first calculation unit employ generic programmed-computer components performing their ordinary functions of receiving data, executing stored instructions, performing calculations, and producing calculated results. Such generic computer implementation is well-understood, routine, and conventional (WURC) activity and does not provide an inventive concept. See MPEP § 2106.05(d). Further, requiring a generic processor to execute the mathematical initialization and cumulative-intensity calculations amounts to mere instructions to apply the judicial exception using a computer as a tool. See MPEP § 2106.05(f).
“a monotonic neural network to which the plurality of weights is applied” – the specification acknowledges that neural networks were known for modeling points processes and further acknowledges previously proposed monotonic neural networks. Thus, use of a monotonic neural network, at the level of generality recited in claim 1, constitutes use of known neural-network technology and its well-understood, routine, and conventional (WURC) in the pertinent field. See MPEP § 2106.05(d). Additionally, claim 1 does not recite a particular neural-network architecture, number of arrangement of layers, particular activation function, particular hardware implementation, particular training process, or other specific technological configuration that itself provides an inventive concept. Rather, the monotonic neural network is the environment in which the mathematically initialized weights are applied and form which the output used in the mathematical cumulative-intensity calculation is obtained. Accordingly, the monotonic neural network, as broadly recited, generally links the use of the judicial exception to a particular technological environment, which does not amount to significantly more than the exception. See MPEP § 2106.05(h).
Regarding claim 2
Claim 2 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is to a machine.
Claim 2 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“further comprising a second calculation unit configured to calculate an intensity function related to a point process based on the calculated cumulative intensity function.” – this limitation recites a mathematical concept under MPEP § 2106.04(a)(2)(I). It requires mathematically deriving one function from another calculated function and therefore constitutes a mathematical relationship and mathematical calculation.
Claim 2 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements:
“a second calculation unit.” – under the 112(f) interpretation set forth above, the second calculation unit is programmed computer structure performing automatic differentiation of the cumulative intensity function. It therefore uses the computer as a tool to perform another mathematical operation and amounts to implementation of the exception rather than a meaningful practical application. See MPEP § 2106.05(f).
Claim 2 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are:
“a second calculation unit.” – the second calculation unit is implemented using the same programmed control circuit as the other processing units. Its automatic-differentiation function performs the mathematical derivation recited in the exception. The additional computer implementation therefore does not provide an inventive concept separate from the exception. See MPEP §§ 2106.05(d) and 2106.05(f).
Regarding claim 3
Claim 3 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is to a machine.
Claim 3 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“wherein the initialization unit is configured to initialize each of the plurality of weights with a positive fixed value.” – this limitation further limits the mathematical initialization calculation by assigning a specified class of numerical constants to the weights. It therefore recites a mathematical calculation under MPEP § 2106.04(a)(2)(I).
Claim 3 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no new additional elements that integrate the judicial exception into a practical application. The recited “initialization unit” is the same structural additional element identified in claim 1.
Claim 3 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no new additional elements that amount to significantly more than the judicial exception. The recited “initialization unit” is the same structural additional element identified in claim 1.
Regarding claim 4
Claim 4 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is to a machine.
Claim 4 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“wherein the initialization unit is configured to generate a random number based on the distribution, and initialize each of the plurality of weights with the random number.” – the operations of generating a numerical value according to a mathematical/statical distribution and assigning the generated numerical value to a model weight recite mathematical relationships and mathematical calculations under MPEP § 2106.04(a)(2)(I).
Claim 4 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no new additional elements that integrate the judicial exception into a practical application. The recited “initialization unit” is the same structural additional element identified in claim 1.
Claim 4 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no new additional elements that amount to significantly more than the judicial exception. The recited “initialization unit” is the same structural additional element identified in claim 1.
Regarding claim 5
Claim 5 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is to a machine.
Claim 5 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“wherein the distribution is a normal distribution in which an average is a first value and a standard deviation is a second value,” – this limitation expressly recites a statistical/mathematical distribution and mathematical parameters defining that distribution. It therefore recites a mathematical relationship under MPEP § 2106.04(a)(2)(I).
“the first value is a positive real number,” – this further mathematically constrains the numerical value of the distribution mean and therefore forms part of the mathematical relationship under MPEP § 2106.04(a)(2)(I).
“and the second value is a positive square root of a value obtained by dividing a positive real number by a number of nodes of the monotonic neural network.” – this expressly recites a mathematical formula involving division and a square-root calculation. It therefore recites a mathematical formula/equation and mathematical calculation under MPEP § 2106.04(a)(2)(I).
Claim 5 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no additional elements that integrate the judicial exception into a practical application.
Claim 5 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no additional elements that amount to significantly more than the judicial exception.
Regarding claim 6
Claim 6 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is to a machine.
Claim 6 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“wherein the distribution is a uniform distribution in which a minimum value is a third value and a maximum value is a fourth value,” – this recites a mathematical/statistical distribution defined by numerical boundary relationships and therefore constitutes a mathematical relationship under MPEP § 2106.04(a)(2)(I).
“the third value is a real number of 0 or more,” and “the fourth value is a positive real number.” – these limitations mathematically constrain the range-defining values of the uniform distribution and therefore further define the mathematical relationship constituting the exception under MPEP § 2106.04(a)(2)(I).
Claim 6 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no additional elements that integrate the judicial exception into a practical application.
Claim 6 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no additional elements that amount to significantly more than the judicial exception.
Regarding claim 7
Claim 7– Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is to a process.
Claim 7 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“initializing a plurality of weights based on of a distribution with a positive average;” – for purposes of examination, “based on of a distribution” is understood as “based on a distribution”, consistent with the separate claim objection. This step recites a mathematical concept under MPEP § 2106.04(a)(2)(I) because it initializes numerical weights according to a statical distribution characterized by a positive mathematical average.
“calculating an output according to a monotonically increasing function from a monotonic neural network to which the plurality of weights is applied;” – the recited operation expressly requires evaluation of a mathematical function and therefore recites a mathematical relationship and mathematical calculation under MPEP § 2106.04(a)(2)(I).
“and calculating a cumulative intensity function based on the calculated output.” – this step requires a calculation of a mathematical function from another calculated numerical output and therefore recites a mathematical relationship and mathematical calculation under MPEP § 2106.04(a)(2)(I).
Claim 7 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no additional elements that integrate the judicial exception into a practical application. The additional elements:
“a monotonic neural network to which the plurality of weights is applied” – the specification expressly describes the monotonic neural-network as a mathematical model modeled to calculate a monotonically increasing function. Thus, as claimed, the neural network provides the mathematical-environment in which the identified mathematical calculations are carried out. See MPEP § 2106.05(h).
Claim 7 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no additional elements that amount to significantly more than the judicial exception. The additional elements are:
“a monotonic neural network to which the plurality of weights is applied” – the specification’s background states that neural networks were known for modeling point processes and that monotonic neural networks had previously been proposed. This provides express factual support for treating the neural-network environment, at the level of generality claimed, as well-understood, routine, and conventional (WURC). See MPEP § 2106.05(d). Further, the monotonic neural network merely supplies the environment in which the identified mathematical concepts are performed and therefore does not provide an inventive concept separate from that exception. See MPEP § 2106.05(h).
Regarding claim 8
Claim 8 – Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
“A storage medium storing a program for causing a computer to execute” – the claim does not limit the storage medium to a non-transitory storage medium. The Specification states broadly that storage medium 15 is a medium configured to accumulate information by electrical, magnetic, optical, mechanical, or chemical action, but does not expressly limit the claimed term to non-transitory media or expressly disclaim propagating signals. Accordingly, the broadest reasonable interpretation encompasses a transitory signal carrying the program. A transitory propagating signal is not a process, machine, manufacture, or composition of matter. See In re Nuijten, 500 F.3d 1346, 1354, 84 USPQ2d 1495, 1500 (Fed. Cir. 2007); MPEP § 2106.03. Therefore, claim 8 fails Step 1.
For completeness, the following Step 2A/Step SB analysis assumes a statutory non-transitory embodiment of the claimed storage medium.
Claim 8 – Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
“initializing a plurality of weights based on of a distribution with a positive average” – this recites the same mathematical relationship and mathematical calculations under MPEP § 2106.04(a)(2)(I) discussed with respect to claim 7.
“calculating an output according to a monotonically increasing function from a monotonic neural network to which the plurality of weights is applied;” – the operation of calculating an output according to a monotonically increasing function recites a mathematical relationship and mathematical calculation under MPEP § 2106.04(a)(2)(I).
“and calculating a cumulative intensity function based on the cumulated output.” – this recites a mathematical relationship and mathematical calculation under MPEP § 2106.04(a)(2)(I).
Claim 8 – Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. There are no additional elements that integrate the judicial exception into a practical application.
Claim 8 – Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. There are no additional elements that amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 4, 7, and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Takahiro Omi (Fully Neural Network based Model for General Temporal Point Processes) in view of Seungil You (Deep Lattice Networks and Partial Monotonic Functions) and further in view of Oleksandr Shchur (Neural Temporal Point Processes: A Review).
Regarding claim 1, Omi in view of You and further in view of Shchur, teach an information processing apparatus comprising:
“” – Omi teaches this limitation in part. Omi teaches a plurality of neural-network weights associated with connections in the cumulative-hazard-function network and requires particular weights to have positive values:
“the weights of the particular network connections are constrained to be positive” (Omi, pg. 4, § 2.4 The proposed model)
Omi further teaches that the positive-constrained weights include the connections from elapsed time to the first hidden layer and the connections from the hidden layers. (See Omi, pg. 4, § 2.4 The proposed model, Figure 1).
Omi further teaches:
“In order to enforce the weights to be positive, if a weight is updated to be a negative value during training, we replace it with its absolute value.” (Omi, pg. 4, n.3)
Thus, Omi teaches the plurality of weights to which the claimed initialization is applied. However, Omi does not teach the distribution or algorithm used to initialize those weights and therefore does not disclose the corresponding Rule Y initialization algorithm.
“a monotonic neural network to which the plurality of weights is applied;” – Omi teaches this limitation. Omi teaches a feedforward cumulative-hazard-function network configured to produce an output that is monotonically increasing with respect to elapsed time:
“we model the cumulative hazard function using a feedforward neural network (a cumulative hazard function network; Fig. 1) for flexible modeling.” (Omi, pg. 4, § 2.4 The proposed model)
Thus, Omi teaches that positive-constrained weights are applied to a monotonic neural network, satisfying this limitation.
Omi does not teach these limitations and/or portions of:
“an initialization unit configured to initialize … based on a distribution with a positive average;”
“and a first calculation unit configured to calculate a cumulative intensity function based on an output from the monotonic neural network.”
You, however, teaches these limitations and/or portions of:
“an initialization unit configured to initialize … based on a distribution with a positive average;” – You teaches a computer-implemented monotonic deep-learning model having linear-embedding layers subject to element-wise nonnegativity constraints:
“we initialize each component in the linear embedding matrix with IID Gaussian noise N(2, 1)” (You, pg. 6, § 5 Numerical Optimization Details for the DLN – Initialization)
The Gaussian distribution N(2, 1) has an average or mean of 2, which is positive. You expressly explains that the positive mean biases the initial parameters toward positive values so that the parameters are not set to zero during the first monotonicity-enforcement projection:
“The initial mean of 2 is to bias the initial parameters to be positive so that they are not clipped to zero by the first monotonicity projection.” (You, pg. 6, § 5 Numerical Optimization Details for the DLN – Initialization)
Thus, You teaches a programmed initialization algorithm that generates respective initialization values according to a normal distribution having a positive mean and applies the values to a plurality of monotonicity-constrained weights. That algorithm is the same as or equivalent to the positive-mean normal-distribution branch of the Rule Y corresponding algorithm.
Neither Omi or You teach these remaining limitations and/or portions of:
“and a first calculation unit configured to calculate a cumulative intensity function based on an output from the monotonic neural network.”
Shchur, in combination with Omi, renders obvious these remaining limitations and/or portions of:
“and a first calculation unit configured to calculate a cumulative intensity function based on an output from the monotonic neural network.” – Shchur teaches that a cumulative hazard function defining a valid inter-event-time distribution must be differentiable, increasing, and satisfy the zero-time condition:
“First, the inter-event times are by definition strictly positive, so it should hold that
∫
-
∞
0
f
i
*
u
d
u
=
0
or, equivalently,
Φ
i
*
0
=
0
.” (Shchur, § 3.3 Predicting the Time of the Next Event)
Shchur expressly establishes that a valid cumulative hazard function must satisfy the zero-time condition
Φ
i
*
0
=
0
. Omi, however, sets
Φ
τ
h
i
=
Z
i
(
τ
)
without requiring
Z
i
0
=
0
. Accordingly, applying Shchur’s expressly stated zero-time condition to Omi’s parameterization reveals that Omi’s cumulative function may fail to equal zero at
τ
=
0
.
It would have been obvious to one of ordinary skill in the art to modify Omi’s cumulative-output calculation to satisfy the zero-time condition expressly identified by Shchur.
Omi itself supplies the mathematical relationship that determines the modification. Omi teaches both:
“
Φ
τ
|
h
i
=
∫
0
τ
ϕ
s
h
i
d
s
” (Omi, pg. 3, eq. 9)
Omi then models the cumulative function using neural-network output
Z
i
(
τ
)
, and obtains the intensity by differentiating that output:
∂
Z
i
(
τ
)
∂
τ
(See Omi, pg. 4, Eqs. (12)-(13); pg. 5).
Omi nevertheless sets
Φ
τ
h
i
=
Z
i
(
τ
)
, without subtracting
Z
i
(
0
)
. See Omi, pg. 3-4, Eq. (9), (12), and (13).
The subtraction is derived from Omi’s equations:
Φ
τ
|
h
i
=
∫
0
τ
ϕ
s
|
h
i
d
s
,
and:
ϕ
s
|
h
i
=
∂
Z
i
(
s
)
∂
s
.
Substituting and evaluating the definite integral gives:
Φ
τ
|
h
i
=
∫
0
τ
∂
Z
i
(
s
)
∂
s
d
s
=
Z
i
τ
-
Z
i
(
0
)
.
The endpoint-subtraction rule was indisputably conventional mathematical knowledge before the filing date.
This modification:
forces
Φ
0
=
0
;
preserves monotonicity because
Z
τ
≥
Z
0
for
τ
≥
0
;
preserves the intensity because
d
d
τ
Z
τ
-
Z
0
=
Z
'
(
τ
)
;
leaves Omi’s architecture and automatic-differentiation procedure unchanged; and
performs the same subtraction algorithm disclosed for the application’s first calculation unit.
The resulting operation
Z
i
τ
-
Z
i
(
0
)
performs the same algorithmic sequence as the corresponding Formula (3) algorithm disclosed in Applicant’s specification at ¶¶ [0241]-[0243]: evaluating the monotonic-neural-network output at the current time, evaluating the output at time zero, and subtracting the time-zero output from the current-time output.
The modification therefore would not have required changing Omi’s neural-network architecture, positive-weight constraint, or automatic-differentiation procedure. One of ordinary skill would have reasonably expected the modification to succeed because it is the direct endpoint evaluation of Omi’s own definite integral and produces the predictable result of anchoring the cumulative function at zero. Accordingly, Omi as modified in view of Shchur discloses a programmed computer performing the identical claimed first-calculation function through the same Formula (3) corresponding algorithm.
It would have been obvious to one of ordinary skill in the art before the effective filing date to initialize Omi’s positive-constrained weights using You’s positive-mean Gaussian distribution. Omi requires the weights to remain positive but does not specify their initial values, while You teaches positive-mean initialization to bias monotonicity-constrained parameters toward the permitted nonnegative region. Applying You’s known initialization procedure would not have altered Omi’s architecture or positive-weight constraint and would predictably have reduced immediate correction of negative initial values. Because both references concern computer-implemented, gradient-trained models having monotonicity-related weight constraints, one of ordinary skill would have reasonably expected the modification to succeed.
Regarding claim 2, Omi in view of You and further in view of Shchur, teach the information processing apparatus according to claim 1, further comprising
“a second calculation unit configured to calculate an intensity function related to a point process based on the calculated cumulative intensity function.” – Omi teaches this limitation. Omi characterizes a temporal point process using conditional intensity function:
“
λ
t
H
t
=
ϕ
t
-
t
i
h
i
)
” (Omi, pg. 3, Eq. 5)
Omi defines cumulative hazard function
Φ
τ
|
h
i
as the integral of the hazard or conditional-intensity function:
“
Φ
τ
|
h
i
=
∫
0
τ
ϕ
s
h
i
d
s
” (Omi, pg. 3, Eq. 9)
and calculates the hazard or intensity function by differentiating the cumulative function:
“
ϕ
τ
h
i
=
∂
∂
τ
Φ
(
τ
|
h
i
)
” (Omi, pg. 4, Eq. 10)
Omi further expresses the cumulative function using neural-network output
Z
i
(
τ
)
and calculates the intensity according to:
“
ϕ
τ
h
i
=
∂
Z
i
(
τ
)
∂
τ
” (Omi, pg. 4, Eq. 13)
Omi expressly states that this derivative is computed using automatic differentiation, which can be:
“easily carried out using neural network libraries such as TensorFlow and
PyTorch.” (Omi, pg. 5, §2.4 The proposed model)
As explained regarding claim 1, Omi’s cumulative function would have been modified in view of Shchur to calculate
Z
i
τ
-
Z
i
(
0
)
. Omi’s automatic-differentiation procedure applies to that modified cumulative function without further change because
Z
i
(
0
)
is constant with respect to
τ
:
∂
∂
τ
Z
i
τ
-
Z
i
0
=
∂
Z
i
τ
∂
τ
=
ϕ
τ
h
i
)
.
Under the previously stated interpretation to 35 U.S.C. § 112(f), Omi therefore discloses the same corresponding algorithm: receiving or accessing the calculated cumulative intensity function, automatically differentiating that function with respect to time, and outputting the resulting point-process intensity function. Omi’s programmed computer performs the identical claimed function in the same way and produces the same result as the corresponding algorithm disclosed for automatic differentiation units 24-3 and 44-3.
Regarding claim 4, Omi in view of You and further in view of Shchur, teach the information processing apparatus according to claim 1, wherein
“the initialization unit is configured to generate a random number based on the distribution, and initialize each of the plurality of weights with the random number.” – Omi does not teach this limitation. You, however, teaches this limitation. You teaches initializing each component of a linear-embedding weight matrix with an independent random value generated according to a Gaussian distribution:
“For linear embedding layers, we initialize each component in the linear embedding matrix with IID Gaussian noise N(2, 1)” (You, pg. 6, § 5 Numerical Optimization Details for the DLN – Initialization)
For purposes of examination, the phrases “generate a random number based on the distribution” and “initialize each of the plurality of weights with the random number” are interpreted as encompassing generation of a respective random number for each weight and initialization of each weight with its corresponding generated random number.
You teaches initializing each component of its linear-embedding weight matrix with an independent random value generated according to Gaussian distribution N(2, 1). The matrix components are trainable weights and collectively constitute a plurality of weights. Each IID Gaussian value is therefore a random number generated according to the distribution and applied to its corresponding weight.
Under the previously stated §112(f) interpretation, You’s programmed initialization performs the same random-number branch of the Rule Y corresponding algorithm: generating respective random initialization values according to a positive-mean distribution and applying the generated values to the respective weights. Thus, You teaches the limitation added by claim 4.
The rationale for modifying Omi using You’s initialization procedure stated with respect to claim 1 applies equally here. Claim 4’s random-number operations are part of the same initialization procedure already relied upon for claim 1; therefore, no additional reference or separate rationale beyond that stated for claim 1 is required.
Regarding claims 7 and 8
Claims 7 and 8 recite method and stored-program limitations corresponding generally to the weight-initialization, monotonic-network-output, and cumulative-intensity-calculation operations discussed above. For purposes of examination, the phrase “based on of a distribution” in claims 7 and 8 is understood as “based on a distribution”, consistent with the separate claim objection below.
Unlike claim 1, claims 7 and 8 do not recite the nonce-term “unit” limitations construed above under 35 U.S.C. § 112(f), and their process limitations do not otherwise invoke §112(f). Accordingly, the cumulative-intensity limitations of claims 7 and 8 are not limited to the corresponding Formula (3) subtraction algorithm.
Claim 7 recites initializing a plurality of weights based on a distribution having a positive average, calculating a monotonically increasing output from a monotonic neural network to which the weights are applied, and calculating a cumulative intensity function based on that output.
Omi teaches a plurality of positive-constrained weights applied to a neural network whose output
Z
i
(
τ
)
is monotonically increasing with respect to elapsed time. Omi further teaches calculating cumulative hazard function
Φ
τ
h
i
)
based on that output according to
Φ
τ
h
i
)
=
Z
i
(
τ
)
.
Shchur further establishes that, in temporal point processes, the cumulative hazard function is the integral of the hazard function and that the hazard function is commonly referred to as the intensity function in the temporal-point-process literature. Thus, Omi’s cumulative-hazard function corresponds to the claimed cumulative intensity function.
Therefore, Omi in view of Schur teaches calculating a cumulative intensity function based on the calculated output.
You teaches initializing each component of a monotonicity-constrained linear-embedding weight matrix with IID Gaussian noise N(2, 1), which is a distribution having a positive average or mean of 2. (You, pg. 6, §5 Numerical Optimization Details for the DLN – Initialization)
It would have been obvious to initialize Omi’s positive-constrained weights using You’s positive-mean Gaussian distribution for the reasons discussed above with respect to claim 1. Applying You’s known initialization procedure would have placed Omi’s initial weights preferentially within the positive region required to preserve monotonicity and would have predictably reduced immediate correction of negative initial values.
Accordingly, claim 7 is obvious over Omi in view of You and further in view of Shchur.
Claim 8 recites a storage medium storing a program that causes a computer to perform the same initialization, monotonic-network-output, and cumulative-intensity-calculation operations recited in claim 7. Omi expressly teaches computer-executable source code implementing its model and execution using neural-network software such as TensorFlow and PyTorch.
Omi in view of Schur, for the same reasons discussed above with respect to claims 1 and 7, teaches “calculating a cumulative intensity function based on the calculated output”.
It would have been obvious to store Omi’s executable program instructions, as modified to include You’s initialization procedure, on a computer-readable storage medium so that the instructions could be loaded and executed by a computer. The storage-medium implementation changes only the manner in which the previously established instructions are embodied and predictably permits the disclosed computer to execute those instructions.
Accordingly, claim 8 is obvious over Omi in view of You and further in view of Shchur.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Omi in view of Shchur and further in view of Alexander Keller (US20220129755A1).
Regarding claim 3, Omi in view of Shchur and further in view of Keller, teach the information processing apparatus according to claim 1. In particular, Omi teaches the plurality of weights and monotonic neural networks, and Omi as modified in view of Shchur, performs the previously identified Formula (3) corresponding algorithm for the first calculation unit.
The combination of Omi and Shchur does not teach these remaining limitations and/or portions of:
“an initialization unit configured to initialize ”
“wherein the initialization unit is configured to initialize each of the plurality of weights with a positive fixed value.”
Keller, however, teaches these remaining limitations and/or portions of:
“an initialization unit configured to initialize ” – Keller teaches initializing selected neural-network weights before training by assigning a small positive constant to each selected weight. Keller’s Table 2 sets InitialWeight = 0.01f and assigns that value to each selected path weight:
“const float InitialWeight = 0.01f;” (Keller, pg. 18, Table 2)
“for (int i = 0; i < Paths; ++i)
Weight [1] [i] = InitialWeight;” (Keller, pg. 18, Table 2)
The specification expressly identifies positive-fixed-value initialization as an implementation of the claimed initialization based on a distribution having a positive average. Keller therefore teaches the inherited initialization limitation of claim 1.
“wherein the initialization unit is configured to initialize each of the plurality of weights with a positive fixed value.” – For the same reasons stated in the inherited claim 1 limitation, Keller also teaches the additional limitation of claim 3. Keller further teaches that the selected connection weights may be initialized to a positive constant, such as the inverse of the number of connections of a neural unit, while unselected connections are implicitly zero. (Keller, pg. 18, ¶[0214], Table 2; pgs. 19-20, ¶¶[0231]-[0235]).
Under the previously stated §112(f) interpretation, Keller’s programmed computer performs the same positive-fixed-value branch of the Rule Y corresponding algorithm: identifying the selected plurality of weights, assigning the same positive fixed value to each weight, and applying the initialized weights to the neural network. Keller’s value of 0.01 is positive, and initialization of the plurality with that value has a positive average of 0.01. Keller therefore teaches both the inherited initialization limitation and the additional limitation of claim 3.
It would have been obvious to one of ordinary skill in the art before the effective filing date to implement Omi’s positive-constrained monotonic neural network using Keller’s selected-path representation and initialize each selected positive-constrained weight with Keller’s small positive constant. Omi requires the relevant weights to remain positive to preserve monotonicity but does not specify their initial values. Keller teaches that its selected-path representation reduces network complexity and permits the selected weights to be initialized deterministically with the same positive constant.
The modification would place each selected Omi weight initially within Omi’s required positive region. It also would preserve Omi’s monotonic operation because the selected path weights remain positive, while omitted or zero-valued connections introduce no negative dependence on elapsed time. Keller further states that the resulting pattern of zero and positive nonzero connection values is sufficiently uniformly distributed to permit convergence during training.
One of ordinary skill therefore would have reasonably expected the modification to succeed and to produce the predictable result of a monotonic neural network in which each weight of the selected plurality begins at a positive fixed value.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Omi in view of You further in view of Shchur and further in view of Kaiming He (Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification (arXiv:1502.01852v1)).
Regarding claim 5, Omi in view of You further in view of Shchur and further in view of He, teach the information processing apparatus according to claim 1, wherein
“the distribution is a normal distribution in which an average is a first value and a standard deviation is a second value, the first value is a positive real number, ” – Omi does not teach this limitation. You, however, teaches this limitation in part. You teaches initializing each component of a linear-embedding weight matrix with IID Gaussian noise N(2, 1):
“we initialize each component in the linear embedding matrix with IID Gaussian noise N(2, 1)” (You, pg. 6, § 5 Numerical Optimization Details for the DLN – Initialization)
The initialization distribution is therefore a normal distribution having an average or mean of 2, which is a positive real number, and a standard deviation defining the spread of the distribution.
You does not teach these limitations and/or portions of:
“”
He, however, teaches these remaining limitations and/or portions of:
”… and the second value is a positive square root of a value obtained by dividing a positive real number by a number of nodes of the monotonic neural network.” – He teaches selecting the standard deviation of a Gaussian weight-initialization distribution according to the number of inputs or connections associated with a layer. He defines (n1) as the number of connections associated with a response in layer (l) and derives:
“
1
2
n
l
V
a
r
w
l
=
1
,
∀
l
” (He, pg. 4, Eq. (10))
Solving gives:
V
a
r
w
l
=
2
n
l
and therefore:
σ
l
=
V
a
r
w
l
=
2
n
l
which leads to a zero-mean Gaussian distribution whose:
“standard deviation (std) is
2
n
l
” (He, pg. 4, § Forward Propagation Case; Eq. (10))
He also explains that the Xavier initialization may be implemented as a zero-mean Gaussian distribution having standard deviation:
“The ‘Xavier’ initialization [7] … only considers the linear case, and its result is given by
n
l
V
a
r
w
l
=
1
(the forward case), which can be implemented as a zero-mean Gaussian distribution whose std is
1
n
l
.” (He, pg. 5, § Comparisons with “Xavier” Initialization [7])
Thus, He, teaches a standard deviation equal to the positive square root of a positive real number (1 or 2, dived by
n
l
). For purposes of examination, and without withdrawing the rejection under 35 U.S.C. 112(b), “a number of nodes of the monotonic neural network” is interpreted as the number of input nodes, or fan-in, of the layer whose weights are initialized. In a fully connected layer, that node count corresponds to the number of input connections
n
l
identified by He.
It would have been obvious to retrain You’s positive mean while adopting He’s fan-independent standard deviation. You teaches selecting a positive mean to bias monotonicity-constrained parameters toward their permitted positive region. He teaches making the spread of a Gaussian initialization depend on the layer’s fan-in account for layer size and address changes in signal and gradient magnitudes through successive layers. Because the man and standard deviation are separately selectable parameters of a Gaussian distribution, one of ordinary skill would have retained You’s positive mean while using He’s known fan-in-dependent spread.
The resulting distribution would have the form:
W
=
μ
+
c
n
l
Z
,
where:
μ
>
0
,
c
>
0
,
Z
~
N
(
0
,
1
)
.
Accordingly:
E
W
=
μ
>
0
and:
S
t
d
W
=
c
n
l
This predictably produces a normal initialization distribution having both the positive average taught by You and the layer-size-dependent standard deviation taught by He.
Under the previously stated §112(f) interpretation, the combined references perform the same or equivalent Rule Y algorithm disclosed at ¶[0237]: generating normal-distribution initialization values having a positive average and a standard deviation equal to a positive square root of a positive real number divided by a layer-specific node count, and applying the generated values to the plurality of weights.
The modification changes only the parameters of the Gaussian initialization and predictably produces positive-biased initialization values having the claimed fan-in dependent spread. Therefore, Omi in view of You, further in view of Shchur, and further in view of He teaches or renders obvious every limitation of claim 5.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Omi in view of Xavier Glorot ("Understanding the Difficulty of Training Deep Feedforward Neural Networks") and further in view of Shchur.
Regarding claim 6, Omi in view of Glorot and further in view of Shchur, teach the information processing apparatus according to claim 1, wherein
“the distribution is a uniform distribution in which a minimum value is a third value and a maximum value is a fourth value, the third value is a real number of 0 or more, and the fourth value is a positive real number.” – Omi in view of Shchur teaches or renders obvious the limitations inherited from claim 1 other than the claimed initialization, for the reasons discussed regarding claim 1. Omi does not expressly teach the recited nonnegative uniform initialization. Glorot, however, teaches that limitation in part (“an initialization unit configured to initialize ”), as well as teaching or rendering obvious the added limitations of claim 6.
Glorot teaches generating each initial neural-network weight
W
according to the bounded uniform distribution:
“We initialized the biases to be 0 and the weights
W
i
j
at each layer with the following commonly used heuristic:
W
i
j
~
U
[
-
1
n
,
1
n
]
, where
U
[
-
a
,
a
]
is the uniform distribution in the interval (-a, a) and n is the size of the previous layer (the number of columns of W).” (Glorot, pg. 251, §2.3 Experimental Setting
Glorot further teaches an initialization procedure, called ‘normalized initialization’:
“
W
~
U
-
6
n
j
+
n
j
+
1
,
6
n
j
+
n
j
+
1
“ (Glorot, pg. 253, §4.2.1, Eq. 16)
where
n
j
and
n
j
+
1
are the sizes of adjacent layers. Glorot’s distribution has a positive upper bound, but its lower bound is negative and its average is zero.
Letting
a
=
6
n
j
+
n
j
+
1
, Glorot’s Eq. 16 may be written as:
W
~
U
-
a
,
a
,
a
=
6
n
j
+
n
j
+
1
>
0
Omi constrains the relevant weights of its monotonic neural network to be positive and states that, if a constrained weight is updated to a negative value during training, the weight is replaced with its absolute value. (Omi, pg. 4, §2.4, Figure 1 and n.3). Under the proposed modification, the same positivity-enforcement operation would be applied immediately after Glorot’s initialization and before the initialized weight is used by Omi’s monotonic neural network:
W
'
=
W
,
-
W
,
,
W
≥
0
,
W
<
0
,
=
W
The resulting distribution of
W
'
can be determined directly.
For
0
≤
y
≤
a
:
P
W
'
≤
y
=
P
W
≤
y
=
P
(
-
y
≤
W
≤
y
)
=
(
2
y
)
(
2
a
)
=
y
a
Because
y
a
is the cumulative distribution function of
U
0
,
a
:
W
'
=
~
U
0
,
a
Consequently:
min
W
'
=
0
,
max
W
'
=
a
>
0
,
and:
E
W
'
=
0
+
a
2
=
a
2
>
0
.
The resulting distribution therefore has a nonnegative minimum, a positive maximum, and a positive average. This satisfies both the inherited positive-average limitation and claim 6’s minimum-and-maximum limitations.
It would have been obvious to one or ordinary skill in the art before the effective filing date to initialize Omi’s positive-constrained weights using Glorot’s bounded uniform initialization and to apply Omi’s existing absolute-value operation to any negative initial value. Omi requires the relevant weights to be positive to preserve monotonicity, while Glorot’s initialization can generate negative values. Beginning training with weights satisfying the same constraint enforced during training would have been a predictable implementation choice. The modification would retain a bounded uniform random initialization and would not alter Omi’s network architecture or monotonicity constraint.
Under the previously stated interpretation pursuit to 35 U.S.C. § 112(f), the modified programmed computer performs the identical claimed initialization function and uses an algorithm interchangeable with the uniform-distribution branch of the disclosed Rule Y algorithm. Directly generating values from (
U
[
0
,
a
]
), and generating values from (
U
[
-
a
,
a
]
) followed by the absolute-value operation, produce final initialization values having the same distribution and apply those values to the selected weights. The difference between the two generation procedures is therefore insubstantial for the claimed function and result.
For the reasons discussed regarding claim 1, Shchur teaches that a valid cumulative intensity must satisfy the zero-time condition. One of ordinary skill also would have modified Omi to calculate the cumulative function as:
Φ
τ
h
i
)
=
Z
i
τ
-
Z
i
(
0
)
,
thereby satisfying the condition and performing the corresponding Formula (3) algorithm associated with the claimed first calculation unit.
Accordingly, Omi in view of Glorot and further in view of Shchur teaches or renders obvious every limitation of claim 6, and claim 6 would have been obvious to one of ordinary skill in the art before the effective filing date.
Claim Objections
Claims 7 and 8 are objected to because of the following informalities:
Each of claims 7 and 8 recites:
“initializing a plurality of weights based on of a distribution with a positive average.”
The phrase “based on of a distribution” is grammatically improper because it contains an extraneous occurrence of the word “of”. The intended meaning is otherwise reasonably ascertainable, and the informality does not render the claims indefinite. Appropriate correction is required, for example, by amending “based on of a distribution” to “based on a distribution”.
Specification
The disclosure is objected to because of the following informalities:
Paragraph [0412] – Incorrect reference sign
Paragraph [0412] identifies:
“53B-2 Automatic differentiation unit”
Reference sign 53B-2 is already identified in paragraph [0411] as a cumulative intensity function calculation unit. Figures 22 and 23 identify the automatic differentiation unit associated with the second intensity-function calculation unit as 53B-3. Paragraph [0412] should therefore be corrected by changing the final reference sign “53B-2” in the automatic-differentiation-unit listing to “53B-3”. Appropriate correction is required.
Paragraph [0417] – incorrect reference sign
Paragraph [0417] identifies:
“25-3 … Optimization unit.”
Figure 2 and the corresponding detailed description identify the evaluation-function calculation unit as 25-1 and the optimization unit as 25-2. The reference to “25-3” in paragraph [0417] is therefore inconsistent with the remainder of the disclosure. Paragraph [0417] should be corrected by changing “25-3” to “25-2”. Appropriate correction is required.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Paul Coleman whose telephone number is (571)272-4687. The examiner can normally be reached Mon-Fri.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAUL COLEMAN/ Examiner, Art Unit 2126
/DAVID YI/ Supervisory Patent Examiner, Art Unit 2126