DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. IN202241073661, filed on 12/19/2022.
Information Disclosure Statement
The information disclosure statement(s) (IDS) submitted on 02/01/2024, 06/03/2024, 02/14/2025, 10/16/2025, and 02/19/2026 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) is/are being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claims 1-7 are directed to a process. Claims 8-20 are directed to a machine or an article of manufacture.
With respect to claim(s) 1, 8, and 16:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
determining/determine […] a predicted probability for each of the plurality of contents of the input data in an output of the AI model; (Mathematical concepts – The broadest reasonable interpretation for this limitation can be an AI model outputting a probability prediction for each of the plurality of contents of the input data, and thus involves mathematical calculations – see MPEP § 2106.04(a)(2)(I))
(Claims 1 and 16) determining/determine […] a neural loss of the AI model by comparing the predicted probability for each of the plurality of contents with a predefined desired probability for each of the plurality of contents; (Mathematical concepts – determining a neural loss by comparing a predicted probability with predefined desired probability (e.g., ground truth) involves mathematical calculations (see paragraph [0043]) – see MPEP § 2106.04(a)(2)(I))
(Claims 1 and 16) determining/determine […] a symbolic loss for the AI model by comparing the predicted probability for each of the plurality of contents with a pre-determined undesired probability for each of the plurality of contents; (Mathematical concepts – determining a symbolic loss by comparing a predicted probability with a pre-determined undesired probability involves mathematical calculations (see paragraph [0044]) – see MPEP § 2106.04(a)(2)(I))
(Claim 8) determine a training loss of the AI model as a measure of the neural loss and the symbolic loss (Mathematical concepts – Determining a training loss as a measure of the neural loss and the symbolic loss involves mathematical calculations – see MPEP § 2106.04(a)(2)(I))
determining/determine […] weights of a plurality of layers of the AI model; (Mathematical concepts – The broadest reasonable interpretation for this limitation can be computing the weights for the layers of the AI model, and thus involves mathematical calculations – see MPEP § 2106.04(a)(2)(I))
updating/update […] the weights of the plurality of layers of the AI model based on the neural loss and the symbolic loss. (Mathematical concepts – Updating the weights of layers of the AI model based on losses involves mathematical calculations – see MPEP § 2106.04(a)(2)(I))
If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process, but for the recitation of generic computer components, then the claim limitations fall within the mathematical or mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 1) A method for neuro-symbolic learning of an artificial intelligence (AI) model, the method comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 8) An electronic device for neuro-symbolic learning of an artificial intelligence (AI) model, the electronic device comprising: a processor; communicator; a neuro-symbolic AI controller; and memory storing one or more programs including computer-executable instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 16) One or more non-transitory computer-readable storage media storing one or more computer programs including computer-executable instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform operations, the operations comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 1) receiving, by an electronic device, input data comprising a plurality of contents for the neuro-symbolic learning of the AI model; (Mere data gathering – Adding insignificant extra-solution activity of mere data gathering to the judicial exception – see § MPEP2106.05(g).)
(Claims 8 and 16) receive input data comprising a plurality of contents for the neuro-symbolic learning of the AI model (Mere data gathering – Adding insignificant extra-solution activity of mere data gathering to the judicial exception – see § MPEP2106.05(g).)
(Claim 1) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea.
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 1) A method for neuro-symbolic learning of an artificial intelligence (AI) model, the method comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 8) An electronic device for neuro-symbolic learning of an artificial intelligence (AI) model, the electronic device comprising: a processor; communicator; a neuro-symbolic AI controller; and memory storing one or more programs including computer-executable instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 16) One or more non-transitory computer-readable storage media storing one or more computer programs including computer-executable instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform operations, the operations comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 1) receiving, by an electronic device, input data comprising a plurality of contents for the neuro-symbolic learning of the AI model; (Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception (WURC)- see MPEP § 2106.05(d)(ll)(i) - Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information).)
(Claims 8 and 16) receive input data comprising a plurality of contents for the neuro-symbolic learning of the AI model (Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception (WURC)- see MPEP § 2106.05(d)(ll)(i) - Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information).)
(Claim 1) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
With respect to claim(s) 2, 9, and 17:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
(Claim 2) wherein the updating […] the weights of the plurality of layers of the AI model based on the neural loss and the symbolic loss comprises: (Mathematical concepts – Updating the weights of layers of the AI model based on losses involves mathematical calculations – see MPEP § 2106.04(a)(2)(I))
determining/determine […] a training loss of the AI model as a measure of the neural loss and the symbolic loss; (Mathematical concepts – Determining a training loss as a measure of the neural loss and the symbolic loss involves mathematical calculations – see MPEP § 2106.04(a)(2)(I))
updating/update […] the weights of the plurality of layers of the AI model in proportion to the determined training loss. (Mathematical concepts – Updating the weights of layers of the AI model in proportion to the determined training loss involves mathematical calculations – see MPEP § 2106.04(a)(2)(I))
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 2) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 9) wherein the one or more programs further include instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 17) the operations further comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 2) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 9) wherein the one or more programs further include instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 17) the operations further comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 3, 10, and 18:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
(Claim 3) wherein the pre-determined undesired probability for each of the plurality of contents in the input data is determined […], wherein the determining of the pre-determined undesired probability for each of the plurality of contents comprises: (Mental process – A person can determine an undesired probability for each of the plurality of contents via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
selecting/select […] an external symbolic knowledge graph comprising a plurality of common-sense facts; (Mental process – A person can mentally select an external symbolic knowledge graph or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
constructing/construct […] a negative knowledge graph comprising a plurality of violating common-sense facts from the external symbolic knowledge graph; (Mental process – A person can mentally construct (think of) a negative knowledge graph comprising a plurality of violating common-sense facts from the external symbolic knowledge graph or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
identifying/identify […] a set of maximally violating facts from the plurality of violating common-sense facts; (Mental process – A person can mentally identify a set of maximally violating facts from the plurality of violating common-sense facts – see MPEP § 2106.04(a)(2)(III))
(Claims 3 and 18) generating […] symbolic labels in a form of probabilities for the maximally violating facts, and wherein the maximally violating facts comprise the undesired probability. (Mental process – A person can mentally generate (think of) symbolic labels in a form of probabilities for the maximally violating facts, and wherein the maximally violating facts comprise the undesired probability or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
(Claim 10) generate symbolic labels for the maximally violating facts, wherein the maximally violating facts are undesired, and wherein the symbolic labels are provided in terms of probabilities. (Mental process – A person can mentally generate (think of) symbolic labels in a form of probabilities for the maximally violating facts, and wherein the maximally violating facts comprise the undesired probability or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 3) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 10) wherein the one or more programs further include instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 18) the operations further comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 3) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 10) wherein the one or more programs further include instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 18) the operations further comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 4, 11, and 19:
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
wherein the plurality of common-sense facts comprise at least one of a set of pre-defined rules or common-sense facts associated with a real-world. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the plurality of common-sense facts comprise at least one of a set of pre-defined rules or common-sense facts associated with a real-world. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 5 and 12:
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
wherein the maximally violating facts [(Claim 5) further] comprise facts which are against the plurality of common-sense facts with respect to the input data. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the maximally violating facts [(Claim 5) further] comprise facts which are against the plurality of common-sense facts with respect to the input data. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 6, 13, and 20:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
determining/determine […] the plurality of contents and relationships between the plurality of contents; (Mental process – A person can mentally determine the plurality of contents and relationships between the plurality of contents – see MPEP § 2106.04(a)(2)(III))
determining/determine […] at least one scene graph based on the plurality of contents and the relationships between the plurality of contents, wherein the at least one scene graph comprises a structural representation of the plurality of contents and the relationships between the plurality of contents in the input data. (Mental process – A person can mentally determine at least one scene graph based on the plurality of contents and the relationships between the plurality of contents or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 6) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 13) wherein the one or more programs further include instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 20) the operations further comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 6) […] by an electronic device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 13) wherein the one or more programs further include instructions that, when executed by the processor, cause the electronic device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
(Claim 20) the operations further comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 7 and 14:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
wherein the plurality of contents is determined […] and the relationships between the plurality of contents is determined […], adhering to common-sense of a real world through an external symbolic knowledge graph. (Mental process – A person can determine the plurality of contents and the relationship between the contents, adhering to common-sense of a real world through an external symbolic knowledge graph via mind or by using a pen and paper – see MPEP § 2106.04(a)(2)(III))
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
[…] using neural AI […] using symbolic AI […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
[…] using neural AI […] using symbolic AI […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 15:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
wherein the training loss is determined by applying various functions on both the neural loss and the symbolic loss, and wherein the various functions include addition, subtraction, multiplication, and division. (Mathematical concepts – Determining a training loss by applying various functions that include addition, subtraction, multiplication, and division involve mathematical calculations – see MPEP § 2106.04(a)(2)(I))
Additionally, the claim(s) do not recite any new additional elements that would amount to an integration of the abstract idea into a practical application (individually or in combination) or significantly more than the judicial exception.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 6-9, 13-17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over ZAREIAN ("Bridging Knowledge Graphs to Generate Scene Graphs") in view of KIM ("NLNL: Negative Learning for Noisy Labels") and MCCARTHY (US 20210216881 A1), hereafter ZAREIAN, KIM, and MCCARTHY respectively.
Regarding Claim 1:
ZAREIAN teaches:
A method for neuro-symbolic learning of an artificial intelligence (AI) model, the method comprising: (ZAREIAN [page 14, section 6 Conclusion] teaches: “We proposed a new method for Scene Graph Generation that incorporates external commonsense knowledge in a novel, graphical neural framework (i.e., neuro-symbolic).” ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1. Perhaps the most important aspect of computation is the run time when deploying the model on new images. To this end, we run each trained model on the entire test set of Visual Genome (VG), i.e. 26446 images, and get the average run time over all images in terms of seconds. […] We record the time it takes to train each model (i.e., method for neuro-symbolic learning of an artificial intelligence (AI) model) on one epoch of the VG training set, i.e. 56224 images, and get the average over 10 training epochs.” ZAREIAN [page 2, Fig. 1.] teaches: “An example of a Visual Genome image and its ground truth scene graph.”)
receiving, by an electronic device, input data comprising a plurality of contents for the neuro-symbolic learning of the AI model; (ZAREIAN [page 7, section 4 Method] teaches: “Given an image (i.e., receiving […] input data), our model first applies a Faster R-CNN [36] to detect objects (i.e., comprising a plurality of contents), and represents them as scene entity (SE) nodes.” ZAREIAN [page 14, section 6 Conclusion] teaches: “We proposed a new method for Scene Graph Generation that incorporates external commonsense knowledge in a novel, graphical neural framework (i.e., neuro-symbolic).” ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU (i.e., by the electronic device) of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1. Perhaps the most important aspect of computation is the run time when deploying the model on new images. To this end, we run each trained model on the entire test set of Visual Genome (VG), i.e. 26446 images, and get the average run time over all images in terms of seconds. […] We record the time it takes to train each model (i.e., for the neuro-symbolic learning of the AI model) on one epoch of the VG training set, i.e. 56224 images (i.e., input data comprising a plurality of contents), and get the average over 10 training epochs.”)
determining, by the electronic device, a predicted probability for each of the plurality of contents of the input data in an output of the AI model; (ZAREIAN [page 3, section 1 Introduction] teaches: “A scene graph node represents an entity or predicate instance in a specific image, while a commonsense graph node represents an entity or predicate class, which is a general concept independent of the image. Similarly, a scene graph edge indicates the participation of an entity instance (e.g. as a subject or object) in a predicate instance in a scene, while a commonsense edge states a general fact about the interaction of two concepts in the world.” ZAREIAN [page 10, section 4.2 Successive message passing and bridging] teaches: “To this end, we compute a pairwise similarity from each SE to all CE nodes, and from each SP to all CP nodes.
a
i
j
E
B
=
exp
x
i
S
E
,
x
j
C
E
E
B
∑
j
'
exp
x
i
S
E
,
x
j
'
C
E
E
B
,
w
h
e
r
e
x
,
y
E
B
=
ϕ
a
t
t
S
E
x
T
ϕ
a
t
t
C
E
y
,
(
11
)
and similarly for predicates,
a
i
j
P
B
=
exp
x
i
S
P
,
x
j
C
P
P
B
∑
j
'
exp
x
i
S
P
,
x
j
'
C
P
P
B
,
w
h
e
r
e
x
,
y
P
B
=
ϕ
a
t
t
S
P
x
T
ϕ
a
t
t
C
P
y
,
(
12
)
[…] We use each
q
i
j
E
B
to set the edge weight of the classifiedTo edge from
x
i
S
E
to
x
j
C
E
, as well as the hasInstance edge from
x
j
C
E
to
x
i
S
E
. Similarly we use each
a
i
j
P
B
to set the weight of edges between
x
i
S
P
and
x
j
C
P
. […] The final values of
a
i
j
E
B
and
a
i
j
P
B
are the outputs of our model, which can be used to classify each entity (i.e., determining […] a predicted probability) and predicate in the scene graph.” ZAREIAN [page 10, section 4.3 Training] teaches: “Then we use the output probability scores of each node (i.e., determining […] a predicted probability for each of the plurality of contents of the input data in an output of the AI model) to define a cross-entropy loss.” ZAREIAN [page 7, section 4 Method] teaches: “Given an image, our model first applies a Faster R-CNN [36] to detect objects (i.e., the plurality of contents of the input data), and represents them as scene entity (SE) nodes.”)
determining, by the electronic device, a neural loss of the AI model by comparing the predicted probability for each of the plurality of contents with a predefined desired probability for each of the plurality of contents; (ZAREIAN [page 10, section 4.3 Training] teaches: “We closely follow [2] which itself follows [52] for training procedure. Specifically, given the output and ground truth graphs, we align output entities and predicates to ground truth counterparts (i.e., a predefined desired probability for each of the plurality of contents). To align entities we use IoU and predicates will be aligned naturally since they correspond to aligned pairs of entities. Then we use the output probability scores of each node to define a cross-entropy loss (i.e., determining […] a neural loss of the AI model by comparing the predicted probability for each of the plurality of contents with a predefined desired probability for each of the plurality of contents). The sum of all node-level loss values will be the objective function to be minimized using Adam [18].” ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU (i.e., by the electronic device) of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1.”)
determining, by the electronic device, weights of a plurality of layers of the AI model; (ZAREIAN [page 12, section 5.2 Implementation details] teaches: “We use three-layer fully connected networks with ReLU activation for all trainable networks
ϕ
i
n
i
t
,
ϕ
s
e
n
d
,
ϕ
r
e
c
e
i
v
e
and
ϕ
a
t
t
.” ZAREIAN [page 10, section 4.3 Training] teaches: “Then we use the output probability scores of each node to define a cross-entropy loss. The sum of all node-level loss values will be the objective function to be minimized using Adam [18].” Examiner’s note: Adam updates the trainable parameters (i.e., determining […] weights) for the three-layer fully connected networks (i.e., of a plurality of layers of the AI model) based on the computed loss values.)
updating, by the electronic device, the weights of the plurality of layers of the AI model based on the neural loss […]. (ZAREIAN [page 10, section 4.3 Training] teaches: “Then we use the output probability scores of each node to define a cross-entropy loss. The sum of all node-level loss (i.e., based on the neural loss) values will be the objective function to be minimized (i.e., updating […] the weights of the plurality of the layers of the AI model) using Adam [18].”)
ZAREIAN is not relied upon for teaching:
updating […] the weights of the plurality of layers of the AI model based on […] the symbolic loss.
determining, by the electronic device, a symbolic loss for the AI model by comparing the predicted probability for each of the plurality of contents with a pre-determined undesired probability for each of the plurality of contents;
However, KIM teaches: determining […] a symbolic loss for the AI model by comparing the predicted probability […] with a pre-determined undesired probability […]; (KIM [page 103, section 3.1 Negative Learning] teaches: “In contrast, with NL, the CNNs are trained that “input image does not belong to this complementary label.” […] We consider the problem of c-class classification. Let
x
∈
X
be an input,
y
,
y
-
∈
Y
=
{
1
,
.
.
.
,
c
}
be its label and complementary label, respectively, and
y
,
y
-
∈
0
,
1
c
be their one-hot vector. Suppose the CNN
f
(
x
;
θ
)
maps the input space to the c-dimensional score space
f
:
X
→
R
c
, where
θ
is the set of network parameters. If
f
passes through the softmax function, the output can be interpreted as a probability
p
∈
∆
c
-
1
(i.e., the predicted probability), where
∆
c
-
1
denotes the c-dimensional simplex. When training with PL, the cross entropy loss function (i.e., determining […] a symbolic loss) of the network
f
(i.e., for the AI model) becomes:
L
f
,
y
=
-
∑
k
=
1
c
y
k
log
p
k
(
1
)
where
p
k
denotes the
k
t
h
element of
p
. Eq. 1 is suitable for optimizing the probability value corresponding to the given label as 1
(
p
y
→
1
)
, satisfying the purpose of PL. However, NL differs from PL as it optimizes the output probability corresponding to the complementary label to be far from 1 (to reach 0 in the end
(
p
y
-
→
0
)
). Therefore, we propose a loss function as follows:
L
f
,
y
-
=
-
∑
k
=
1
c
y
-
k
log
1
-
p
k
(
2
)
This complementary label is completely random in that it is selected randomly from the labels of all classes except for the given label y for every iteration during training (Algorithm 1). Eq. 2 enables the probability value of the complementary label to be optimized as zero, resulting in an increase in the probability values of other classes, meeting the purpose of NL.” Examiner’s note: KIM’s Equation (2) computes a loss
L
based on the predicted probability
p
y
-
assigned to the selected complementary class
y
-
k
. Under BRI, comparing the predicted probability […] with a pre-determined undesired probability can be interpreted as applying
-
log
1
-
p
y
-
, where the undesired probability is 0 and the predicted probability is
p
y
-
. KIM’s Negative Learning optimizes for
p
y
-
to reach 0 in the end.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of ZAREIAN and KIM before them, to include KIM’s negative learning techniques in ZAREIAN’s method for training a model using scene graphs and commonsense graphs. One of ordinary skill could apply KIM’s equation (2) to each of the nodes in ZAREIAN when generating scene graphs for training to prevent overfitting and achieve higher accuracy due to filtering based on KIM’s Selective Negative Learning and Positive Learning (SelNLPL), as disclosed in KIM [page 105, section 3.5 Semi-supervised learning]. One would have been motivated to make such a combination in order to perform Selective Negative Learning and Positive Learning (SelNLPL) to filter noisy data from training data and prevent overfitting, and train a model to determine images that do not belong to the complementary label (KIM [page 103, section 3. Method]).
ZAREIAN in view of KIM is not relied upon for teaching, but MCCARTHY teaches: updating […] the weights of the plurality of layers of the AI model based on the […] loss and the symbolic loss. (MCCARTHY [0029] teaches: “The symbolic loss function 442 quantifies symbolic prediction errors by comparing predictions of the symbolic task pipeline 204 to actual symbolic relationships indicated in training symbolic triples (in the embedding space). The numerical loss function 444 may be constructed as a regression loss that quantifies numerical prediction errors by comparing numerical predictions from the neural networks of the numerical pipeline 206 to actual numerals indicated in the training numerical triples (in value).” MCCARTHY [0030] teaches: “As further shown by 440 in FIG. 4, the symbolic loss function 442 (i.e., based on […] the symbolic loss) and numerical loss function 444 (i.e., based on the […] loss) are aggregated to calculate a joint loss. Such a joint loss is minimized iteratively by an optimizer 450 based on, for example, stochastic gradient descent techniques (e.g., Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, AMSGrad, and the like) (i.e., updating […] the weights of the plurality of layers of the AI model). The iterations shown by arrow 451 are performed through all triples of the training triple set, and for each triple of the training set triples, the optimization process iterates for adjusting the training parameters of the MTKGPM 202 to minimize the joint loss.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of ZAREIAN, KIM, and MCCARTHY before them, to include MCCARTHY’s joint training using numerical and symbolic loss functions in ZAREIAN and KIM’s method for training a model using scene graphs and commonsense graphs. One would have been motivated to make such a combination in order to provide an integrated two-layer multi-task prediction model trained from an existing knowledge graph and to jointly train the model by minimizing an aggregation of a symbolic loss function and a numerical regression loss function, thus predict new symbolic triples that may be previously unknown (or missing) so that they can be added to the knowledge graph to expand existing knowledge (MCCARTHY [0015] and [0023]).
Regarding Claim 2:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 1 as outlined above. ZAREIAN further teaches:
[…] neural loss […] (ZAREIAN [page 10, section 4.3 Training] teaches: “We closely follow [2] which itself follows [52] for training procedure. Specifically, given the output and ground truth graphs, we align output entities and predicates to ground truth counterparts. To align entities we use IoU and predicates will be aligned naturally since they correspond to aligned pairs of entities. Then we use the output probability scores of each node to define a cross-entropy loss (i.e., neural loss). The sum of all node-level loss values will be the objective function to be minimized using Adam [18].”)
MCCARTHY further teaches: wherein the updating, by the electronic device, the weights of the plurality of layers of the AI model based on the […] loss and the symbolic loss comprises: (MCCARTHY [0029] teaches: “The symbolic loss function 442 quantifies symbolic prediction errors by comparing predictions of the symbolic task pipeline 204 to actual symbolic relationships indicated in training symbolic triples (in the embedding space). The numerical loss function 444 may be constructed as a regression loss that quantifies numerical prediction errors by comparing numerical predictions from the neural networks of the numerical pipeline 206 to actual numerals indicated in the training numerical triples (in value).” MCCARTHY [0030] teaches: “As further shown by 440 in FIG. 4, the symbolic loss function 442 (i.e., based on […] the symbolic loss) and numerical loss function 444 (i.e., based on the […] loss) are aggregated to calculate a joint loss. Such a joint loss is minimized iteratively by an optimizer 450 based on, for example, stochastic gradient descent techniques (e.g., Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, AMSGrad, and the like) (i.e., updating […] the weights of the plurality of layers of the AI model). The iterations shown by arrow 451 are performed through all triples of the training triple set, and for each triple of the training set triples, the optimization process iterates for adjusting the training parameters of the MTKGPM 202 to minimize the joint loss.”)
determining, by the electronic device, a training loss of the AI model as a measure of the neural loss and the symbolic loss; (MCCARTHY [0030] teaches: “the symbolic loss function 442 and numerical loss function 444 are aggregated to calculate a joint loss (i.e., determining […] a training loss of the AI model as a measure of the neural loss and the symbolic loss).” MCCARTHY [0064] teaches: “the embedding vectors of the set of entities and the set of symbolic predicates in the embedding space, and other model parameters of the symbolic predictive pipeline and the numerical predictive pipeline are jointly trained by optimizing a joint loss function comprising a weighted sum of a symbolic loss function and a numerical loss function.”)
updating, by the electronic device, the weights of the plurality of layers of the AI model in proportion to the determined training loss. (MCCARTHY [0030] teaches: “Such a joint loss is minimized iteratively by an optimizer 450 based on, for example, stochastic gradient descent techniques (e.g., Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, AMSGrad, and the like). The iterations shown by arrow 451 are performed through all triples of the training triple set, and for each triple of the training set triples, the optimization process iterates for adjusting the training parameters of the MTKGPM 202 to minimize the joint loss (i.e., updating […] the weights of the plurality of layers of the AI model).” MCCARTHY [0065] teaches: “the embedding vectors of the set of entities and the set of symbolic predicates in the embedding space, and other model parameters of the symbolic predictive pipeline and the numerical predictive pipeline are jointly trained by optimizing a joint loss function based on a stochastic gradient descent (i.e., updating […] the weights […] in proportion to the determined training loss.” Examiner’s note: MCCARTHY’s optimizer, as shown in [0030] and FIG. 4, adjusts the training parameters by stochastic gradient descent. Under BRI, in proportion to the determined training loss can be interpreted as updating the weights according to a gradient calculated from the determined joint loss.)
Regarding Claim 6:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 1 as outlined above. ZAREIAN further teaches:
determining, by the electronic device, the plurality of contents and relationships between the plurality of contents; (ZAREIAN [page 3, section 1 Introduction] teaches: “A scene graph node represents an entity or predicate instance in a specific image, while a commonsense graph node represents an entity or predicate class, which is a general concept independent of the image.” ZAREIAN [page 7, section 4 Method] teaches: “Given an image, our model first applies a Faster R-CNN [36] to detect objects (i.e., determining […] the plurality of contents), and represents them as scene entity (SE) nodes. It also creates a scene predicate (SP) node for each pair of entities (i.e., determining […] relationships between the plurality of contents), which forms a scene graph proposal, yet to be classified. Given this graph and a background commonsense graph, each with fixed internal connectivity, our goal is to create bridge edges between the two graphs that connect each instance (SE and SP node) to its corresponding class (CE and CP node).
determining, by the electronic device, at least one scene graph based on the plurality of contents and the relationships between the plurality of contents, (ZAREIAN [page 7, section 4 Method] teaches: “Given an image, our model first applies a Faster R-CNN [36] to detect objects, and represents them as scene entity (SE) nodes. It also creates a scene predicate (SP) node for each pair of entities, which forms a scene graph proposal (i.e., determining […] at least one scene graph), yet to be classified.” Examiner’s note” ZAREIAN’s scene graph is created by using entity nodes (i.e., based on the plurality of contents) and the scene predicates for each node (i.e., and the relationship between the plurality of contents).)
wherein the at least one scene graph comprises a structural representation of the plurality of contents and the relationships between the plurality of contents in the input data. (ZAREIAN [page 2, Fig. 1.] shows a scene graph (i.e., at least one scene graph) representation from a given image (i.e., the input data) showing entities (i.e., plurality of contents) and the predicate for each node (i.e., and the relationships between the plurality of contents).)
Regarding Claim 7:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 6 as outlined above. ZAREIAN further teaches:
wherein the plurality of contents is determined using neural AI and the relationships between the plurality of contents is determined using symbolic AI, adhering to common-sense of a real world through an external symbolic knowledge graph. (ZAREIAN [page 7, section 4 Method] teaches: “Given an image, our model first applies a Faster R-CNN [36] (i.e., the plurality of contents is determined using neural AI) to detect objects, and represents them as scene entity (SE) nodes.” ZAREIAN [page 4, section 2.1 Scene graph generation ] teaches: “we do not classify each object and relation using classifiers, but instead use a pairwise matching mechanism to connect them to corresponding class nodes in the commonsense graph (i.e., and the relationships between the plurality of contents is determined using symbolic AI).” ZAREIAN [page 2, Fig. 1.] teaches a Commonsense Graph that shows real-world facts for guiding the detection of objects using the faster R-CNN model, and thus teaches determining contents and relationships adhering to common-sense of a real world through an external symbolic knowledge graph.)
Regarding Claim 8:
ZAREIAN teaches:
An electronic device for neuro-symbolic learning of an artificial intelligence (AI) model, the electronic device comprising: (ZAREIAN [page 14, section 6 Conclusion] teaches: “We proposed a new method for Scene Graph Generation that incorporates external commonsense knowledge in a novel, graphical neural framework (i.e., neuro-symbolic).” ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1. Perhaps the most important aspect of computation is the run time when deploying the model on new images. To this end, we run each trained model on the entire test set of Visual Genome (VG), i.e. 26446 images, and get the average run time over all images in terms of seconds. […] We record the time it takes to train each model (i.e., method for neuro-symbolic learning of an artificial intelligence (AI) model) on one epoch of the VG training set, i.e. 56224 images, and get the average over 10 training epochs.” ZAREIAN [page 2, Fig. 1.] teaches: “An example of a Visual Genome image and its ground truth scene graph.”)
a processor; (ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1.”)
a neuro-symbolic AI controller; (ZAREIAN [page 3, section 1 Introduction] teaches: “Our Graph Bridging Network (i.e., neuro-symbolic AI controller), GB-Net, successively infers edges and nodes, allowing to simultaneously exploit and refine the rich, heterogeneous structure of the interconnected scene and commonsense graphs.”)
memory storing one or more programs including computer-executable instructions that, when executed by the processor, cause the electronic device to: (ZAREIAN [pages 19-20, section B Computational cost] teaches: “GPU of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory […].” Examiner’s note: One of ordinary skill would recognize that, in order to execute ZAREIAN’s method for training a model using scene graphs and commonsense graphs, memory for storing programs that include computer-executable instructions for training the model would be required.)
receive input data comprising a plurality of contents for the neuro-symbolic learning of the AI model (ZAREIAN [page 7, section 4 Method] teaches: “Given an image (i.e., receive input data), our model first applies a Faster R-CNN [36] to detect objects (i.e., comprising a plurality of contents), and represents them as scene entity (SE) nodes.” ZAREIAN [page 14, section 6 Conclusion] teaches: “We proposed a new method for Scene Graph Generation that incorporates external commonsense knowledge in a novel, graphical neural framework (i.e., neuro-symbolic).” ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1. Perhaps the most important aspect of computation is the run time when deploying the model on new images. To this end, we run each trained model on the entire test set of Visual Genome (VG), i.e. 26446 images, and get the average run time over all images in terms of seconds. […] We record the time it takes to train each model (i.e., for the neuro-symbolic learning of the AI model) on one epoch of the VG training set, i.e. 56224 images (i.e., input data comprising a plurality of contents), and get the average over 10 training epochs.”)
determine a neural loss of the AI model by comparing the predicted probability for each of the plurality of contents with a predefined desired probability for each of the contents (ZAREIAN [page 10, section 4.3 Training] teaches: “We closely follow [2] which itself follows [52] for training procedure. Specifically, given the output and ground truth graphs, we align output entities and predicates to ground truth counterparts (i.e., a predefined desired probability for each of the plurality of contents). To align entities we use IoU and predicates will be aligned naturally since they correspond to aligned pairs of entities. Then we use the output probability scores of each node to define a cross-entropy loss (i.e., determine a neural loss of the AI model by comparing the predicted probability for each of the plurality of contents with a predefined desired probability for each of the plurality of contents). The sum of all node-level loss values will be the objective function to be minimized using Adam [18].”)
determine weights of a plurality of layers of the AI model (ZAREIAN [page 12, section 5.2 Implementation details] teaches: “We use three-layer fully connected networks with ReLU activation for all trainable networks
ϕ
i
n
i
t
,
ϕ
s
e
n
d
,
ϕ
r
e
c
e
i
v
e
and
ϕ
a
t
t
.” ZAREIAN [page 10, section 4.3 Training] teaches: “Then we use the output probability scores of each node to define a cross-entropy loss. The sum of all node-level loss values will be the objective function to be minimized using Adam [18].” Examiner’s note: Adam updates the trainable parameters (i.e., determine weights) for the three-layer fully connected networks (i.e., of a plurality of layers of the AI model) based on the computed loss value.)
update the weights of the plurality of layers of the AI model based on the neural loss […]. (ZAREIAN [page 10, section 4.3 Training] teaches: “Then we use the output probability scores of each node to define a cross-entropy loss. The sum of all node-level loss (i.e., based on the neural loss) values will be the objective function to be minimized (i.e., update the weights of the plurality of the layers of the AI model) using Adam [18].”)
ZAREIAN is not relied upon for teaching:
a communicator;
determine a predicted probability for each of the plurality of contents of the input data in an output of the AI model,
determine a symbolic loss for the AI model by comparing the predicted probability for each of the plurality of contents with a pre-determined undesired probability for each of the plurality of contents,
determine a training loss of the AI model as a measure of the neural loss and the symbolic loss,
update the weights of the plurality of layers of the AI model based on […] the symbolic loss.
However, KIM teaches: determine a predicted probability for each of the plurality of contents of the input data in an output of the AI model (KIM [page 103, section 3.1 Negative Learning] teaches: “In contrast, with NL, the CNNs are trained that “input image does not belong to this complementary label.” […] We consider the problem of c-class classification. Let
x
∈
X
be an input,
y
,
y
-
∈
Y
=
{
1
,
.
.
.
,
c
}
be its label and complementary label, respectively, and
y
,
y
-
∈
0
,
1
c
be their one-hot vector. Suppose the CNN
f
(
x
;
θ
)
maps the input space to the c-dimensional score space
f
:
X
→
R
c
, where
θ
is the set of network parameters. If
f
passes through the softmax function, the output can be interpreted as a probability
p
∈
∆
c
-
1
(i.e., the predicted probability), where
∆
c
-
1
denotes the c-dimensional simplex. When training with PL, the cross entropy loss function (i.e., determine a symbolic loss) of the network
f
(i.e., for the AI model) becomes:
L
f
,
y
=
-
∑
k
=
1
c
y
k
log
p
k
(
1
)
where
p
k
denotes the
k
t
h
element of
p
. Eq. 1 is suitable for optimizing the probability value corresponding to the given label as 1
(
p
y
→
1
)
, satisfying the purpose of PL. However, NL differs from PL as it optimizes the output probability corresponding to the complementary label to be far from 1 (to reach 0 in the end
(
p
y
-
→
0
)
). Therefore, we propose a loss function as follows:
L
f
,
y
-
=
-
∑
k
=
1
c
y
-
k
log
1
-
p
k
(
2
)
This complementary label is completely random in that it is selected randomly from the labels of all classes except for the given label y for every iteration during training (Algorithm 1). Eq. 2 enables the probability value of the complementary label to be optimized as zero, resulting in an increase in the probability values of other classes, meeting the purpose of NL.” Examiner’s note: KIM’s Equation (2) computes a loss
L
based on the predicted probability
p
y
-
assigned to the selected complementary class
y
-
k
. Under BRI, comparing the predicted probability […] with a pre-determined undesired probability can be interpreted as applying
-
log
1
-
p
y
-
, where the undesired probability is 0 and the predicted probability is
p
y
-
. KIM’s Negative Learning optimizes for
p
y
-
to reach 0 in the end.)
determine a symbolic loss for the AI model by comparing the predicted probability […] with a pre-determined undesired probability […] (KIM [page 103, section 3.1 Negative Learning] teaches: “In contrast, with NL, the CNNs are trained that “input image does not belong to this complementary label.” […] We consider the problem of c-class classification. Let
x
∈
X
be an input,
y
,
y
-
∈
Y
=
{
1
,
.
.
.
,
c
}
be its label and complementary label, respectively, and
y
,
y
-
∈
0
,
1
c
be their one-hot vector. Suppose the CNN
f
(
x
;
θ
)
maps the input space to the c-dimensional score space
f
:
X
→
R
c
, where
θ
is the set of network parameters. If
f
passes through the softmax function, the output can be interpreted as a probability
p
∈
∆
c
-
1
(i.e., the predicted probability), where
∆
c
-
1
denotes the c-dimensional simplex. When training with PL, the cross entropy loss function of the network
f
becomes:
L
f
,
y
=
-
∑
k
=
1
c
y
k
log
p
k
(
1
)
where
p
k
denotes the
k
t
h
element of
p
. Eq. 1 is suitable for optimizing the probability value corresponding to the given label as 1
(
p
y
→
1
)
, satisfying the purpose of PL. However, NL differs from PL as it optimizes the output probability corresponding to the complementary label to be far from 1 (to reach 0 in the end
(
p
y
-
→
0
)
). Therefore, we propose a loss function as follows:
L
f
,
y
-
=
-
∑
k
=
1
c
y
-
k
log
1
-
p
k
(
2
)
This complementary label is completely random in that it is selected randomly from the labels of all classes except for the given label y for every iteration during training (Algorithm 1). Eq. 2 enables the probability value of the complementary label to be optimized as zero, resulting in an increase in the probability values of other classes, meeting the purpose of NL.” Examiner’s note: KIM’s Equation (2) computes a loss
L
based on the predicted probability
p
y
-
assigned to the selected complementary class
y
-
k
. Under BRI, comparing the predicted probability […] with a pre-determined undesired probability can be interpreted as applying
-
log
1
-
p
y
-
, where the undesired probability is 0 and the predicted probability is
p
y
-
. KIM’s Negative Learning optimizes for
p
y
-
to reach 0 in the end.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of ZAREIAN and KIM before them, to include KIM’s negative learning techniques in ZAREIAN’s method for training a model using scene graphs and commonsense graphs. One of ordinary skill could apply KIM’s equation (2) to each of the nodes in ZAREIAN when generating scene graphs for training to prevent overfitting and achieve higher accuracy due to filtering based on KIM’s Selective Negative Learning and Positive Learning (SelNLPL), as disclosed in KIM [page 105, section 3.5 Semi-supervised learning]. One would have been motivated to make such a combination in order to perform Selective Negative Learning and Positive Learning (SelNLPL) to filter noisy data from training data and prevent overfitting, and train a model to determine images that do not belong to the complementary label (KIM [page 103, section 3. Method]).
ZAREIAN in view of KIM is not relied upon for teaching, but MCCARTHY teaches: a communicator; (MCCARTHY [0043] teaches: “[0043] The communication interfaces 802 may include wireless transmitters and receivers (herein, “transceivers”) 812 and any antennas 814 used by the transmit-and-receive circuitry of the transceivers 812. The transceivers 812 and antennas 814 may support WiFi network communications, for instance, under any version of IEEE 802.11, e.g., 802.11n or 802.11ac, or other wireless protocols such as Bluetooth, Wi-Fi, WLAN, cellular (4G, LTE/A). The communication interfaces 802 may also include serial interfaces, such as universal serial bus (USB), serial ATA, IEEE 1394, lighting port, I.sup.2C, slimBus, or other serial interfaces. The communication interfaces 802 may also include wireline transceivers 816 to support wired communication protocols. The wireline transceivers 816 may provide physical layer interfaces for any of a wide range of communication protocols, such as any type of Ethernet, optical networking protocols, data over cable service interface specification (DOCSIS), digital subscriber line (DSL), Synchronous Optical Network (SONET), or other protocol.”)
determine a training loss of the AI model as a measure of the neural loss and the symbolic loss (MCCARTHY [0030] teaches: “the symbolic loss function 442 and numerical loss function 444 are aggregated to calculate a joint loss (i.e., determine a training loss of the AI model as a measure of the neural loss and the symbolic loss).” MCCARTHY [0064] teaches: “the embedding vectors of the set of entities and the set of symbolic predicates in the embedding space, and other model parameters of the symbolic predictive pipeline and the numerical predictive pipeline are jointly trained by optimizing a joint loss function comprising a weighted sum of a symbolic loss function and a numerical loss function.”)
update the weights of the plurality of layers of the AI model based on the […] loss and the symbolic loss. (MCCARTHY [0029] teaches: “The symbolic loss function 442 quantifies symbolic prediction errors by comparing predictions of the symbolic task pipeline 204 to actual symbolic relationships indicated in training symbolic triples (in the embedding space). The numerical loss function 444 may be constructed as a regression loss that quantifies numerical prediction errors by comparing numerical predictions from the neural networks of the numerical pipeline 206 to actual numerals indicated in the training numerical triples (in value).” MCCARTHY [0030] teaches: “As further shown by 440 in FIG. 4, the symbolic loss function 442 (i.e., based on […] the symbolic loss) and numerical loss function 444 (i.e., based on the […] loss) are aggregated to calculate a joint loss. Such a joint loss is minimized iteratively by an optimizer 450 based on, for example, stochastic gradient descent techniques (e.g., Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, AMSGrad, and the like) (i.e., update the weights of the plurality of layers of the AI model). The iterations shown by arrow 451 are performed through all triples of the training triple set, and for each triple of the training set triples, the optimization process iterates for adjusting the training parameters of the MTKGPM 202 to minimize the joint loss.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of ZAREIAN, KIM, and MCCARTHY before them, to include MCCARTHY’s communication interface and its joint training using numerical and symbolic loss functions in ZAREIAN and KIM’s method for training a model using scene graphs and commonsense graphs. One would have been motivated to make such a combination in order to provide an integrated two-layer multi-task prediction model trained from an existing knowledge graph and to jointly train the model by minimizing an aggregation of a symbolic loss function and a numerical regression loss function, thus predicting new symbolic triples that may be previously unknown (or missing) so that they can be added to the knowledge graph to expand existing knowledge (MCCARTHY [0015] and [0023]).
Regarding Claim 9:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 8 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale.
Regarding Claim 13:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 8 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale.
Regarding Claim 14:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 13 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale.
Regarding Claim 15:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 8 as outlined above. ZAREIAN further teaches:
[…] neural loss […] (ZAREIAN [page 10, section 4.3 Training] teaches: “We closely follow [2] which itself follows [52] for training procedure. Specifically, given the output and ground truth graphs, we align output entities and predicates to ground truth counterparts. To align entities we use IoU and predicates will be aligned naturally since they correspond to aligned pairs of entities. Then we use the output probability scores of each node to define a cross-entropy loss (i.e., neural loss). The sum of all node-level loss values will be the objective function to be minimized using Adam [18].”)
wherein the various functions include […] division. (ZAREIAN [page 10, section 4.3 Training] teaches: “More specifically, we use the following loss function for each predicate node:
L
i
P
=
-
1
-
β
1
-
β
n
j
log
a
i
j
P
B
,
(
13
)
where
j
is the class index of the ground truth predicate aligned with
i
,
n
j
is the frequency of class
j
in training data, and
β
is a hyperparameter.”)
KIM further teaches: wherein the various functions include […] subtraction […] (KIM [page 103, section 3.1 Negative Learning] teaches: “Therefore, we propose a loss function as follows:
L
f
,
y
-
=
-
∑
k
=
1
c
y
-
k
log
1
-
p
k
(
2
)
This complementary label is completely random in that it is selected randomly from the labels of all classes except for the given label y for every iteration during training (Algorithm 1).”)
MCCARTHY further teaches: wherein the training loss is determined by applying various functions on both the […] loss and the symbolic loss, and wherein the various functions include addition, […], multiplication […] (MCCARTHY [0064] teaches: “In any of the implementations above, the embedding vectors of the set of entities and the set of symbolic predicates in the embedding space, and other model parameters of the symbolic predictive pipeline and the numerical predictive pipeline are jointly trained by optimizing a joint loss function comprising a weighted sum (i.e., applying various functions and wherein the various functions include addition, […], multiplication […]) of a symbolic loss function and a numerical loss function (i.e., on both the symbolic loss and the […] loss).”)
Regarding Claim 16:
The claim recites similar limitations as corresponding claims 1 and 8 and is rejected for similar reasons as claims 1 and 8 using similar teachings and rationale. ZAREIAN further teaches:
One or more non-transitory computer-readable storage media storing one or more computer programs including computer-executable instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform operations, the operations comprising: (ZAREIAN [page 14, section 6 Conclusion] teaches: “We proposed a new method for Scene Graph Generation that incorporates external commonsense knowledge in a novel, graphical neural framework.” ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1. Perhaps the most important aspect of computation is the run time when deploying the model on new images. To this end, we run each trained model on the entire test set of Visual Genome (VG), i.e. 26446 images, and get the average run time over all images in terms of seconds. […] We record the time it takes to train each model on one epoch of the VG training set, i.e. 56224 images, and get the average over 10 training epochs.” ZAREIAN [page 2, Fig. 1.] teaches: “An example of a Visual Genome image and its ground truth scene graph.” ZAREIAN [page 3, section 1 Introduction] teaches: “Our Graph Bridging Network, GB-Net, successively infers edges and nodes, allowing to simultaneously exploit and refine the rich, heterogeneous structure of the interconnected scene and commonsense graphs.” Examiner’s note: One of ordinary skill would recognize that, in order to execute ZAREIAN’s method for training a model using scene graphs and commonsense graphs, storage media storing programs with computer-executable instructions would be required.)
Regarding Claim 17:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 16 as outlined above. Additionally, the claim recites similar limitations as corresponding claims 2 and 9 and is rejected for similar reasons as claims 2 and 9 using similar teachings and rationale.
Regarding Claim 20:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 16 as outlined above. Additionally, the claim recites similar limitations as corresponding claims 6 and 13 and is rejected for similar reasons as claims 6 and 13 using similar teachings and rationale.
Claims 3-5, 10-12, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over ZAREIAN in view of KIM and MCCARTHY, as applied respectively above to claims 1, 8, and 16, and further in view of SAFAVI ("NEGATER: Unsupervised Discovery of Negatives in Commonsense Knowledge Bases") and PAI (US 20210174217 A1), hereafter SAFAVI and PAI respectively.
Regarding Claim 3:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 1 as outlined above. ZAREIAN further teaches:
selecting, by the electronic device, an external symbolic knowledge graph comprising a plurality of common-sense facts; (ZAREIAN [page 7, section 4 Method] teaches: “Given (i.e., selecting) this graph and a background commonsense graph (i.e., external symbolic knowledge graph comprising a plurality of common-sense facts), each with fixed internal connectivity, our goal is to create bridge edges between the two graphs that connect each instance (SE and SP node) to its corresponding class (CE and CP node).” ZAREIAN [page 12, section 5.2 Implementation details] teaches: “In our commonsense graph, the nodes are the 151 entity classes and 51 predicate classes that are fixed by [44], including background. We use the GloVE [33] embedding of category titles to initialize their node representation (via
ϕ
i
n
i
t
), and fix GloVE during training. We compile our commonsense edges from three sources, WordNet [30], ConceptNet [27], and Visual Genome. To summarize, there are three groups of edge types in our commonsense graph.” ZAREIAN [pages 19-20, section B Computational cost] teaches: “We compute the training and test speed of our method and compare to KERN [2] using identical hardware, with one GPU (i.e., by the electronic device) of type NVIDIA GeForce GTX 1080 Ti with 11 gigabytes of memory, and summarize the results in Table 1.)
ZAREIAN in view of KIM and MCCARTHY is not relied upon for teaching, but SAFAVI teaches: wherein the pre-determined undesired probability for each of the plurality of contents in the input data is determined by the electronic device (SAFAVI [page 5633, Abstract] teaches: “this paper proposes NegatER, a framework that ranks potential negatives (i.e., the pre-determined undesired probability for each of the plurality of contents in the input data is determined) in commonsense KBs using a contextual language model (LM).” SAFAVI [page 5644, Software and hardware] teaches: “We implement our LMs with the Transformers PyTorch library (Wolf et al., 2020) and run all experiments on a NVIDIA Tesla V100 GPU (i.e., by the electronic device) with 16 GB of RAM. Both BERT and RoBERTa take around 1.5 hours/epoch to train on the ConceptNet benchmark.” SAFAVI [page 5638, NegatER
-
θ
r
] teaches: “We rank candidates using fine-tuned BERT’s classification scores.”)
wherein the determining of the pre-determined undesired probability for each of the plurality of contents comprises: […] identifying, by the electronic device, a set of maximally violating facts from the plurality of violating common-sense facts; (SAFAVI [page 5635, section 3 Framework] teaches: “Then, a set of grammatical (R1) and topically consistent (R2) out-of-KB candidate statements are fed to the LM and ranked by the degree to which they “contradict” the LM’s finetuned positive beliefs (R3), such that the higher-ranking statements are more likely to be negative.” SAFAVI [page 5636, section 3.2 Ranking out-of-KB statements] teaches: “Finally, to meet requirement R3, we rank the remaining out-of-KB candidates by the degree to which they “contradict” the positive beliefs of the fine-tuned LM. These ranked statements can be then taken in order of rank descending as input to any discriminative KB reasoning task requiring negative examples […].” Examiner’s note: Under BRI, identifying […] a set of maximally violating facts from the plurality of violating common-sense facts can be interpreted as the contradicting ranked statements in descending order, the top statements being the most contradicting (i.e., maximally violating) to positive beliefs in the set.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of ZAREIAN, KIM, MCCARTHY, and SAFAVI before them, to include SAFAVI’s ranking of candidates that contradict positive beliefs in ZAREIAN, KIM, and MCCARTHY’s method for training a model using scene graphs and commonsense graphs. One would have been motivated to make such a combination so that discriminative models that operate over structured knowledge learn good decision boundaries (SAFAVI [page 5633, section 1 Introduction]).
ZAREIAN in view of KIM, MCCARTHY, and SAFAVI is not relied upon for teaching, but PAI teaches: constructing, by the electronic device, a negative knowledge graph comprising a plurality of violating common-sense facts from the external symbolic knowledge graph; (PAI [0005] teaches: “The processor is further adapted to execute the executable instructions stored in the memory to determine a positive structural score for each triple in the knowledge graph, adjust each positive structural score based on each corresponding significance parameter, generate a synthetic negative graph-based dataset based on the graph-based dataset (i.e., constructing, by the electronic device, a negative knowledge graph), the synthetic negative graph-based dataset comprising a set of synthetic negative triples (i.e., comprising a plurality of violating common-sense facts from the external symbolic knowledge graph), and determine a negative structural score for each synthetic negative triple of the synthetic negative graph-based dataset.”)
generating, by the electronic device, symbolic labels in a form of probabilities for the maximally violating facts, and (PAI [0005] teaches: “determine (i.e., generating) a negative structural score (i.e., symbolic labels in a form of probabilities) for each synthetic negative triple (i.e., for the maximally violating facts) of the synthetic negative graph-based dataset.”)
wherein the maximally violating facts comprise the undesired probability. (PAI [0074] teaches: “In some implementations, a loss function may be minimized in order to learn optimal parameters that best discriminate positive statements from negative statements (i.e., undesired probability). Exemplary optimizers may also be utilized, including the following optimization algorithms: stochastic gradient descent (SGD), adaptive moment estimation (Adam), and the gradient-based optimization algorithm (Adagrad).” PAI [0080] teaches: “The generated negative structure score may be utilized as input to a non-linear function 855, for example, a sigmoid function to generate a normalized negative structure score. The normalized negative structure score may be utilized as input to a multiplication function 857.” Examiner’s note: The generated negative structure score for each negative triplet is input into multiplication function 857 to then determine a KGE loss in step 860, according to PAI [FIG. 8]. Further, PAI’s optimization algorithms learn optimal parameters to discriminate positive statements from negative statements. Under BRI, the undesired probability can be interpreted as the negative structure scores from each negative triplet for which the model is being trained to discriminate.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of ZAREIAN. KIM, MCCARTHY, SAFAVI, and PAI before them, to include PAI’s generation of a negative graph-based dataset and generation of negative structural scores for each negative triple in the negative graph-based dataset in ZAREIAN, KIM, MCCARTHY, and SAFAVI’s method for training a model using scene graphs and commonsense graphs. SAFAVI teaches ranking candidates in a knowledge graph that contradict positive beliefs. PAI teaches generating negative graph-based datasets and scoring the structure of each negative triplet. Therefore, one of ordinary skill in the art could use PAI’s method to generate the dataset using SAFAVI’s ranked negative triplets and then score each triplet to provide a structural score for minimizing the loss of the model in discriminative tasks that operate over knowledge-structured datasets/graphs. One would have been motivated to make such a combination in order to train machine learning models on both true and false statements/facts and predict, with high accuracy, missing links of a specific type in a knowledge graph with numerical values associated to the known links (PAI [0027] and [0067]).
Regarding Claim 4:
ZAREIAN in view of KIM, MCCARTHY, SAFAVI, and PAI teaches the elements of claim 3 as outlined above. ZAREIAN further teaches:
wherein the plurality of common-sense facts comprise at least one of a set of pre-defined rules or common-sense facts associated with a real-world. (ZAREIAN [page 2, Fig. 1.] teaches a Commonsense Graph with real-world facts.)
Regarding Claim 5:
ZAREIAN in view of KIM, MCCARTHY, SAFAVI, and PAI teaches the elements of claim 3 as outlined above. ZAREIAN further teaches:
[…] with respect to the input data. (ZAREIAN [page 7, section 4 Method] teaches: “Given an image (i.e., with respect to the input data), our model first applies a Faster R-CNN [36] to detect objects, and represents them as scene entity (SE) nodes.” ZAREIAN [page 14, section 6 Conclusion] teaches: “We proposed a new method for Scene Graph Generation that incorporates external commonsense knowledge in a novel, graphical neural framework.” ZAREIAN [page 10, section 4.3 Training] teaches: “Then we use the output probability scores of each node to define a cross-entropy loss.”)
SAFAVI further teaches: wherein the maximally violating facts further comprise facts which are against the plurality of common-sense facts […] . (SAFAVI [page 5635, section 3 Framework] teaches: “Then, a set of grammatical (R1) and topically consistent (R2) out-of-KB candidate statements are fed to the LM and ranked by the degree to which they “contradict” the LM’s finetuned positive beliefs (R3), such that the higher-ranking statements are more likely to be negative.” SAFAVI [page 5636, section 3.2 Ranking out-of-KB statements] teaches: “Finally, to meet requirement R3, we rank the remaining out-of-KB candidates by the degree to which they “contradict” the positive beliefs of the fine-tuned LM. These ranked statements can be then taken in order of rank descending as input to any discriminative KB reasoning task requiring negative examples […].” Examiner’s note: Under BRI, maximally violating facts further comprise facts which are against the plurality of common-sense facts can be interpreted as the contradicting ranked statements in descending order, the top statements being the most contradicting (i.e., maximally violating) to positive beliefs in the set.)
Regarding Claim 10:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 8 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale.
Regarding Claim 11:
ZAREIAN in view of KIM, MCCARTHY, SAFAVI, and PAI teaches the elements of claim 10 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale.
Regarding Claim 12:
ZAREIAN in view of KIM, MCCARTHY, SAFAVI, and PAI teaches the elements of claim 10 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale.
Regarding Claim 18:
ZAREIAN in view of KIM and MCCARTHY teaches the elements of claim 16 as outlined above. Additionally, the claim recites similar limitations as corresponding claims 3 and 10 and is rejected for similar reasons as claims 3 and 10 using similar teachings and rationale.
Regarding Claim 19:
ZAREIAN in view of KIM, MCCARTHY, SAFAVI, and PAI teaches the elements of claim 18 as outlined above. Additionally, the claim recites similar limitations as corresponding claims 4 and 11 and is rejected for similar reasons as claims 4 and 11 using similar teachings and rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
JAIN (US 20220383143 A1) discloses generating negative samples for knowledge graph embedding models by ranking triples that violate ontology constraints.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Alvaro S Laham Bauzo whose telephone number is (571)272-5650. The examiner can normally be reached Mon-Fri 7:30 AM - 11:00 AM | 1:00 PM - 5:30 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.S.L./Examiner, Art Unit 2146
/USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146