DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This Office Action is in response to the communication filed on 16 January 2022
Claims 1-20 are being considered on the merits.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Teney, et. al. (“Graph-structured representations for Visual Question Answering” arXiv:1609.05600v2 [cs.CV] 30 Mar 2017; hereinafter, “Teney”) in view of Zhao, et. al. (US 2021/0248376 A1; hereinafter, “Zhao”)
Claim 1:
A method of automated data extraction, comprising: deploying a single visual question answering (VQA) model (Teney, Abstract and fig 2: “This paper proposes to improve visual question answering (VQA) with structured representations of both scene contents and questions” “Figure 2. Architecture of the proposed neural network. The input is provided as a description of the scene (a list of objects with their visual characteristics) and a parsed question (words with their syntactic relations)”) comprising a single neural network trained through a supervised learning process comprising: (Teney, sec. 1 and 4: “Note that this pretraining and ad hoc processing of the language part mimics a practice common for the image part, in which visual features are usually obtained from a fixed CNN, itself pretrained on a larger dataset and with a different (supervised classification) objective.” “We now describe a deep neural network suitable for pro cessing the question and scene graphs to infer an answer.”)
receiving outputs from the single VQA model based on processing, through nodes of the single VQA model, a training data set (Teney, sec. 5.1: “We trained our model with limited subsets of the training data (see Fig. 4).”) that was generated by injecting, using a processor (Zhao, para. 0144: “Computer-executable instructions comprise, for example, instructions and data which, when executed by a processor, cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions.”), respective text of a respective set of possible answers into given text (Teney, sec. 1: “This produces a graph representation of the question in which each node represents a word and each edge a particular type of dependency (e.g. determiner, nominal subject, direct object, etc.). Second, we associate each word (node) with a vector embedding pretrained on large corpora of text data [20]. This embedding maps the words to a space in which distances are semantically meaningful. Consequently, this essentially regularizes the remainder of the network to share learned concepts among related words and synonyms. This particularly helps in dealing with rare words, and also allows questions to include words absent from the training questions/answers.”) and given positional data that was extracted from a given image (Teney, sec. 1: “Each object in the scene corresponds to a node in the scene graph, which has an associated feature vector describing its appearance. The graph is fully connected, with each edge representing the relative position of the objects in the image.”) ; and (Teney, sec. 1: “The two graph representations feed into a deep neural network that we will describe in Section 4. The advantage of this approach with text- and scene-graphs, rather than more typical representations, is that the graphs can capture relationships between words and between objects which are of semantic significance. This enables the GNN to exploit (1) the unordered nature of scene elements (the objects in particular) and (2) the semantic relationships between elements (and the grammatical relationships between words in particular).”)
iteratively adjusting parameters of the single VQA model based on a difference between the outputs and training outputs, wherein the training outputs are based on known correct answers; (Zhao, para. 0101: “In particular, the loss function 632 can return such loss data to the query-response system 106 based upon which the query-response system 106 can adjust various parameters/hyperparameters to improve the quality/accuracy of predicted-visual-feature probabilities in subsequent training iterations—by narrowing the difference between the predicted-graphical objects 628 and the synthetic-ground-truth-graphical objects 630.”)
providing one or more inputs to the single VQA model based on an image and a set of possible answers associated with a question, wherein: (Teney, sec. 1: “The two graph representations feed into a deep neural network that we will describe in Section 4. The advantage of this approach with text- and scene-graphs, rather than more typical representations, is that the graphs can capture relationships between words and between objects which are of semantic significance” Examiner notes Teney teaches providing a training question/answer set i.e. a set of possible answers associated with a question and a scene representation into a deep neural network used for VQA)
the question is related to the image; (Teney, fig. 1: Examiner notes Figure 1 teaches a question related to an image).
a data set to be processed by the VQA model is generated by injecting, using one or more processors (Zhao, para. 0144: “Computer-executable instructions comprise, for example, instructions and data which, when executed by a processor, cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions.”), text data comprising the set of possible answers into text and positional data extracted from the image; and (Teney, sec. 1: “ This produces a graph representation of the question in which each node represents a word and each edge a particular type of dependency (e.g. determiner, nominal subject, direct object, etc.). Second, we associate each word (node) with a vector embedding pretrained on large corpora of text data [20]. This embedding maps the words to a space in which distances are semantically meaningful. Consequently, this essentially regularizes the remainder of the network to share learned concepts among related words and synonyms. This particularly helps in dealing with rare words, and also allows questions to include words absent from the training questions/answers.” Examiner notes for examination purposes “injecting…text data comprising the set of possible answers into text” is interpreted to mean labeling)
the data set is processed through nodes of the VQA model until the VQA model generates an output indicating an answer of the set of possible answers; and (Teney sec. 1 and fig. 1: “Our network uses multiple layers that iterate over the features associated with every node, then ultimately identifies a soft matching between nodes from the two graphs. This matching reflects the correspondences between the words in the question and the objects in the image. The features of the matched nodes then feed into a classifier to infer the answer to the question” Examiner notes figure 1 illustrates a VQA neural network outputting an answer among a possible set of answers)
performing one or more actions within a software application based on the answer indicated in the output. (Zhao, para. 0035: “Upon selecting a response, the query-response system can provide an audio reproduction of the identified response or provide, for display within a user interface of a client device, the identified response to the question.” Examiner notes Zhao teaches audio reproduction or display of output within the query-response system)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney. Teney teaches an improved visual question answering (VQA) with structured representations of both scene contents and questions; Zhao teaches generating a response to a question received from a user during display or playback of a video segment by utilizing a query-response-neural network. One of ordinary skill would have been motivated to combine the teachings of Zhao into Teney in order to enable users to interactively query a video for(Zhao paras. 0001 and 0084.
Claim 2:
The method of Claim 1, wherein the text of the set of possible answers and the text data extracted from the image are analyzed by the single VQA model in a single forward pass. (Teney, sec. 2: “The image is passed through a convolutional neural network (CNN) pretrained for image classification, from which intermediate features are extracted to describe the image. The question is typically passed through a recurrent neural network (RNN) such as an LSTM, which produces a fixed-size vector representing the sequence of words. These two representations are mapped to a joint space by one or several non-linear layers. They can then be fed into a classifier over an output vocabulary, predicting the final answer.” Examiner notes Teney teaches use of an RNN and LSTM, passing data through the neural network using at least a single forward pass.)
Claim 3:
The method of Claim 2, wherein the single VQA model further analyzes the positional data extracted from the image during the single forward pass. (Teney, sec. 2 and fig 1: “The image is passed through a convolutional neural network (CNN) pretrained for image classification, from which intermediate features are extracted to describe the image. The question is typically passed through a recurrent neural network (RNN) such as an LSTM, which produces a fixed-size vector representing the sequence of words” “The graph is fully connected, with each edge representing the relative position of the objects in the image.”)
Claim 4:
The method of Claim 1, wherein the set of possible answers is not present in text form in the image. (Teney, fig 1: Examiner notes fig 1 illustrates a set of possible text answers where none of the text is present in image)
Claim 5:
The method of Claim 4, further comprising: determining a different question related to the image, wherein a given answer of a different set of possible answers associated with the different question is present in text form in the image; (Zhao, para. 0038 and 0105: “For example, unlike other systems, the query-response system can accurately respond to a question about content visibly depicted in a video segment, even though the narration during the video segment corresponds to some other content or another topic. Similarly, unlike conventional video systems, the query-response system can accurately respond to a question having a subject or predicate that depends on the visual context (e.g., “Is there a shortcut for that?”). In this manner, the query-response system can provide increased accuracy to responses to questions received during playback of a video segment.” “To extract the textual-feature embeddings 644 from the object 614, for example, the graphical-object-matching engine 606 can use one or more word-vector-representation models, as described above in relation to the foregoing figures. Additionally or alternatively, the graphical-object-matching engine 606 can perform an optical character recognition process to convert image data of the detected pop-up dialogue 612 (or just the object 614) to textual data” Examiner notes Zhao teaches a determining a question about the content visibly depicted in an image and converting image object data into textual data i.e. the text form of the object present in image)
providing one or more different inputs to the single VQA model based on the image; and (Teney, sec. 2: “As in our model, the DMN models interactions between different parts of the input. Our method can additionally take, as input, features characterizing arbitrary relations between parts of the input (the edge features in our graphs).” Examiner notes Teney teaches taking features characterizing arbitrary relations between parts of the input where each feature is different than other features and also teaches providing to the DMN an input and features of the input which are different inputs from the input).
receiving, from the single VQA model in response to the one or more different inputs, one or more different outputs indicating the given answer of the different set of possible answers. (Zhao, para. 0038 and 0105: “For example, unlike other systems, the query-response system can accurately respond to a question about content visibly depicted in a video segment, even though the narration during the video segment corresponds to some other content or another topic. Similarly, unlike conventional video systems, the query-response system can accurately respond to a question having a subject or predicate that depends on the visual context (e.g., “Is there a shortcut for that?”). In this manner, the query-response system can provide increased accuracy to responses to questions received during playback of a video segment.” “ To extract the textual-feature embeddings 644 from the object 614, for example, the graphical-object-matching engine 606 can use one or more word-vector-representation models, as described above in relation to the foregoing figures. Additionally or alternatively, the graphical-object-matching engine 606 can perform an optical character recognition process to convert image data of the detected pop-up dialogue 612 (or just the object 614) to textual data” Examiner notes for examination purposes only, “receiving” is an action by a user. Zhao teaches a determining a question about the content visibly depicted in an image and converting image object data into textual data i.e. the text form of the object present in image where both the inputs and outputs are different than those taught above by Teney).
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 6:
The method of Claim 1, wherein performing the one or more actions within the software application based on the answer indicated in the output comprises assigning a value to a variable in the software application based on the answer. (Zhao, para. 0035: “In some embodiments, the query-response system generates a matching score for each query-response pairing between a query-context vector with a respective candidate-response. In turn, the query-response system can select a response to the question based on a particular query-response pairing having a matching score that satisfies a threshold matching score.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 7:
The method of Claim 1, wherein performing the one or more actions within the software application based on the answer indicated in the output comprises displaying the answer to a user via a user interface. (Zhao, para. 0035: “In some embodiments, the query-response system generates a matching score for each query-response pairing between a query-context vector with a respective candidate-response. In turn, the query-response system can select a response to the question based on a particular query-response pairing having a matching score that satisfies a threshold matching score. As described below, the query-response system may generate matching scores represented as matching probabilities (e.g., as output from fully-connected layers and/or a softmax function), cosine similarity values, or Euclidean distances. Upon selecting a response, the query-response system can provide an audio reproduction of the identified response or provide, for display within a user interface of a client device, the identified response to the question.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 8:
The method of Claim 1, wherein the set of possible answers comprises true and false. (Teney, sec. 1: “A particularly attractive VQA dataset was introduced in [31] byselecting only the questions with binary answers (e.g. yes/no) and pairing each (synthetic) image with a minimally-different complementary version that elicits the opposite (no/yes) answer (see examples in Fig. 5, bottom rows).” Examiner notes Teney teaches binary answers where yes correlates to true and no correlates to false)
Claim 9:
A method for training a model, comprising: providing training inputs to a single visual question answering (VQA) model, wherein: the training inputs are based on a plurality of images and a set of possible answers related to the plurality of images; (Teney, sec. 1: “The two graph representations feed into a deep neural network that we will describe in Section 4. The advantage of this approach with text- and scene-graphs, rather than more typical representations, is that the graphs can capture relationships between words and between objects which are of semantic significance” Examiner notes Teney teaches providing a training question/answer set i.e. a set of possible answers associated with a question and a scene representation into a deep neural network used for VQA)
a respective data set to be processed by the VQA model for each respective image of the plurality of images is generated by injecting, using one or more processors, (Zhao, para. 0144: “Computer-executable instructions comprise, for example, instructions and data which, when executed by a processor, cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions.”), text data comprising the set of possible answers into text and positional data extracted from the respective image; and Teney, sec. 1: “This produces a graph representation of the question in which each node represents a word and each edge a particular type of dependency (e.g. determiner, nominal subject, direct object, etc.). Second, we associate each word (node) with a vector embedding pretrained on large corpora of text data [20]. This embedding maps the words to a space in which distances are semantically meaningful. Consequently, this essentially regularizes the remainder of the network to share learned concepts among related words and synonyms. This particularly helps in dealing with rare words, and also allows questions to include words absent from the training questions/answers.” Examiner notes for examination purposes “injecting…text data comprising the set of possible answers into text” is interpreted to mean labeling)
each respective image of the plurality of images is processed through nodes of the single VQA model until the single VQA model generates an output indicating an answer of the set of possible answers; and (Teney sec. 1 and fig. 1: “Our network uses multiple layers that iterate over the features associated with every node, then ultimately identifies a soft matching between nodes from the two graphs. This matching reflects the correspondences between the words in the question and the objects in the image. The features of the matched nodes then feed into a classifier to infer the answer to the question” Examiner notes figure 1 illustrates a VQA neural network outputting an answer among a possible set of answers)
iteratively adjusting one or more parameters of the single VQA model based on a difference between the output and a training output, wherein the training output is based on a known answer to a question associated with a respective image. (Zhao, para. 0101: “ In particular, the loss function 632 can return such loss data to the query-response system 106 based upon which the query-response system 106 can adjust various parameters/hyperparameters to improve the quality/accuracy of predicted-visual-feature probabilities in subsequent training iterations—by narrowing the difference between the predicted-graphical objects 628 and the synthetic-ground-truth-graphical objects 630.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 10:
The method of Claim 9, wherein the text of the set of possible answers and the text data extracted from the respective image are analyzed by the single VQA model in a single forward pass. (Teney, sec. 2: “The image is passed through a convolutional neural network (CNN) pretrained for image classification, from which intermediate features are extracted to describe the image. The question is typically passed through a recurrent neural network (RNN) such as an LSTM, which produces a fixed-size vector representing the sequence of words. These two representations are mapped to a joint space by one or several non-linear layers. They can then be fed into a classifier over an output vocabulary, predicting the final answer.” Examiner notes Teney teaches use of an RNN and LSTM, passing data through the neural network using at least a single forward pass.)
Claim 11:
The method of Claim 10, wherein the single VQA model further analyzes the positional data extracted from the respective image during the single forward pass. (Teney, sec. 2 and fig 1: “The image is passed through a convolutional neural network (CNN) pretrained for image classification, from which intermediate features are extracted to describe the image. The question is typically passed through a recurrent neural network (RNN) such as an LSTM, which produces a fixed-size vector representing the sequence of words” “The graph is fully connected, with each edge representing the relative position of the objects in the image.”)
Claim 12:
The method of Claim 9, wherein the set of possible answers is not present in text form in the plurality of images. (Teney, sec. 1 and fig 1: “Multiple datasets for VQA have been introduced with either real [4, 15, 18, 22, 32] or synthetic images [4, 31]” Examiner notes fig 1 illustrates a set of possible text answers where none of the text is present in image)
Claim 13:
A system, comprising: one or more processors; and a memory comprising instructions that, (Zhao, para. 0128: “Each of the components of the computing device 1002 can include software, hardware, or both. For example, the components of the computing device 1002 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices, such as a client device or server device. When executed by the one or more processors, the computer-executable instructions of the query-response system 106 can cause the computing device(s) (e.g., the computing device 1002, the server(s) 102) to perform the methods described herein.”) when executed by the one or more processors, cause the system to:
deploy a single visual question answering (VQA) model (Teney, abstract and fig 2: “This paper proposes to improve visual question answering (VQA) with structured representations of both scene contents and questions” “Figure 2. Architecture of the proposed neural network. The input is provided as a description of the scene (a list of objects with their visual characteristics) and a parsed question (words with their syntactic relations)”) comprising a single neural network trained through a supervised learning process comprising: (Teney, sec. 1 and 4: “Note that this pretraining and ad hoc processing of the language part mimics a practice common for the image part, in which visual features are usually obtained from a fixed CNN, itself pretrained on a larger dataset and with a different (supervised classification) objective.” “We now describe a deep neural network suitable for processing the question and scene graphs to infer an answer.”)
receiving outputs from the single VQA model based on processing, through nodes of the single VQA model, a training data set (Teney, sec. 5.1: “We trained our model with limited subsets of the training data (see Fig. 4).”) that was generated by injecting, using a processor of the one or more processors (Zhao, para. 0144: “Computer-executable instructions comprise, for example, instructions and data which, when executed by a processor, cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions.”), respective text of a respective set of possible answers into given text (Teney, sec. 1: “ This produces a graph representation of the question in which each node represents a word and each edge a particular type of dependency (e.g. determiner, nominal subject, direct object, etc.). Second, we associate each word (node) with a vector embedding pretrained on large corpora of text data [20]. This embedding maps the words to a space in which distances are semantically meaningful. Consequently, this essentially regularizes the remainder of the network to share learned concepts among related words and synonyms. This particularly helps in dealing with rare words, and also allows questions to include words absent from the training questions/answers.”) and given positional data that was extracted from a given image (Teney, sec. 1: “Each object in the scene corresponds to a node in the scene graph, which has an associated feature vector describing its appearance. The graph is fully connected, with each edge representing the relative position of the objects in the image.”); and
iteratively adjusting parameters of the single VQA model based on a difference between the outputs and training outputs, wherein the training outputs are based on known correct answers; (Zhao, para. 0101: “ In particular, the loss function 632 can return such loss data to the query-response system 106 based upon which the query-response system 106 can adjust various parameters/hyperparameters to improve the quality/accuracy of predicted-visual-feature probabilities in subsequent training iterations—by narrowing the difference between the predicted-graphical objects 628 and the synthetic-ground-truth-graphical objects 630.”)
provide one or more inputs to the single VQA model based on an image and a set of possible answers associated with a question, wherein: (Teney, sec. 1: “The two graph representations feed into a deep neural network that we will describe in Section 4. The advantage of this approach with text- and scene-graphs, rather than more typical representations, is that the graphs can capture relationships between words and between objects which are of semantic significance” Examiner notes Teney teaches providing a training question/answer set i.e. a set of possible answers associated with a question and a scene representation into a deep neural network used for VQA)
the question is related to the image; (Teney, fig. 1: Examiner notes Figure 1 teaches a question related to an image).
a data set to be processed by the VQA model is generated by injecting, using one or more of the one or more processors (Zhao, para. 0144: “Computer-executable instructions comprise, for example, instructions and data which, when executed by a processor, cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions.”), text data comprising the set of possible answers into text and positional data extracted from the image; and (Teney, sec. 1: “ This produces a graph representation of the question in which each node represents a word and each edge a particular type of dependency (e.g. determiner, nominal subject, direct object, etc.). Second, we associate each word (node) with a vector embedding pretrained on large corpora of text data [20]. This embedding maps the words to a space in which distances are semantically meaningful. Consequently, this essentially regularizes the remainder of the network to share learned concepts among related words and synonyms. This particularly helps in dealing with rare words, and also allows questions to include words absent from the training questions/answers.” Examiner notes for examination purposes “injecting…text data comprising the set of possible answers into text” is interpreted to mean labeling)
the data set is processed through nodes of the VQA model until the VQA model generates an output indicating an answer of the set of possible answers; and (Teney sec. 1 and fig. 1: “Our network uses multiple layers that iterate over the features associated with every node, then ultimately identifies a soft matching between nodes from the two graphs. This matching reflects the correspondences between the words in the question and the objects in the image. The features of the matched nodes then feed into a classifier to infer the answer to the question” Examiner notes figure 1 illustrates a VQA neural network outputting an answer among a possible set of answers)
perform one or more actions within a software application based on the answer indicated in the output. (Zhao, para. 0035: “Upon selecting a response, the query-response system can provide an audio reproduction of the identified response or provide, for display within a user interface of a client device, the identified response to the question.” Examiner notes Zhao teaches audio reproduction or display of output within the query-response system)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 14:
The system of Claim 13, wherein the text of the set of possible answers and the text data extracted from the image are analyzed by the single VQA model in a single forward pass. (Teney, sec. 2: “The image is passed through a convolutional neural network (CNN) pretrained for image classification, from which intermediate features are extracted to describe the image. The question is typically passed through a recurrent neural network (RNN) such as an LSTM, which produces a fixed-size vector representing the sequence of words. These two representations are mapped to a joint space by one or several non-linear layers. They can then be fed into a classifier over an output vocabulary, predicting the final answer.” Examiner notes Teney teaches use of an RNN and LSTM, passing data through the neural network using at least a single forward pass.)
Claim 15:
The system of Claim 14, wherein the single VQA model further analyzes the positional data extracted from the image during the single forward pass. (Teney, sec. 1 and fig 1: “The image is passed through a convolutional neural network (CNN) pretrained for image classification, from which intermediate features are extracted to describe the image. The question is typically passed through a recurrent neural network (RNN) such as an LSTM, which produces a fixed-size vector representing the sequence of words” “The graph is fully connected, with each edge representing the relative position of the objects in the image.”)
Claim 16:
The system of Claim 13, wherein the set of possible answers is not present in text form in the image. (Teney, fig 1: Examiner notes fig 1 illustrates a set of possible text answers where none of the text is present in image)
Claim 17:
The system of Claim 16, wherein the instructions, when executed by the one or more processors, further cause the system to: determine a different question related to the image, wherein a given answer of a different set of possible answers associated with the different question is present in text form in the image; (Zhao, para. 0038 and 0105: “For example, unlike other systems, the query-response system can accurately respond to a question about content visibly depicted in a video segment, even though the narration during the video segment corresponds to some other content or another topic. Similarly, unlike conventional video systems, the query-response system can accurately respond to a question having a subject or predicate that depends on the visual context (e.g., “Is there a shortcut for that?”). In this manner, the query-response system can provide increased accuracy to responses to questions received during playback of a video segment.” “To extract the textual-feature embeddings 644 from the object 614, for example, the graphical-object-matching engine 606 can use one or more word-vector-representation models, as described above in relation to the foregoing figures. Additionally or alternatively, the graphical-object-matching engine 606 can perform an optical character recognition process to convert image data of the detected pop-up dialogue 612 (or just the object 614) to textual data” Examiner notes Zhao teaches a determining a question about the content visibly depicted in an image and converting image object data into textual data i.e. the text form of the object present in image)
provide one or more different inputs to the single VQA model based on the image; and (Teney, sec. 2: “As in our model, the DMN models interactions between different parts of the input. Our method can additionally take, as input, features characterizing arbitrary relations between parts of the input (the edge features in our graphs).” Examiner notes Teney teaches taking features characterizing arbitrary relations between parts of the input where each feature is different than other features and also teaches providing to the DMN an input and features of the input which are different inputs from the input).
receive, from the single VQA model in response to the one or more different inputs, one or more different outputs indicating the given answer of the different set of possible answers. . (Zhao, para. 0038 and 0105: “For example, unlike other systems, the query-response system can accurately respond to a question about content visibly depicted in a video segment, even though the narration during the video segment corresponds to some other content or another topic. Similarly, unlike conventional video systems, the query-response system can accurately respond to a question having a subject or predicate that depends on the visual context (e.g., “Is there a shortcut for that?”). In this manner, the query-response system can provide increased accuracy to responses to questions received during playback of a video segment.” “ To extract the textual-feature embeddings 644 from the object 614, for example, the graphical-object-matching engine 606 can use one or more word-vector-representation models, as described above in relation to the foregoing figures. Additionally or alternatively, the graphical-object-matching engine 606 can perform an optical character recognition process to convert image data of the detected pop-up dialogue 612 (or just the object 614) to textual data” Examiner notes for examination purposes only, “receiving” is an action by a user. Zhao teaches a determining a question about the content visibly depicted in an image and converting image object data into textual data i.e. the text form of the object present in image where both the inputs and outputs are different than those taught above by Teney).
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 18:
The system of Claim 13, wherein performing the one or more actions within the software application based on the answer indicated in the output comprises assigning a value to a variable in the software application based on the answer. (Zhao, para. 0035: “In some embodiments, the query-response system generates a matching score for each query-response pairing between a query-context vector with a respective candidate-response. In turn, the query-response system can select a response to the question based on a particular query-response pairing having a matching score that satisfies a threshold matching score.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 19:
The system of Claim 13, wherein performing the one or more actions within the software application based on the answer indicated in the output comprises displaying the answer to a user via a user interface. (Zhao, para. 0035: “In some embodiments, the query-response system generates a matching score for each query-response pairing between a query-context vector with a respective candidate-response. In turn, the query-response system can select a response to the question based on a particular query-response pairing having a matching score that satisfies a threshold matching score. As described below, the query-response system may generate matching scores represented as matching probabilities (e.g., as output from fully-connected layers and/or a softmax function), cosine similarity values, or Euclidean distances. Upon selecting a response, the query-response system can provide an audio reproduction of the identified response or provide, for display within a user interface of a client device, the identified response to the question.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Zhao into Teney, as modified, as set forth above with respect to claim 1.
Claim 20:
The system of Claim 13, wherein the set of possible answers comprises true and false. (Teney, sec. 1: “A particularly attractive VQA dataset was introduced in [31] by selecting only the questions with binary answers (e.g. yes/no) and pairing each (synthetic) image with a minimally-different complementary version that elicits the opposite (no/yes) answer (see examples in Fig. 5, bottom rows).” Examiner notes Teney teaches binary answers where yes correlates to true and no correlates to false)
Search Notes
PE2E Search best results CPC + key terms: VQA
IP.com best search results based on application as main search idea with priority date filter
Google scholar search best results from: visual question answering "supervised learning" positional data image output VQA
Prior Art
M. Stefanini, M. Cornia, L. Baraldi and R. Cucchiara, ("A Novel Attention-based Aggregation Function to Combine Vision and Language," 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 2021, pp. 1212-1219, doi: 10.1109/ICPR48806.2021.9413269.) teaches an approach to VQA comprising a set of scores for each element of each modality employing a novel variant of cross-attention, and performs a learnable and cross-modal reduction, which can be used for both classification and ranking
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sally T. Ley whose telephone number is (571)272-3406. The examiner can normally be reached Monday - Thursday, 10:00am - 6:00pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/STL/Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147