Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner’s Note
The Examiner encourages Applicant to schedule an interview (via Automated Interview Request (AIR)) to discuss issues related to, for example, the rejections noted below under 35 U.S.C § 101 and § 103, for moving toward allowance. (e.g., integrating dependent claims, elaborating claim languages with specificities, and elaborating claims toward a practical application)
Providing supporting paragraph(s) for each limitation of amended/new claim(s) in Remarks is strongly requested for clear and definite claim interpretations by Examiner (e.g., to avoid rejections under 35 U.S.C § 112(a) “Lack of written description”)
Applicant can schedule interviews (via Automated Interview Request (AIR)) at any stage of the prosecution (e.g., Non-Final, Final, and After-Final) to discuss any issues related to, for example, rejections under 35 U.S.C § 101 and § 102/103, for moving toward allowance.
If a limitation has bold brackets (i.e. [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards.
If a limitation has one or more bold underlines, the one or more bold underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references.
Priority
Acknowledgment is made of applicant's claim for the present application filed on 09/05/2023.
Response to Arguments
Applicant's arguments filed on 06/22/2026 have been fully considered but they are not persuasive.
In Remarks, regarding 35 USC § 101, Applicant contends:
The Applicant respectfully submits that claims 1-18 are directed to statutory subject matter. For example, it is not possible to perform in the mind the element of "executing, at the computer system, the at least one machine learning model relevant to the problem using the selected computing infrastructure to generate at least one recommendation relevant to the problem." In addition, the element of "wherein the selected infrastructure includes at least one framework configured to perform model tuning and hyperparameters optimization using at least Bayesian optimization and metaheuristics, at least one container configured to run stateful services, and at least one graphics processing unit" is an example of an additional element that is sufficient to amount to more than the judicial exception. Accordingly, these features of claims 1, 8, and 15 do recite patentable subject matter. Further, these features of claims 1, 7, and 13 cause the claim as a whole to be integrated into a practical application, such as an improvement to the functioning of a computer.
Examiner’s response:
The examiner understands the applicant’s assertion.
However, it appears that each processing step is just applying the abstract idea to a general field of endeavor with additional elements. In addition, improvements to technology or technical field are not necessarily reflected in the claims. Thus, the claim does not integrate the judicial exception into a practical application, and the claim does not amount to significantly more than the judicial exception.
The examiner understands the applicant’s assertion “These operations are performed by computer systems operating on enterprise-network multimedia data and cannot be performed mentally. The claims therefore integrate any alleged abstract idea into a practical application that improves the way network-related multimedia resources are processed and delivered for enterprise-network actions. Identifying the right multimedia tutorials associated with specific enterprise network operation procedures (e.g., troubleshooting a network) can be inefficient without considering context. The features of the independent claim constitute a technological improvement to such enterprise network operations to automatically and efficiently provide the distilled multimedia data related to resolving a network configuration or operation. Accordingly, the claims are not directed to an abstract idea and satisfy Section 101.”
However, it appears that “improves the way network-related multimedia resources are processed and delivered for enterprise-network actions” may be interpreted as an improvement to abstract ideas. Currently, an LM and GNN are used together to perform an actionable task, but the claim languages are not detailed yet, and the actionable task is not specified yet. In addition, “inefficient without considering context” is not persuasive yet since it is not clear what kind of context is for the actionable task. Furthermore, “resolving a network configuration or operation” is not specified yet. For example, elaborating the last limitation of claim 8 and integrating the feature into the independent claims may help provide specificities in resolving a network configuration or operation toward a practical application.
Currently, the limitations do not clearly show e.g., improvements in computer technology and improvements to other technical fields. Rather, the improvements in Remarks are about just improving the abstract ideas of the independent claims. It doesn’t seem that the specification and/or the independent claims clearly show how the inventive concept of the claims enables improvements and how they are tied together. The applicant may need to amend the claims to show how the claim languages and improvements are tied together.
To find a valid improvement to a technology, MPEP 2106.04(d)(1) says the specification must explain the improvement and that the claim must reflect the disclosed improvement. Furthermore, the improvement should not be merely a consequence of the abstract idea. See MPEP 2106.05(a). An improvement in the abstract idea itself is not an improvement to technology.
For at least these reasons, Applicant's arguments are not convincing.
The Examiner encourages Applicant to schedule an interview to discuss issues related to, for example, the rejections noted below under 35 U.S.C § 101.
Applicant’s arguments regarding 35 USC § 102/103 with respect to the independent claims have been considered but are moot because the arguments are directed to amended limitation(s) that has/have not been previously examined.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-18, 20-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“…comprising:
…;
… decompose the multimedia data into network entities and network events, and determine semantic relationships in the multimedia data based on the network entities and the network events;
segmenting the multimedia data into a plurality of multimedia slices based on the semantic relationships;
… based on user persona and context to generate a multi-modal knowledge graph, each multimedia slice of the plurality of multimedia slices representing a node in the multi-modal knowledge graph;
selecting, based on the context, two or more multimedia slices of the plurality of multimedia slices within the multi-modal knowledge graph, and stitching the two or more multimedia slices to generate a distilled multimedia data set; and
performing an actionable task associated with configuration or operation of the enterprise network based on the distilled multimedia data set”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). In particular, the claim recites an additional element(s) (“A computer-implemented”, “executing at least one language model to”, “executing a graph neural network”) – using a device and/or a model to process data. The device and the model in each step are recited at a high-level of generality (i.e., as a generic computer performing a generic computer function of processing data) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
In particular, the claim recites an additional element(s) (“obtaining multimedia data from one or more data sources related to operation or configuration of an enterprise network”) – the act of receiving data. The claim is adding an insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). The act of receiving data is recited at a high-level of generality (i.e., as a generic act of receiving performing a generic act function of receiving data) such that it amounts no more than a mere act to apply the exception using a generic act of receiving. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, with respect to integration of the abstract idea into a practical application, the additional elements of using a generic computer component to perform each step amount to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible. MPEP 2106.05(f).
As discussed above, the claim recites the additional element(s) of receiving data at a high-level of generality and is adding an insignificant extra-solution activity – see MPEP 2106.05(g). However, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood, routine, and conventional. See MPEP 2106.05(d)(II) – “Receiving or transmitting data over a network” or “Storing and retrieving information in memory”. Accordingly, this additional element does not provide an inventive concept and significantly more than the abstract idea. Thus, the claim is not patent eligible.
Regarding claim 2
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“…, and wherein generating the distilled multimedia data set includes:
generating a reduced graph from the multi-modal knowledge graph based on the network persona and a user input; and
selecting at least two multimedia slices for the distilled multimedia data set using the reduced graph.”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element (“wherein the user persona is a network persona”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 3
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“…, and further comprising:
determining the network persona based on past activities performed by the user with respect to the enterprise network.”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element (“wherein the network persona includes a skill level of a user”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 4
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“…, and the method further comprising:
encoding the user input to generate an input embedding; and
generating the reduced graph from the multi-modal knowledge graph based on the input embedding”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element (“wherein the actionable task is configuring the enterprise network”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 5
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“…;
encoding the user input to generate a plurality of input embeddings, …; and
generating the reduced graph from the multi-modal knowledge graph based on the plurality of input embeddings to provide interactive learning using the distilled multimedia data set”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element(s) (“obtaining the user input comprising at least two search queries”) – the act of receiving data. The claim is adding an insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). The act of receiving data is recited at a high-level of generality (i.e., as a generic act of receiving performing a generic act function of receiving data) such that it amounts no more than a mere act to apply the exception using a generic act of receiving. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
In particular, the claim recites an additional element (“each of the plurality of input embeddings being specific to one of the at least two search queries”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, the claim recites the additional element(s) of receiving data at a high-level of generality and is adding an insignificant extra-solution activity – see MPEP 2106.05(g). However, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood, routine, and conventional. See MPEP 2106.05(d)(II) – “Receiving or transmitting data over a network” or “Storing and retrieving information in memory”. Accordingly, this additional element does not provide an inventive concept and significantly more than the abstract idea. Thus, the claim is not patent eligible.
This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 6
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“converting an audio portion of the multimedia data to text;
determining semantic relationships in the text, …; and
generating the multi-modal knowledge graph based on the semantic relationships in which the multimedia data is segmented into the plurality of multimedia slices represented by respective nodes in the multi-modal knowledge graph, …”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element (“wherein the semantic relationships include entities and events in the text”, “wherein at least one of the plurality of multimedia slices includes a portion of the text, at least one respective semantic relationship, a corresponding audio portion of the multimedia data, and a corresponding video portion of the multimedia data”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 7
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“segmenting a video of the multimedia data into a plurality of video slices to map the entities and the events in the text to the video of the multimedia data”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible.
Regarding claim 8
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element (“wherein the multimedia data comprises one or more of: a plurality of network related video learning seminars for configuring one or more network devices in the enterprise network; a plurality of network related video tutorials for obtaining operating data of the one or more network devices in the enterprise network; a plurality of network related videos for progressing a network technology along an adoption lifecycle; and a plurality of troubleshooting videos that address one or more network issues by performing the actionable task associated with the enterprise network that changes a configuration of one or more affected network devices in the enterprise network”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 9
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element(s) (“wherein the multimedia data is obtained from different data sources that provide video recordings associated with the operation or the configuration of the enterprise network”) – the act of receiving data. The claim is adding an insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). The act of receiving data is recited at a high-level of generality (i.e., as a generic act of receiving performing a generic act function of receiving data) such that it amounts no more than a mere act to apply the exception using a generic act of receiving. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, the claim recites the additional element(s) of receiving data at a high-level of generality and is adding an insignificant extra-solution activity – see MPEP 2106.05(g). However, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood, routine, and conventional. See MPEP 2106.05(d)(II) – “Receiving or transmitting data over a network” or “Storing and retrieving information in memory”. Accordingly, this additional element does not provide an inventive concept and significantly more than the abstract idea. Thus, the claim is not patent eligible.
Regarding claim 10
The claim recites “An apparatus comprising: a memory; a network interface configured to enable network communications; and a processor, wherein the processor is configured to perform a method comprising:” to perform precisely the method of Claim 1. As performance of an abstract idea on generic computer components (see MPEP 2106.05(f)) cannot integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself, the claim is rejected for reasons set forth in the rejection of Claim 1.
Regarding claim 11
The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 12
The claim is rejected for the reasons set forth in the rejection of Claim 3 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 13
The claim is rejected for the reasons set forth in the rejection of Claim 4 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 14
The claim is rejected for the reasons set forth in the rejection of Claim 5 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 15
The claim is rejected for the reasons set forth in the rejection of Claim 6 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 16
The claim is rejected for the reasons set forth in the rejection of Claim 7 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 17
The claim recites “One or more non-transitory computer readable storage media encoded with software comprising computer executable instructions that, when executed by a processor, cause the processor to perform a method including:” to perform precisely the method of Claim 1. As performance of an abstract idea on generic computer components (see MPEP 2106.05(f)) and “Storing and retrieving information in memory” (see MPEP 2106.05(g) on Insignificant Extra-Solution Activity, and MPEP 2106.05(d) on Well-Understood, Routine, Conventional Activity) cannot integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself, the claim is rejected for reasons set forth in the rejection of Claim 1.
Regarding claim 18
The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 20
The claim is rejected for the reasons set forth in the rejection of Claim 4 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 21
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element(s) (“obtaining a user input indicating a troubleshooting query related to operation of the enterprise network; and providing the distilled multimedia data set, which is specific to the user input and the semantic relationships in the multimedia data, for troubleshooting one or more network issues associated with the enterprise network”) – the act of receiving/transmitting data. The claim is adding an insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). The act of receiving/transmitting data is recited at a high-level of generality (i.e., as a generic act of performing a generic act function of receiving/transmitting data) such that it amounts no more than a mere act to apply the exception using a generic act of transmitting. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, the claim recites the additional element(s) of receiving/transmitting data at a high-level of generality and is adding an insignificant extra-solution activity – see MPEP 2106.05(g). However, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood, routine, and conventional. See MPEP 2106.05(d)(II) – “Receiving or transmitting data over a network” or “Storing and retrieving information in memory”. Accordingly, this additional element does not provide an inventive concept and significantly more than the abstract idea. Thus, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 9-14, 17-18, 20-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (VQA-GNN: Reasoning with Multimodal Semantic Graph for Visual Question Answering) in view of Chang et al. (KID: Knowledge Graph-Enabled Intent-Driven Network with Digital Twin)
Regarding claim 1
(Note: Hereinafter, if a limitation has bold brackets (i.e. [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards.)
Wang teaches
A [computer]-implemented method comprising:
(Wang [fig(s) 2] “Overview of VQA-GNN: we first perform the image-level KG retrieval (§4.1) and the concept-level KG retrieval (§4.1) to build a multimodal semantic graph (§4.2), then we perform Multi-relation GNN (§4.3) to joint reason correct answer of the scene (§4.4). Here, "C-SG" and "I-SG" indicate Concept-level and Image-level semantic graph, both are subgraph of the multimodal semantic graph.”;)
obtaining multimedia data from one or more data sources related to operation or configuration of an enterprise [network];
(Wang [sec(s) 4] “Given an image and its related question with an answer choice, first we obtain explicit concept-level and image-level KGs from the visual and textual context, respectively (§4.1). Then we introduce a QA node to connect the concept-level and image level KGs to build a multimodal semantic graph so that we can reason about the correct answer over the joint knowledge resources (§4.2).” [sec(s) 3] “In contrast, we extract a scene graph from the image in order to provide explicit relation information in the image, and build a GNN model with the scene graph to achieve an interpretable reasoning. It does not need extra visual-language dataset for pretraining, and outperforms existing models without pretraining process (Table 1).” [sec(s) 2] “VQA-GNN extracts a scene graph from the given input image using a off-the-shelf scene graph generator to obtain a high-level structured representation of the visual context, and then retrieves relevant linguistic/visual subgraphs from massive knowledge graphs including ConceptNet (Speer et al., 2017) and VisualGenome (Krishna et al., 2017).” [sec(s) 5.1] “We evaluate VQA-GNN on the Visual Common sense Reasoning dataset (VCR). VCR consists of two tasks: visual question answering (Q→A), answer justification (QA→R).”;)
executing at least one language model to decompose the multimedia data into network entities and network events, and determine semantic relationships in the multimedia data based on the network entities and the network events;
(Wang [fig(s) 1] [fig(s) 3-4] [sec(s) 4] “Given an image and its related question with an answer choice, first we obtain explicit concept-level and image-level KGs from the visual and textual context, respectively (§4.1). Then we introduce a QA node to connect the concept-level and image level KGs to build a multimodal semantic graph so that we can reason about the correct answer over the joint knowledge resources (§4.2).” [sec(s) 3] “In contrast, we extract a scene graph from the image in order to provide explicit relation information in the image, and build a GNN model with the scene graph to achieve an interpretable reasoning. It does not need extra visual-language dataset for pretraining, and outperforms existing models without pretraining process (Table 1).” [sec(s) 2] “VQA-GNN extracts a scene graph from the given input image using a off-the-shelf scene graph generator to obtain a high-level structured representation of the visual context, and then retrieves relevant linguistic/visual subgraphs from massive knowledge graphs including ConceptNet (Speer et al., 2017) and VisualGenome (Krishna et al., 2017).” [sec(s) 5.1] “We evaluate VQA-GNN on the Visual Common sense Reasoning dataset (VCR). VCR consists of two tasks: visual question answering (Q→A), answer justification (QA→R).”;)
segmenting the multimedia data into a plurality of multimedia slices based on the semantic relationships;
(Wang [fig(s) 1] “SceneGraph”, “Visual knowledge retrieval”, “Textual knowledge retrieval” [sec(s) Abs] “Specifically, given a question image pair, we build a scene graph from the image, retrieve a relevant linguistic subgraph from ConceptNet and visual subgraph from VisualGenome, and unify these three graphs and the question into one joint graph, multimodal semantic graph.” [sec(s) 1] “Finally, VQA-GNN performs message passing and reasons on this structured multimodal semantic graph with graph neural networks (GNNs) to score each answer candidate.” [sec(s) 4] “A diagram of VQA-GNN is shown in Figure 2. Given an image and its related question with an answer choice, first we obtain explicit concept-level and image-level KGs from the visual and textual context, respectively (§4.1). Then we introduce a QA node to connect the concept-level and image level KGs to build a multimodal semantic graph so that we can reason about the correct answer over the joint knowledge resources (§4.2). Here, node z represents the QA textual context that is the concatenation of a question and an answer choice. In addition to z, we concatenate the retrieved answer choice-relevant global/local image context to get a QA visual context node p. We then propose an attention-based GNN to pass messages across different nodes on the joint multimodal semantic graph (§4.3) for effective reasoning. Finally, we train the model to make the prediction using the concatenation of QA text, image, z, p and pooled working graphs representations (§4.4)” [sec(s) 4.1] In addition to the local image context, with an intuition that the global image context of the correct choice is assumed to be similar to the local image context, we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). … “Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing). Hence, we can obtain a concept-level semantic graph Gacp = (Vacp,Eacp) to represent concept-level knowledge.”;)
executing a graph neural network based on user persona and context to generate a multi-modal knowledge graph, each multimedia slice of the plurality of multimedia slices representing a node in the multi-modal knowledge graph;
(Wang [fig(s) 1-3] [sec(s) Abs] “Specifically, given a question image pair, we build a scene graph from the image, retrieve a relevant linguistic subgraph from ConceptNet and visual subgraph from VisualGenome, and unify these three graphs and the question into one joint graph, multimodal semantic graph.” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph). Finally, VQA-GNN performs message passing and reasons on this structured multimodal semantic graph with graph neural networks (GNNs) to score each answer candidate.” [sec(s) 4.1] “In addition to the local image context, with an intuition that the global image context of the correct choice is assumed to be similar to the local image context, we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). … Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing).” [sec(s) 4.4] “For the training process, we apply the cross entropy loss to optimize the VQA-GNN model end-to-end.” [sec(s) 5.1] “We joint train VQA GNN on Q→A and QA→R, with a common LM encoder, the multimodal semantic graph for Q→A, concept-level semantic graph for QA→R.”;)
selecting, based on the context, two or more multimedia slices of the plurality of multimedia slices within the multi-modal knowledge graph, and stitching the two or more multimedia slices to generate a distilled multimedia data set; and
(Wang [fig(s) 1] [sec(s) Abs] “Specifically, given a question image pair, we build a scene graph from the image, retrieve a relevant linguistic subgraph from ConceptNet and visual subgraph from VisualGenome, and unify these three graphs and the question into one joint graph, multimodal semantic graph.” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph). Finally, VQA-GNN performs message passing and reasons on this structured multimodal semantic graph with graph neural networks (GNNs) to score each answer candidate.” [sec(s) 4.1] “In addition to the local image context, with an intuition that the global image context of the correct choice is assumed to be similar to the local image context, we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). … Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing).” [sec(s) 4.4] “For the training process, we apply the cross entropy loss to optimize the VQA-GNN model end-to-end.” [sec(s) 5.1] “We joint train VQA GNN on Q→A and QA→R, with a common LM encoder, the multimodal semantic graph for Q→A, concept-level semantic graph for QA→R.”;)
performing an actionable task associated with configuration or operation of the enterprise [network] based on the distilled multimedia data set.
(Wang [fig(s) 1] [sec(s) Abs] “Specifically, given a question image pair, we build a scene graph from the image, retrieve a relevant linguistic subgraph from ConceptNet and visual subgraph from VisualGenome, and unify these three graphs and the question into one joint graph, multimodal semantic graph.” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph). Finally, VQA-GNN performs message passing and reasons on this structured multimodal semantic graph with graph neural networks (GNNs) to score each answer candidate.” [sec(s) 4.1] “In addition to the local image context, with an intuition that the global image context of the correct choice is assumed to be similar to the local image context, we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). … Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing).” [sec(s) 4.4] “For the training process, we apply the cross entropy loss to optimize the VQA-GNN model end-to-end.” [sec(s) 5.1] “We joint train VQA GNN on Q→A and QA→R, with a common LM encoder, the multimodal semantic graph for Q→A, concept-level semantic graph for QA→R.”;)
However, Wang does not appear to explicitly teach:
A [computer]-implemented method comprising:
obtaining multimedia data from one or more data sources related to operation or configuration of an enterprise [network];
performing an actionable task associated with configuration or operation of the enterprise [network] based on the distilled multimedia data set.
(Note: Hereinafter, if a limitation has one or more bold underlines, the one or more underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references.)
Chang teaches
A computer-implemented method comprising:
(Chang [fig(s) 4] “The process of forming the IKG” [sec(s) III.A] “Considering the characteristics of the KG technology itself, we design a generic intent refinement model on the basis of the KG. Specifically, its working mechanism is as follows: At first, the system pre-processes the intent data after receiving it. The program reads the data of intent input and loads it into memory. Following that, the regular matching rules are invoked to eliminate special symbols. Finally, the cleaned and filtered intent data is stored in the file.” [sec(s) IV.B] “We create two virtual machines based on Ubuntu 16.04 system. One is used to run the KG-based intent translation system. And the other is used to build the network simulation environment and construct the NSKG.”;)
obtaining multimedia data from one or more data sources related to operation or configuration of an enterprise network;
(Chang [fig(s) 1] [sec(s) II.D] “The physical network layer is a real-world network environment. It is composed primarily of the physical entities that comprise the end-to-end network, such as terminals, switches, routers, and other network element devices” [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “To improve the accuracy in representing intents, the initial IKG must be extended. As a result of intent expansion, the corresponding conditions and attributes are added to the intent entities. For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added. It’s worth noting that the intent expansion module will expand the IKG in accordance with the information from the digital twin layer.” [sec(s) III.B] “The diagram of the interaction between the IKG and the NSKG is shown in Fig. 3. The system gradually implements intent expansion, parameter mapping, and intent verification through interaction between both KGs. It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
performing an actionable task associated with configuration or operation of the enterprise network based on the distilled multimedia data set.
(Chang [sec(s) II.D] “The physical network layer is a real-world network environment. It is composed primarily of the physical entities that comprise the end-to-end network, such as terminals, switches, routers, and other network element devices” [sec(s) I] “The NSKG is capable of dynamically updating the attribute information of nodes and relations to reflect changes in network status.” [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added.” [sec(s) III.B] “Then, the system dynamically updates the attributes of nodes and relations in the NSKG based on the indicator data collected in real-time. … It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Wang with the enterprise network of Chang.
One of ordinary skill in the art would have been motived to combine in order to merge the semantic gap between user intents, policy generation, and the underlying network toward feasibility and effectiveness, and contribute to the ability of the intent-driven network to adapt to the network autonomously.
(Chang [sec(s) I] “• We present an IDN framework named the KID, aiming to merge the semantic gap between user intents, policy generation, and the underlying network. • We develop an intent language specification and a generic intent refinement model. Additionally, an intelligent translation technique is designed to achieve intelligent intent translation which takes user intents into consideration as well as network status. • Finally, we provide a use case of the presented KID framework to demonstrate the feasibility and effectiveness of the proposed scheme.” [sec(s) II.B] “It contributes to the ability of the IDN to adapt to the network autonomously”)
Regarding claim 2
The combination of Wang, Chang teaches claim 1.
Wang further teaches
wherein the user persona is a [network] persona, and wherein generating the distilled multimedia data set includes:
(Wang [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph). Finally, VQA-GNN performs message passing and reasons on this structured multimodal semantic graph with graph neural networks (GNNs) to score each answer candidate.” [sec(s) 4.1] “In addition to the local image context, with an intuition that the global image context of the correct choice is assumed to be similar to the local image context, we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). … Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing). Hence, we can obtain a concept-level semantic graph Gacp = (Vacp,Eacp) to represent concept-level knowledge.”;)
generating a reduced graph from the multi-modal knowledge graph based on the [network] persona and a user input; and
(Wang [fig(s) 2] “Overview of VQA-GNN: we first perform the image-level KG retrieval (§4.1) and the concept-level KG retrieval (§4.1) to build a multimodal semantic graph (§4.2), then we perform Multi-relation GNN (§4.3) to joint reason correct answer of the scene (§4.4). Here, "C-SG" and "I-SG" indicate Concept-level and Image-level semantic graph, both are subgraph of the multimodal semantic graph” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph).” [sec(s) 4.1] “Step 2-1: We use grounded phases to retrieve their 1-hop neighbor nodes from the ConceptNet KG. Step 2-2: As many retrieved concept nodes that are semantically irrelevant to the answer choice, in spired by QA-GNN (Yasunaga et al., 2021), we introduce relevance score to prune irrelevance nodes. We use a word2vec model released by the spaCy library1 to get relevance score between concept node candidates and answer choices. As a result, given an answer choice, we can retrieve a relevance subgraph from ConceptNet KG based on the relevance score. To better comprehend concept knowledge from the image as well, Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing). Hence, we can obtain a concept-level semantic graph Gacp = (Vacp, Eacp) to represent concept-level knowledge.”;)
selecting at least two multimedia slices for the distilled multimedia data set using the reduced graph.
(Wang [fig(s) 2] “Here, "C-SG" and "I-SG" indicate Concept-level and Image-level semantic graph, both are subgraph of the multimodal semantic graph” [sec(s) 4.1] “we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). Here we experimentally use top 10 of retrieved result and embed them using a same object detector for local image context, and then take their mean to get a QA visual context node p. … Step 2-1: We use grounded phases to retrieve their 1-hop neighbor nodes from the ConceptNet KG. Step 2-2: As many retrieved concept nodes that are semantically irrelevant to the answer choice, in spired by QA-GNN (Yasunaga et al., 2021), we introduce relevance score to prune irrelevance nodes. We use a word2vec model released by the spaCy library1 to get relevance score between concept node candidates and answer choices. As a result, given an answer choice, we can retrieve a relevance subgraph from ConceptNet KG based on the relevance score. To better comprehend concept knowledge from the image as well, Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing). Hence, we can obtain a concept-level semantic graph Gacp = (Vacp, Eacp) to represent concept-level knowledge.”;)
Chang teaches
wherein the user persona is a network persona
(Chang [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “To improve the accuracy in representing intents, the initial IKG must be extended. As a result of intent expansion, the corresponding conditions and attributes are added to the intent entities. For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added. It’s worth noting that the intent expansion module will expand the IKG in accordance with the information from the digital twin layer.” [sec(s) III.B] “The diagram of the interaction between the IKG and the NSKG is shown in Fig. 3. The system gradually implements intent expansion, parameter mapping, and intent verification through interaction between both KGs. It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
The combination of Wang, Change is combinable with Chang for the same rationale as set forth above with respect to claim 1.
Regarding claim 3
The combination of Wang, Change teaches claim 2.
Chang further teaches
wherein the network persona includes a skill level of a user, and further comprising:
(Chang [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “To improve the accuracy in representing intents, the initial IKG must be extended. As a result of intent expansion, the corresponding conditions and attributes are added to the intent entities. For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added. It’s worth noting that the intent expansion module will expand the IKG in accordance with the information from the digital twin layer.” [sec(s) III.B] “The diagram of the interaction between the IKG and the NSKG is shown in Fig. 3. The system gradually implements intent expansion, parameter mapping, and intent verification through interaction between both KGs. It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
determining the network persona based on past activities performed by the user with respect to the enterprise network.
(Chang [sec(s) I] “The NSKG is capable of dynamically updating the attribute information of nodes and relations to reflect changes in network status.” [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added.” [sec(s) III.B] “Then, the system dynamically updates the attributes of nodes and relations in the NSKG based on the indicator data collected in real-time. … It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
The combination of Wang, Change is combinable with Chang for the same rationale as set forth above with respect to claim 1.
Regarding claim 4
The combination of Wang, Chang teaches claim 2.
Wang further teaches
encoding the user input to generate an input embedding; and
(Wang [fig(s) 2] “LM encoding” [sec(s) 4.2] “As transformer-based method is powerful for textual representation learning (Su et al., 2020; Lu et al., 2019), we employ the RoBERTa LM (Liu et al., 2019) as the encoder of QA textual context node z and finetune it with VQA-GNN. z =LM_encoder(Q,A)” [sec(s) 5.1] “We joint train VQA GNN on Q→A and QA→R, with a common LM encoder, the multimodal semantic graph for Q→A, concept-level semantic graph for QA→R. We use a pretrained RoBERTa Large model to embed QA textual context node, and finetune it with GNN components for 50 epoch by using learning rate 1e-5 and 1e-4 respectively.”;)
generating the reduced graph from the multi-modal knowledge graph based on the input embedding.
(Wang [fig(s) 2] “Overview of VQA-GNN: we first perform the image-level KG retrieval (§4.1) and the concept-level KG retrieval (§4.1) to build a multimodal semantic graph (§4.2), then we perform Multi-relation GNN (§4.3) to joint reason correct answer of the scene (§4.4). Here, "C-SG" and "I-SG" indicate Concept-level and Image-level semantic graph, both are subgraph of the multimodal semantic graph” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph).” [sec(s) 4.1] “Step 2-1: We use grounded phases to retrieve their 1-hop neighbor nodes from the ConceptNet KG. Step 2-2: As many retrieved concept nodes that are semantically irrelevant to the answer choice, in spired by QA-GNN (Yasunaga et al., 2021), we introduce relevance score to prune irrelevance nodes. We use a word2vec model released by the spaCy library1 to get relevance score between concept node candidates and answer choices. As a result, given an answer choice, we can retrieve a relevance subgraph from ConceptNet KG based on the relevance score. To better comprehend concept knowledge from the image as well, Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing). Hence, we can obtain a concept-level semantic graph Gacp = (Vacp, Eacp) to represent concept-level knowledge.”;)
Chang further teaches
wherein the user input includes an actionable task related to the configuration of the enterprise network, and further comprising:
(Chang [sec(s) II.D] “The physical network layer is a real-world network environment. It is composed primarily of the physical entities that comprise the end-to-end network, such as terminals, switches, routers, and other network element devices” [sec(s) I] “The NSKG is capable of dynamically updating the attribute information of nodes and relations to reflect changes in network status.” [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added.” [sec(s) III.B] “Then, the system dynamically updates the attributes of nodes and relations in the NSKG based on the indicator data collected in real-time. … It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
The combination of Wang, Change is combinable with Chang for the same rationale as set forth above with respect to claim 1.
Regarding claim 5
The combination of Wang, Chang teaches claim 2.
obtaining the user input comprising at least two search queries;
(Wang [fig(s) 2] “QA Text” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph).” [sec(s) 4] “A diagram of VQA-GNN is shown in Figure 2. Given an image and its related question with an answer choice, first we obtain explicit concept-level and image-level KGs from the visual and textual context, respectively (§4.1). Then we introduce a QA node to connect the concept-level and image level KGs to build a multimodal semantic graph so that we can reason about the correct answer over the joint knowledge resources (§4.2). Here, node z represents the QA textual context that is the concatenation of a question and an answer choice. In addition to z, we concatenate the retrieved answer choice-relevant global/local image context to get a QA visual context node p. We then propose an attention-based GNN to pass messages across different nodes on the joint multimodal semantic graph (§4.3) for effective reasoning. Finally, we train the model to make the prediction using the concatenation of QA text, image, z, p and pooled working graphs representations (§4.4).”;)
encoding the user input to generate a plurality of input embeddings, each of the plurality of input embeddings being specific to one of the at least two search queries; and
(Wang [fig(s) 2] “LM encoding” [sec(s) 4.2] “As transformer-based method is powerful for textual representation learning (Su et al., 2020; Lu et al., 2019), we employ the RoBERTa LM (Liu et al., 2019) as the encoder of QA textual context node z and finetune it with VQA-GNN. z =LM_encoder(Q,A)” [sec(s) 5.1] “We joint train VQA GNN on Q→A and QA→R, with a common LM encoder, the multimodal semantic graph for Q→A, concept-level semantic graph for QA→R. We use a pretrained RoBERTa Large model to embed QA textual context node, and finetune it with GNN components for 50 epoch by using learning rate 1e-5 and 1e-4 respectively.”;)
generating the reduced graph from the multi-modal knowledge graph based on the plurality of input embeddings to provide interactive learning using the distilled multimedia data set.
(Wang [fig(s) 2] “Overview of VQA-GNN: we first perform the image-level KG retrieval (§4.1) and the concept-level KG retrieval (§4.1) to build a multimodal semantic graph (§4.2), then we perform Multi-relation GNN (§4.3) to joint reason correct answer of the scene (§4.4). Here, "C-SG" and "I-SG" indicate Concept-level and Image-level semantic graph, both are subgraph of the multimodal semantic graph” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph).” [sec(s) 4.1] “we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). … Step 2-1: We use grounded phases to retrieve their 1-hop neighbor nodes from the ConceptNet KG. Step 2-2: As many retrieved concept nodes that are semantically irrelevant to the answer choice, in spired by QA-GNN (Yasunaga et al., 2021), we introduce relevance score to prune irrelevance nodes. We use a word2vec model released by the spaCy library1 to get relevance score between concept node candidates and answer choices. As a result, given an answer choice, we can retrieve a relevance subgraph from ConceptNet KG based on the relevance score. To better comprehend concept knowledge from the image as well, Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing). Hence, we can obtain a concept-level semantic graph Gacp = (Vacp, Eacp) to represent concept-level knowledge.”;)
Regarding claim 9
The combination of Wang, Chang teaches claim 1.
Chang further teaches
wherein the multimedia data is obtained from different data sources that provide video recordings associated with the operation or the configuration of the enterprise network.
(Chang [sec(s) IV.A] “In this use case, the entity "video service" has two conditions in the NSKG: "bandwidth" and "delay", with default attribute values. Therefore, when the intent knowledge is imported into Neo4j, the system appends corresponding conditions to the "video service". Meanwhile, the attribute values of both conditions are updated into the IKG. When the system detects that there is no intent entity can be expanded according to the NSKG, the intent expansion is completed” [sec(s) IV.B] “Finally, the physical performance of the KID is evaluated in terms of ensuring video service. We use the video lan client to simulate end-to-end video transmission in Mininet. And the network topology is identical to that shown in Fig. 2. Additionally, the link’s maximum bandwidth is set to 20M, and background traffic is pre-injected into the simulation network. Then, as the intent is issued in the simulation network, the system plans the path for intent execution. The routers follow the control path issued by the ONOS when forwarding packets for end-to-end video streams. Then, in the process of ensuring video service, iPerf is used to test the size of the link packet loss rate. Eventually, we compare the performance of the KG-based path algorithm against the shortest path algorithm in terms of reducing link packet loss rate in Fig. 6(b). Fig. 6(b) illustrates the effect of the KG-based path algorithm and the shortest path algorithm on the packet loss rate links. As seen in Fig. 6(b), the packet loss rate of links is roughly at a stable level after the end-to-end video service is established. When the system forwards packets using the shortest path algorithm, the link packet loss rate is approximately 45%. The link packet loss rate is approximately 13% when the system forwards packets along the paths planned by the KG-based path algorithm. Clearly, the KID has good physical performance when it comes to ensuring intent execution.”;)
The combination of Wang, Change is combinable with Chang for the same rationale as set forth above with respect to claim 1.
Regarding claim 10
The claim is a system claim corresponding to the method claim 1, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
In addition, Chang further teaches
a network interface configured to enable network communications; and
(Chang [fig(s) 1] “Physical Network”, “ONOS”, “Collect Data” [sec(s) II.A] “The user or application intent input module is designed at the application layer. It is primarily aimed at providing interfaces to the users and applications. Meanwhile, this module enables various sources and types of intent data to be entered into the system through front-end interface. It also allows users to declare their intents via text or voice. Then the intent data is sent to the data pre-processing module in the intent layer.” [sec(s) II.D] “The physical network layer is a real-world network environment. It is composed primarily of the physical entities that comprise the end-to-end network, such as terminals, switches, routers, and other network element devices.” [sec(s) IV.B] “Fig. 5 shows the network topology diagram displayed in the ONOS control interface. It illustrates how the KID ensures end-to-end intent.”;)
The combination of Wang, Change is combinable with Chang for the same rationale as set forth above with respect to claim 1.
Regarding claim 11
The claim is a system claim corresponding to the method claim 2, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 12
The claim is a system claim corresponding to the method claim 3, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 13
The claim is a system claim corresponding to the method claim 4, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 14
The claim is a system claim corresponding to the method claim 5, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 17
The claim is a computer readable storage media claim corresponding to the method claim 1, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 18
The claim is a computer readable storage media claim corresponding to the method claim 2, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 20
The claim is a computer readable storage media claim corresponding to the method claim 4, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 21
The combination of Wang, Chang teaches claim 1.
Wang further teaches
providing the distilled multimedia data set, which is specific to the user input and the semantic relationships in the multimedia data, for [troubleshooting one or more network issues associated with the enterprise network].
(Wang [fig(s) 1] [sec(s) Abs] “Specifically, given a question image pair, we build a scene graph from the image, retrieve a relevant linguistic subgraph from ConceptNet and visual subgraph from VisualGenome, and unify these three graphs and the question into one joint graph, multimodal semantic graph.” [sec(s) 1] “Given the inferred scene graph and the retrieved knowledge subgraphs, we combine them into a multimodal semantic graph by introducing super nodes that link relevant concepts from the graphs. The multimodal semantic graph consists of QA-context aware concepts (from the scene graph) as well as the QA-context-agnostic concepts (from the knowledge graph). Finally, VQA-GNN performs message passing and reasons on this structured multimodal semantic graph with graph neural networks (GNNs) to score each answer candidate.” [sec(s) 4.1] “In addition to the local image context, with an intuition that the global image context of the correct choice is assumed to be similar to the local image context, we use a pretrained sentence-BERT model to calculate the similarity between each answer choice with all region descriptions of region image in the Visual Genome dataset so that we can get relevance region images to represent the global image context of each choice (Reimers and Gurevych, 2019). … Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing).” [sec(s) 4.4] “For the training process, we apply the cross entropy loss to optimize the VQA-GNN model end-to-end.” [sec(s) 5.1] “We joint train VQA GNN on Q→A and QA→R, with a common LM encoder, the multimodal semantic graph for Q→A, concept-level semantic graph for QA→R.”;)
Chang further teaches
obtaining a user input indicating a troubleshooting query related to operation of the enterprise network; and
(Chang [sec(s) II.D] “The physical network layer is a real-world network environment. It is composed primarily of the physical entities that comprise the end-to-end network, such as terminals, switches, routers, and other network element devices” [sec(s) I] “The NSKG is capable of dynamically updating the attribute information of nodes and relations to reflect changes in network status.” [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added.” [sec(s) III.B] “Then, the system dynamically updates the attributes of nodes and relations in the NSKG based on the indicator data collected in real-time. … It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
providing the distilled multimedia data set, which is specific to the user input and the semantic relationships in the multimedia data, for troubleshooting one or more network issues associated with the enterprise network.
(Chang [sec(s) II.D] “The physical network layer is a real-world network environment. It is composed primarily of the physical entities that comprise the end-to-end network, such as terminals, switches, routers, and other network element devices” [sec(s) I] “The NSKG is capable of dynamically updating the attribute information of nodes and relations to reflect changes in network status.” [sec(s) IV] “We input the intent as "Transmit an important video service from Xian user A to Beijing user B, the time requirement is from 10:30 on June 5, 2020, to 12:30 on June 5, 2020". … When the system reads the intent entity "Beijing user B" and queries the node "Beijing user B" in the NSKG, it updates the attribute values of the node, such as IP address 10.0.0.4 and port number 1257, to the node "Beijing user B" in the IKG. Furthermore, similar steps are repeated until corresponding network parameters is added for all intent entities and relations.” [sec(s) III.A] “For instance, if the user types "establish a reassurance level voice service from A to B", the phrase "reassurance-level voice service" will be parsed. The necessary requirements for "reassurance-level voice service", such as delay, bandwidth, and other parameters, are added.” [sec(s) III.B] “Then, the system dynamically updates the attributes of nodes and relations in the NSKG based on the indicator data collected in real-time. … It is in these processes that network policies are gradually formed. After implementing the parameter mapping, the system parses the IKG to understand the requirements of user intents. These intent requirements are expressed in terms of parameter information such as bandwidth, delay, and packet loss rate.”;)
The combination of Wang, Chang is combinable with Chang for the same rationale as set forth above with respect to claim 1.
Claim(s) 6-7, 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (VQA-GNN: Reasoning with Multimodal Semantic Graph for Visual Question Answering) in view of Chang et al. (KID: Knowledge Graph-Enabled Intent-Driven Network with Digital Twin) in view of Lei et al. (TVQA: Localized, Compositional Video Question Answering)
Regarding claim 6
The combination of Wang, Chang teaches claim 1.
Wang further teaches
generating the multi-modal knowledge graph based on the semantic relationships in which the multimedia data is segmented into the plurality of multimedia slices represented by respective nodes in the multi-modal knowledge graph, [wherein at least one of the plurality of multimedia slices includes a portion of the text, at least one respective semantic relationship, a corresponding audio portion of the multimedia data, and a corresponding video portion of the multimedia data].
(Wang [fig(s) 1] “SceneGraph”, “Visual knowledge retrieval”, “Textual knowledge retrieval” [sec(s) Abs] “Specifically, given a question image pair, we build a scene graph from the image, retrieve a relevant linguistic subgraph from ConceptNet and visual subgraph from VisualGenome, and unify these three graphs and the question into one joint graph, multimodal semantic graph.” [sec(s) 1] “Finally, VQA-GNN performs message passing and reasons on this structured multimodal semantic graph with graph neural networks (GNNs) to score each answer candidate.” [sec(s) 4] “A diagram of VQA-GNN is shown in Figure 2. Given an image and its related question with an answer choice, first we obtain explicit concept-level and image-level KGs from the visual and textual context, respectively (§4.1). Then we introduce a QA node to connect the concept-level and image level KGs to build a multimodal semantic graph so that we can reason about the correct answer over the joint knowledge resources (§4.2). Here, node z represents the QA textual context that is the concatenation of a question and an answer choice. In addition to z, we concatenate the retrieved answer choice-relevant global/local image context to get a QA visual context node p. We then propose an attention-based GNN to pass messages across different nodes on the joint multimodal semantic graph (§4.3) for effective reasoning. Finally, we train the model to make the prediction using the concatenation of QA text, image, z, p and pooled working graphs representations (§4.4)” [sec(s) 4.1] “Step 3: In addition to linking adjacent object entities in the ConceptNet KGdomain, we also combine them with retrieved subgraph by matching and linking relevance concept entities, e.g., (bottle, atlocation, beverage), (book, relatedto, thing). Hence, we can obtain a concept-level semantic graph Gacp = (Vacp,Eacp) to represent concept-level knowledge.”;)
However, the combination of Wang, Chang does not appear to explicitly teach:
converting an audio portion of the multimedia data to text;
determining semantic relationships in the text, wherein the semantic relationships include entities and events in the text; and
generating the multi-modal knowledge graph based on the semantic relationships in which the multimedia data is segmented into the plurality of multimedia slices represented by respective nodes in the multi-modal knowledge graph, [wherein at least one of the plurality of multimedia slices includes a portion of the text, at least one respective semantic relationship, a corresponding audio portion of the multimedia data, and a corresponding video portion of the multimedia data].
Lei teaches
converting an audio portion of the multimedia data to text;
(Lei [fig(s) 1] “some questions can be answered using subtitles or videos alone, while some require information from both modalities” [fig(s) 4] “For brevity, we only show regional visual features (upper) and subtitle (bottom) streams.” [fig(s) 6] “Top row shows correct predictions, bottom row shows failure cases. Ground truth answers are in green, and the model predictions are indicated by
PNG
media_image1.png
36
32
media_image1.png
Greyscale
.” [fig(s) 10] “Ground truth answers are highlighted in green, and model predictions are indicated by
PNG
media_image1.png
36
32
media_image1.png
Greyscale
.” [sec(s) 1] “With video-QA in particular, as opposed to image-QA, the video itself often comes with as sociated natural language in the form of (subtitle) dialogue. We argue that this is an important area to study because it reflects the real world, where people interact through language, and where many computational systems like robots or other intelligent agents will ultimately have to operate. As such, systems will need to combine information from what they see with what they hear, to pose and answer questions about what is happening. … we provide the dialogue (character name + subtitle) for each QA video clip. Understanding the relationship between the provided dialogue and the question-answer pairs is crucial for correctly answering many of the collected questions. Fourth, our questions are compositional, requiring algorithms to localize relevant moments (START and END points are provided for each question).”;)
determining semantic relationships in the text, wherein the semantic relationships include entities and events in the text; and
(Lei [fig(s) 4] “Illustration of our multi-stream model for Multi-Modal Video QA. Our full model takes different con textual sources (regional visual features, visual concept features, and subtitles) along with question-answer pair as inputs to each stream. For brevity, we only show regional visual features (upper) and subtitle (bottom) streams.” [sec(s) Abs] “In this paper, we present TVQA, a large scale video QA dataset based on 6 popular TV shows. TVQA consists of 152,545 QA pairs from 21,793 clips, spanning over 460 hours of video. Questions are designed to be compositional in nature, requiring systems to jointly localize relevant moments within a clip, comprehend subtitle-based dialogue, and recognize relevant visual concepts.” [sec(s) 1] “we provide the dialogue (character name + subtitle) for each QA video clip. Understanding the relationship between the provided dialogue and the question-answer pairs is crucial for correctly answering many of the collected questions. Fourth, our questions are compositional, requiring algorithms to localize relevant moments (START and ENDpoints are provided for each question).” [sec(s) 2] “While questions in these text QA datasets are specifically designed for language understanding, TVQA questions require both vision understanding and language understanding. Although methods developed for text QA are not directly applicable to TVQA tasks, they can provide inspiration for designing suitable models.”;)
generating the multi-modal knowledge graph based on the semantic relationships in which the multimedia data is segmented into the plurality of multimedia slices represented by respective nodes in the multi-modal knowledge graph, wherein at least one of the plurality of multimedia slices includes a portion of the text, at least one respective semantic relationship, a corresponding audio portion of the multimedia data, and a corresponding video portion of the multimedia data.
(Lei [fig(s) 4] “Illustration of our multi-stream model for Multi-Modal Video QA. Our full model takes different con textual sources (regional visual features, visual concept features, and subtitles) along with question-answer pair as inputs to each stream. For brevity, we only show regional visual features (upper) and subtitle (bottom) streams.” [sec(s) Abs] “In this paper, we present TVQA, a large scale video QA dataset based on 6 popular TV shows. TVQA consists of 152,545 QA pairs from 21,793 clips, spanning over 460 hours of video. Questions are designed to be compositional in nature, requiring systems to jointly localize relevant moments within a clip, comprehend subtitle-based dialogue, and recognize relevant visual concepts.” [sec(s) 1] “we provide the dialogue (character name + subtitle) for each QA video clip. Understanding the relationship between the provided dialogue and the question-answer pairs is crucial for correctly answering many of the collected questions. Fourth, our questions are compositional, requiring algorithms to localize relevant moments (START and ENDpoints are provided for each question).” [sec(s) 2] “While questions in these text QA datasets are specifically designed for language understanding, TVQA questions require both vision understanding and language understanding. Although methods developed for text QA are not directly applicable to TVQA tasks, they can provide inspiration for designing suitable models.”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Wang, Chang with the multimedia data of Lei.
One of ordinary skill in the art would have been motived to combine in order to improve the classification performance with a better accuracy using different input representations.
(Lei [sec 5.2] “Rows 11-18 show results of our model with different contextual inputs and features. The model that only uses question-answer pairs (row 11) achieves 43.50% accuracy. Compared to the subtitle model (row 15), adding video as additional sources (row 16-18) improves performance. … Overall, the best performance is achieved by using all the contextual sources, including subtitles and videos (using concept features, row 18).”)
Regarding claim 7
The combination of Wang, Chang teaches claim 6.
However, the combination of Wang, Chang does not appear to explicitly teach:
segmenting a video of the multimedia data into a plurality of video slices to map the entities and the events in the text to the video of the multimedia data.
Lei teaches
segmenting a video of the multimedia data into a plurality of video slices to map the entities and the events in the text to the video of the multimedia data.
(Lei [fig(s) 1] “Examples from the TVQA dataset. All questions and answers are attached to 60-90 seconds long clips. For visualization purposes, we only show a few of the most relevant frames here. As illustrated above, some questions can be answered using subtitles or videos alone, while some require information from both modalities” [fig(s) 4] “Figure 4: Illustration of our multi-stream model for Multi-Modal Video QA. Our full model takes different con textual sources (regional visual features, visual concept features, and subtitles) along with question-answer pair as inputs to each stream. For brevity, we only show regional visual features (upper) and subtitle (bottom) streams.” [fig(s) 6] “Figure 6: Example predictions from our best model. Top row shows correct predictions, bottom row shows failure cases. Ground truth answers are in green, and the model predictions are indicated by
PNG
media_image1.png
36
32
media_image1.png
Greyscale
.” [fig(s) 10] “More examples from TVQA dataset, along with the predictions from our best model. Ground truth answers are highlighted in green, and model predictions are indicated by
PNG
media_image1.png
36
32
media_image1.png
Greyscale
.” [sec(s) 5] “Qualitative Analysis: Fig. 6 shows example predictions from our S+V+Q model (row 18) using full-length video and subtitle. Fig. 6a and Fig. 6b demonstrate its ability to solve both grounded visual questions and textual reasoning question. Bottom row shows two incorrect predictions.” [sec(s) A] “More QA examples: We show more examples using our S+V+Q joint model trained on full-length video and subtitle in Fig. 10, where the top three rows show the correct answers made by our model and the bottom two rows show the wrong ones.”;)
Regarding claim 15
The claim is a system claim corresponding to the method claim 6, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 16
The claim is a system claim corresponding to the method claim 7, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (VQA-GNN: Reasoning with Multimodal Semantic Graph for Visual Question Answering) in view of Chang et al. (KID: Knowledge Graph-Enabled Intent-Driven Network with Digital Twin) in view of Baykal et al. (US 20150083795 A1)
Regarding claim 8
The combination of Wang, Chang teaches claim 1.
However, the combination of Wang, Chang does not appear to explicitly teach:
wherein the multimedia data comprises one or more of:
a plurality of network related video learning seminars for configuring one or more network devices in the enterprise network;
a plurality of network related video tutorials for obtaining operating data of the one or more network devices in the enterprise network;
a plurality of network related videos for progressing a network technology along an adoption lifecycle; and
a plurality of troubleshooting videos that address one or more network issues by performing the actionable task associated with the enterprise network that changes a configuration of one or more affected network devices in the enterprise network.
Baykal teaches
wherein the multimedia data comprises one or more of:
a plurality of network related video learning seminars for configuring one or more network devices in the enterprise network;
a plurality of network related video tutorials for obtaining operating data of the one or more network devices in the enterprise network;
a plurality of network related videos for progressing a network technology along an adoption lifecycle; and
a plurality of troubleshooting videos that address one or more network issues by performing the actionable task associated with the enterprise network that changes a configuration of one or more affected network devices in the enterprise network.
(Baykal [par(s) 24-29] “In addition, data can be received from the back office system via the network interface 209. For example, support materials, such as installation or troubleshooting manuals, videos, tips, etc. can be provided over the network interface 209 and communicated to the user via the HMI 211 for use during installation or maintenance of a device. Additionally, in some embodiments, support materials include information obtained during an install of a customer or network device that is provided from the back office system over the network interface 209 while troubleshooting the customer or network device. For example, photos of the wiring, or orientation of the customer or network device, etc. can be provided to help troubleshoot an error and, thereby, reduce the time needed to troubleshoot the error. … one or more network devices are optionally provisioned by the back office system based on the received device identification and secondary information at block 308. In particular, the one or more network devices are provisioned for the desired service. Additionally, in some embodiments, the back office system optionally provides suggestions to a user based on the received device identification and secondary information at block 310. For example, the back office system can send suggestions regarding additional or alternative services and/or devices for the desired service to the scanning device for communication to the user, as discussed above. At block 312, the back office system optionally transmits support materials (e.g. installation or troubleshooting information) to the scanning device based on the received device identification and secondary information. For example, in some embodiments, based on the device identification information and/or secondary information, the back office system transmits one or more of installation manuals, videos, photographs of the original installation of the device having the scanned matrix barcode, etc., as discussed above”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Wang, Chang with the multimedia data of Baykal.
One of ordinary skill in the art would have been motived to combine in order to provide recommendations for improved available services based on the capabilities of the device being installed and the location of the device.
(Baykal [par(s) 17] “the back office system 104 analyzes the associated secondary and device identification information to provide recommendations for improved or additional services. For example, if the device being installed is not recommended or compatible with the desired service, an alternative device can be suggested. Similarly, the back office system 104 can provide recommendations for additional available services based on the capabilities of the device being installed and the location of the device.”)
Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Raj et al. (Building a Digital Twin Network of SDN Using Knowledge Graphs) teaches passing message over the network.
Tang et al. (COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis) teaches instructional videos.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEHWAN KIM whose telephone number is (571)270-7409. The examiner can normally be reached Mon - Fri 9:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SEHWAN KIM/Examiner, Art Unit 2129