DETAILED ACTION
This office action is in response to the communication filed on June 08, 2026. Claims 21-25, 27-30, and 32-52 are currently pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed on June 08, 2026 have been fully considered but they are not persuasive for the following reasons:
Applicant in Pages 9-11 of the Remarks argues that Chi, Gopalakrishnan, and Arat do not teach or even suggest the amended features “receiving a user interaction”, "based on the user interaction, presenting an input engine configured to receive a user-based text or voice prompt", “receiving the user-based text or voice prompt at the input engine”, “selecting, by a first system, a first data domain comprising at least one of: data from a partial timeframe of the recording, text from a transcript of the recording, or a timestamp associated with the recording”, and “transmitting the user-based text or voice prompt and the first data domain to a second system”, as recited in amended independent claim 21 and similarly recited in amended independent claim 37.
Examiner respectfully disagrees. The cited prior art alone and/or in combination discloses the argued features.
In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video, here the graphical user interfaces present various input options that enable users to interact with an input engine in the graphical interface by clicking or selecting a chip comprising textual description to display an answer data, which is receiving a user interaction.
Therefore, Chi discloses “receiving a user interaction”.
Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video, the graphical user interfaces present various input options that enable users to interact with an input engine in the graphical interface by clicking or selecting a chip comprising textual description to display an answer data.
Chi in [0037], [0077], and [0078] discloses process textual content and a prompt to generate an output, input to machine-learned model(s) can be text or natural language data which can be processed to generate an output, input can be speech data which can be processed to generate an output.
Chi in [0080] discloses input can include audio data representing a spoken utterance and an output comprises a text output mapped to the spoken utterance.
Chi in [0041] and [0044] discloses given an input source video retrieving the video transcript, descriptions, and annotations annotating the frames with time-code to identify faces and text, generating prompts to an LLM for relevant information to the video, using the prompt to generate a set of short-length videos, which is receiving a text or voice prompt based on user interaction or a user-based text or voice prompt.
Therefore, Chi discloses “based on the user interaction, presenting an input engine configured to receive a user-based text or voice prompt”.
Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video, the graphical user interfaces present various input options that enable users to interact with an input engine in the graphical interface by clicking or selecting a chip comprising textual description to display an answer data.
Chi in [0022] and [0037] discloses prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output.
Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model.
Therefore, Chi discloses “receiving the user-based text or voice prompt at the input engine”.
Chi in [0022] and [0037] discloses extracting textual content such as transcripts, metadata, etc. from video content, prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output.
Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model.
Therefore, Chi discloses “selecting, by a first system, a first data…comprising at least one of: data from a partial timeframe of the recording, text from a transcript of the recording, or a timestamp associated with the recording”.
Chi does not explicitly disclose selecting a first data domain, but the Gopalakrishnan reference discloses the feature.
Chi in [0022] and [0037] discloses extracting textual content such as transcripts, metadata, etc. from video content, prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output.
Chi in [0037], [0077], and [0078] discloses process textual content and a prompt to generate an output, input to machine-learned model(s) can be text or natural language data which can be processed to generate an output, input can be speech data which can be processed to generate an output.
Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model.
Therefore, Chi discloses “transmitting the user-based text or voice prompt and the first data…to a second system having access to a machine learning model”.
Chi does not explicitly disclose transmitting the user-based text or voice prompt and the first data domain, but the Gopalakrishnan reference discloses the feature.
Chi discloses transmitting a prompt associated with various data to a system having access to a machine learning model and providing an output, however, Chi does not explicitly disclose:
selecting…a first data domain…;
transmitting the…prompt and the first data domain…;
Gopalakrishnan in [0017], [0018], and [0049] discloses domain specialty prompt instructions are generated and inserted into prompts to perform tasks, transcripts for different specialties labeled with corresponding specialty, performing text analysis tasks on different domain specialties or across multiple domains with multiple domain specialties, model store used to maintain different fine-tuned models for different text analysis tasks or domains.
Gopalakrishnan in [0023] discloses automatic speech recognition and natural language processing.
Gopalakrishnan in [0039] and [0044] discloses identifying domain general entities in input text by performing searches, replacing or modifying domain general entities with domain specialty identifier.
Therefore, Gopalakrishnan discloses “selecting a first data domain” and “transmitting a prompt and the first data domain”.
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having the teachings of Chi and Gopalakrishnan, to have combined Chi and Gopalakrishnan. The motivation to combine Chi and Gopalakrishnan would be to improve use of machine learning models to perform domain-specific text analysis tasks by augmenting a data set for tuning a pre-trained large language model using different domain specialties.
For the above reasons, Examiner states that rejection of the current Office action is proper.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 21-25, 27-30, and 32-52 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
At step 1:
Independent claims 21 and 37 respectively recite a method, one or more processors, and a system, which are directed to a statutory category such as a process, machine, or an article of manufacture.
At step 2A, prong one:
Independent claim 21 and similarly independent claim 37 recites the limitation:
“selecting…a first data domain comprising at least one of: data from a partial timeframe of the recording, text from a transcript of the recording, or a timestamp associated with the recording”;
A person can mentally or using a pen and paper select a first data domain comprising at least one of: data from a partial timeframe of the recording, text from a transcript of the recording, or a timestamp associated with the recording.
The limitation, as recited above, is a processes that, under its broadest reasonable interpretation, cover steps that can be performed in the human mind or by a human using a pen and paper, but for recitation of generic computer components.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
At step 2A, prong two:
This judicial exception is not integrated into a practical application.
Independent claim 21 and similarly independent claim 37 recites the limitations:
“displaying a recording on a display”, which is a step of displaying or outputting data. The step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity (MPEP 2106.05(g)).
“receiving a user interaction”, which is a step of receiving data. The step is recited at a high level of generality, and amounts to mere data gathering, which is a form of insignificant extra-solution activity (MPEP 2106.05(g)).
“based on the user interaction, presenting an input engine configured to receive a user-based text or voice prompt”, which is a step of presenting or outputting data. The step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity (MPEP 2106.05(g)).
“receiving the user-based text or voice prompt at the input engine”, which is a step of receiving data. The step is recited at a high level of generality, and amounts to mere data gathering, which is a form of insignificant extra-solution activity (MPEP 2106.05(g)).
“transmitting the user-based text or voice prompt and the first data domain to a second system having access to a machine learning model”, which is a step of transmitting data. The step is recited at a high level of generality, and amounts to mere data gathering, which is a form of insignificant extra-solution activity (MPEP 2106.05(g)).
“receiving answer data from the machine learning model”, which is a step of receiving data. The step is recited at a high level of generality, and amounts to mere data gathering, which is a form of insignificant extra-solution activity (MPEP 2106.05(g)).
“displaying the answer data in a visual format on the display”, which is a step of displaying or outputting data. The step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity (MPEP 2106.05(g)).
The additional elements “a method for prompting a machine learning model to generate answer data based on a recording, the method comprising:”, “on a display”, “an input engine”, “at the input engine”, “by a first system”, “to a second system having access to a machine learning model”, “from the machine learning model”, and “on the display” in the steps in claim 21 are recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
The additional elements “a non-transitory computer readable medium including instructions that are executable by one or more processors to perform operations comprising:”, “on a display”, “an input engine”, “at the input engine”, “by a first system”, “to a second system having access to a machine learning model”, “from the machine learning model”, and “on the display” in the steps in claim 37 are recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not
integrate the abstract idea into a practical application, even viewing the claims a whole,
because it does not impose any meaningful limits on practicing the abstract idea.
At step 2B:
Independent claims 21 and 37 recite the same additional elements as identified in step 2A prong two above. These additional elements are not sufficient to amount to significantly more than the judicial exception.
Independent claim 21 and similarly independent claim 37 recites the limitations:
“displaying a recording on a display”, which is a step of displaying or outputting data, and is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
“receiving a user interaction”, which is a step of receiving data, and is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i)).
“based on the user interaction, presenting an input engine configured to receive a user-based text or voice prompt”, which is a step of presenting or outputting data, and is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
“receiving the user-based text or voice prompt at the input engine”, which is a step of receiving data, and is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i)).
“transmitting the user-based text or voice prompt and the first data domain to a second system having access to a machine learning model”, which is a step of transmitting data, and is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i)).
“receiving answer data from the machine learning model”, which is a step of receiving data, and is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i)).
“displaying the answer data in a visual format on the display”, which is a step of displaying or outputting data, and is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
Accordingly, the additional limitations are not sufficient to amount to significantly more than the judicial exception. Therefore, the claims are directed to an abstract idea and are not patent eligible.
Dependent claim 22 and similarly dependent claim 38 recites additional limitations, such as:
“wherein the first data domain comprises the text from the transcript of the recording”.
These limitations are directed to the same abstract idea under the mental processes grouping as independent claim 21 and 37, because a person can mentally or using a pen and paper select a first data domain comprising text from a transcript of a recording, and because the limitations do not recite any additional elements that are sufficient to amount to significantly more.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 23 and similarly dependent claim 39 recites additional limitations, such as:
“wherein the recording comprises media having audio and visual components”, which is a step of displaying data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 24 and similarly dependent claim 40 recites additional limitations, such as:
“wherein the recording is displayed using at least one of a browser or a video hosting site”, which is a step of displaying data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
The additional elements “a browser” and “a video hosting site” in the step are recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 25 and similarly dependent claim 41 recites additional limitations, such as:
“wherein the user-based text or voice prompt comprises at least one of audio input or text input”, which is a step of receiving data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data gathering, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i)).
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 27 and similarly dependent claim 43 recites additional limitations, such as:
“displaying a user interface, wherein the user interaction is received at the user interface”, which is a step of displaying data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
The additional elements “a user interface” and “at the user interface” in the step are recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 28 recites additional limitations, such as:
“wherein user devices may only access the machine learning model through an application programming interface (API)”, which is a step reciting additional elements at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 29 and similarly dependent claim 44 recites additional limitations, such as:
“wherein the machine learning model is configured for generative artificial intelligence”, which is a step reciting additional elements at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 30 and similarly dependent claim 45 recites additional limitations, such as:
“wherein the machine learning model is a large language model (LLM)”, which is a step reciting additional elements at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 32 and similarly dependent claim 47 recites additional limitations, such as:
“wherein the second system has access to at least one second data domain with a data scope differing from the first data domain”, which is a step reciting additional elements at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 33 and similarly dependent claim 38 recites additional limitations, such as:
“wherein the at least one second data domain includes information available on the internet”, which is a step reciting additional elements at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 34 recites additional limitations, such as:
“wherein the answer data in the visual format is based on generating natural language corresponding to the answer data”, which is a step of displaying data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 35 and similarly dependent claim 49 recites additional limitations, such as:
“wherein the input engine is presented in response to receiving the user interaction”, which is a step of presenting or outputting data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 36 and similarly dependent claim 50 recites additional limitations, such as:
“wherein the first data domain comprises the data from a partial timeframe of the recording, the text from a transcript of the recording, and the timestamp associated with the recording”.
These limitations are directed to the same abstract idea under the mental processes grouping as independent claims 21 and 37, because a person can mentally or using a pen and paper select a first data domain comprising data from a partial timeframe of a recording, a text from a transcript of the recording, and a timestamp associated with the recording, and because the limitations do not recite any additional elements that are sufficient to amount to significantly more.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 42 recites additional limitations, such as:
“wherein the user interaction is received at a button”, which is a step of receiving data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data gathering, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i)).
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 46 recites additional limitations, such as:
“displaying a feedback interface for user feedback to transmit to the second system or a third system”, which is a step of displaying data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
The additional elements “a feedback interface”, “the second system”, and “a third system” in the step are recited at a high-level of generality, such that it amounts to no more than mere instructions to apply the exception using generic computer components.
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 51 recites additional limitations, such as:
“presenting the input engine on the display”, which is a step of presenting or outputting data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of presenting offers and gathering statistics (MPEP 2106.05(d)(II)(iv)).
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Dependent claim 52 recites additional limitations, such as:
“wherein the user-based text or voice prompt comprises a natural language input”, which is a step of receiving data.
At step 2A prong two, the step is recited at a high level of generality, and amounts to mere data gathering, which is a form of insignificant extra-solution activity.
At step 2B, the step is recognized as a well understood, routine, and conventional activity within the field of computer functions as an element of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i)).
Accordingly, the additional elements, individually or in combination, do not integrate the abstract idea into a practical application, even viewing the claims a whole, because it does not impose any meaningful limits on practicing the abstract idea.
Accordingly, dependent claims 22-25, 27-30, 32-36, and 38-52 are also directed to abstract idea without significantly more and are not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 21-25, 27-30, 32-45, and 47-52 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chi (US Pub 2025/0095690, provisional application filing date 09/14/23) in view of Gopalakrishnan (US Pub 2025/0029603).
With respect to claim 21, Chi discloses a method for prompting a machine learning model to generate answer data based on a recording, the method comprising:
displaying a recording on a display (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video);
receiving a user interaction (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video, here the graphical user interfaces present various input options that enable users to interact with an input engine in the graphical interface by clicking or selecting a chip comprising textual description to display an answer data);
based on the user interaction, presenting an input engine configured to receive a user-based text or voice prompt (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video, the graphical user interfaces present various input options that enable users to interact with an input engine in the graphical interface by clicking or selecting a chip comprising textual description to display an answer data; Chi in [0037], [0077], and [0078] discloses process textual content and a prompt to generate an output, input to machine-learned model(s) can be text or natural language data which can be processed to generate an output, input can be speech data which can be processed to generate an output; Chi in [0080] discloses input can include audio data representing a spoken utterance and an output comprises a text output mapped to the spoken utterance; Chi in [0041] and [0044] discloses given an input source video retrieving the video transcript, descriptions, and annotations annotating the frames with time-code to identify faces and text, generating prompts to an LLM for relevant information to the video, using the prompt to generate a set of short-length videos, which is receiving a text or voice prompt based on user interaction or a user-based text or voice prompt);
receiving the user-based text or voice prompt at the input engine (Chi in [0022] and [0037] discloses prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model);
selecting, by a first system, a first data…comprising at least one of: data from a partial timeframe of the recording, text from a transcript of the recording, or a timestamp associated with the recording (Chi in [0022] and [0037] discloses extracting textual content such as transcripts, metadata, etc. from video content, prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; here Chi does not explicitly disclose selecting a first data domain, but the Gopalakrishnan reference discloses the feature, as discussed below);
transmitting the user-based text or voice prompt and the first data…to a second system having access to a machine learning model (Chi in [0022] and [0037] discloses extracting textual content such as transcripts, metadata, etc. from video content, prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; here Chi does not explicitly disclose transmitting the user-based text or voice prompt and the first data domain, but the Gopalakrishnan reference discloses the feature, as discussed below);
receiving answer data from the machine learning model (Chi in [0022] and [0037] discloses prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output); and
displaying the answer data in a visual format on the display (Chi in [0022] and [0037] discloses prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output).
Chi discloses transmitting a prompt associated with various data to a system having access to a machine learning model, however, Chi does not explicitly disclose:
selecting…a first data domain…;
transmitting the…prompt and the first data domain…;
The Gopalakrishnan reference discloses selecting a first data domain and transmitting a prompt and the first data domain (Gopalakrishnan in [0017], [0018], and [0049] discloses domain specialty prompt instructions are generated and inserted into prompts to perform tasks, transcripts for different specialties labeled with corresponding specialty, performing text analysis tasks on different domain specialties or across multiple domains with multiple domain specialties, model store used to maintain different fine-tuned models for different text analysis tasks or domains; Gopalakrishnan in [0023] discloses automatic speech recognition and natural language processing; Gopalakrishnan in [0039] and [0044] discloses identifying domain general entities in input text by performing searches, replacing or modifying domain general entities with domain specialty identifier).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having the teachings of Chi and Gopalakrishnan, to have combined Chi and Gopalakrishnan. The motivation to combine Chi and Gopalakrishnan would be to improve use of machine learning models to perform domain-specific text analysis tasks by augmenting a data set for tuning a pre-trained large language model using different domain specialties (Gopalakrishnan: [0015] and [0017]).
With respect to claim 22, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the first data domain comprises the text from the transcript of the recording (Chi in [0018] and [0022] discloses extracting textual content associated with videos, such as a transcript of speech that occurs within the video, textual metadata, and other textual information associated with the video; Gopalakrishnan in [0012] discloses obtaining text generates from audio or video transcripts; Gopalakrishnan in [0025] and [0027] receiving a request to generate a transcript and summary of an audio conversation).
With respect to claim 23, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the recording comprises media having audio and visual components (Chi in [0039] discloses content including textual and/or visual content, content can be an audio and/or video content; Gopalakrishnan in [0027] and [0042] discloses receiving an audio file including metadata of a conversation, transcribing text from audio or video sources).
With respect to claim 24, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the recording is displayed using at least one of a browser or a video hosting site (Chi in [0019] and [0026] discloses presenting videos in an interactive player in a user interface; Chi in [0026] and [0030] and Figure 1A discloses the user interface providing side panels for support videos and questions and answers similar to a web document or website).
With respect to claim 25, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the user-based text or voice prompt comprises at least one of audio input or text input (Chi in [0037] discloses a prompt including an instruction to summarize one or more sets of textual content, to explain, to generate pairs of questions and answers, and/or other instructions; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model).
With respect to claim 27, Chi in view of Gopalakrishnan discloses the method of claim 21, further comprising displaying a user interface, wherein the user interaction is received at the user interface (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, user can hover over user-selectable chips in the interface that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video; Chi in [0033] discloses viewers or users can select an information button and the interface can generate a response).
With respect to claim 28, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein user devices may only access the machine learning model through an application programming interface (API) (Chi in [0084], [0086], and [0087] discloses application communicating with device components using an API, application communicates with central intelligence layer including a number of machine learning models; Gopalakrishnan in [0022] and [0045] discloses interface may be one or more graphical user interfaces that implement application programming interfaces (APIs), an API call using inserted domain specialty identifiers in a generated instruction invokes a host system for a pre-trained language model to perform text analysis).
With respect to claim 29, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the machine learning model is configured for generative artificial intelligence (Chi in [0019] and [0022] discloses processing textual content with a generative sequence processing model, such as a large language model, to generate an output, prompt a LLM to generate relevant content, such as summaries, explanations, and/or questions and answers; Chi in [0070] and [0087] discloses training machine learning models using various training or learning techniques, a central intelligence layer including a number of machine learning models; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0015] and [0017] discloses train machine learning models to accept and apply domain-specific information as part of input to perform text analysis tasks, tuning a pre-trained large language model using different domain specialties).
With respect to claim 30, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the machine learning model is a large language model (LLM) (Chi in [0019] and [0022] discloses processing textual content with a generative sequence processing model, such as a large language model, to generate an output, prompt a LLM to generate relevant content, such as summaries, explanations, and/or questions and answers; Chi in [0070] and [0087] discloses training machine learning models using various training or learning techniques, a central intelligence layer including a number of machine learning models; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0015] and [0017] discloses train machine learning models to accept and apply domain-specific information as part of input to perform text analysis tasks, tuning a pre-trained large language model using different domain specialties).
With respect to claim 32, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the second system has access to at least one second data domain with a data scope differing from the first data domain (Chi in [0031] and [0041] discloses retrieving external materials such as URL link to a web tutorial; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript or a URL link, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0018] and [0049] discloses performing text analysis tasks on different domain specialties or across multiple domains with multiple domain specialties, model store used to maintain different fine-tuned models for different text analysis tasks or domains; Gopalakrishnan in [0021] and [0026] discloses a provided network implementing natural language processing service that implements domain specialty instruction generation for performing text analysis tasks, a provider network may be a private or closed system or may be accessible via the internet, clients may convey network-based services requests).
With respect to claim 33, Chi in view of Gopalakrishnan discloses the method of claim 32, wherein the at least one second data domain includes information available on the internet (Chi in [0031] and [0041] discloses retrieving external materials such as URL link to a web tutorial; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript or a URL link, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0018] and [0049] discloses performing text analysis tasks on different domain specialties or across multiple domains with multiple domain specialties, model store used to maintain different fine-tuned models for different text analysis tasks or domains; Gopalakrishnan in [0021] and [0026] discloses a provided network implementing natural language processing service that implements domain specialty instruction generation for performing text analysis tasks, a provider network may be a private or closed system or may be accessible via the internet, clients may convey network-based services requests).
With respect to claim 34, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the answer data in the visual format is based on generating natural language corresponding to the answer data (Chi in [0019] and [0022] discloses processing textual content with a generative sequence processing model, such as a large language model, to generate an output, prompt a LLM to generate relevant content, such as summaries, explanations, and/or questions and answers; Chi in [0070] and [0087] discloses training machine learning models using various training or learning techniques, a central intelligence layer including a number of machine learning models; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0015] and [0017] discloses train machine learning models to accept and apply domain-specific information as part of input to perform text analysis tasks, tuning a pre-trained large language model using different domain specialties).
With respect to claim 35, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the input engine is presented in response to receiving the user interaction (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, user can hover over user-selectable chips in the interface that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video; Chi in [0033] discloses viewers or users can select an information button and the interface can generate a response).
With respect to claim 36, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the first data domain comprises the data from a partial timeframe of the recording, the text from a transcript of the recording, and the timestamp associated with the recording (Chi in [0022] and [0037] discloses extracting textual content such as transcripts, metadata, etc. from video content, prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model).
With respect to claim 37, Chi discloses a non-transitory computer readable medium including instructions that are executable by one or more processors to perform operations (Chi in [0059] discloses one or more non-transitory computer-readable storage media storing data and instructions executed by processor to perform operations) comprising:
displaying a recording on a display (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video);
receiving a user interaction (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video, here the graphical user interfaces present various input options that enable users to interact with an input engine in the graphical interface by clicking or selecting a chip comprising textual description to display an answer data);
based on the user interaction, presenting an input engine configured to receive a user-based text or voice prompt (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, presenting a video in the user interface and enabling access to a set of support videos including question-answer videos by hovering over user-selectable chips that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video, the graphical user interfaces present various input options that enable users to interact with an input engine in the graphical interface by clicking or selecting a chip comprising textual description to display an answer data; Chi in [0037], [0077], and [0078] discloses process textual content and a prompt to generate an output, input to machine-learned model(s) can be text or natural language data which can be processed to generate an output, input can be speech data which can be processed to generate an output; Chi in [0080] discloses input can include audio data representing a spoken utterance and an output comprises a text output mapped to the spoken utterance; Chi in [0041] and [0044] discloses given an input source video retrieving the video transcript, descriptions, and annotations annotating the frames with time-code to identify faces and text, generating prompts to an LLM for relevant information to the video, using the prompt to generate a set of short-length videos, which is receiving a text or voice prompt based on user interaction or a user-based text or voice prompt);
receiving the user-based text or voice prompt at the input engine (Chi in [0022] and [0037] discloses prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model);
selecting, by a first system, a first data…comprising at least one of: data from a partial timeframe of the recording, text from a transcript of the recording, or a timestamp associated with the recording (Chi in [0022] and [0037] discloses extracting textual content such as transcripts, metadata, etc. from video content, prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; here Chi does not explicitly disclose selecting a first data domain, but the Gopalakrishnan reference discloses the feature, as discussed below);
transmitting the user-based text or voice prompt and the first data…to a second system having access to a machine learning model (Chi in [0022] and [0037] discloses extracting textual content such as transcripts, metadata, etc. from video content, prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; here Chi does not explicitly disclose transmitting the user-based text or voice prompt and the first data domain, but the Gopalakrishnan reference discloses the feature, as discussed below);
receiving answer data from the machine learning model (Chi in [0022] and [0037] discloses prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output); and
displaying the answer data in a visual format on the display (Chi in [0022] and [0037] discloses prompt a LLM to generate additional relevant content, such as summaries, explanations, and/or questions and answers, process extracted textual content and a prompt with the LLM to generate an output).
Chi discloses transmitting a prompt associated with various data to a system having access to a machine learning model, however, Chi does not explicitly disclose:
selecting…a first data domain…;
transmitting the…prompt and the first data domain…;
The Gopalakrishnan reference discloses selecting a first data domain and transmitting a prompt and the first data domain (Gopalakrishnan in [0017], [0018], and [0049] discloses domain specialty prompt instructions are generated and inserted into prompts to perform tasks, transcripts for different specialties labeled with corresponding specialty, performing text analysis tasks on different domain specialties or across multiple domains with multiple domain specialties, model store used to maintain different fine-tuned models for different text analysis tasks or domains; Gopalakrishnan in [0023] discloses automatic speech recognition and natural language processing; Gopalakrishnan in [0039] and [0044] discloses identifying domain general entities in input text by performing searches, replacing or modifying domain general entities with domain specialty identifier).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having the teachings of Chi and Gopalakrishnan, to have combined Chi and Gopalakrishnan. The motivation to combine Chi and Gopalakrishnan would be to improve use of machine learning models to perform domain-specific text analysis tasks by augmenting a data set for tuning a pre-trained large language model using different domain specialties (Gopalakrishnan: [0015] and [0017]).
With respect to claim 38, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the first data domain comprises the text from the transcript of the recording (Chi in [0040] discloses associating support videos with one or more timestamps of source video, one or more timestamps correspond to one or more sets of textual content; Chi in [0043] discloses acquiring timecoded sentences, each with a start and end time mapped to the source video; Chi in [0044] discloses analyze video frames to annotate relevant information, annotations can be time-coded; Gopalakrishnan in [0012], [0017], and [0018] discloses generating text from audio or video transcripts and labeling with domain specialty, inserting domain specialty information into prompts to perform domain text analysis tasks; Gopalakrishnan in [0039] and [0044] discloses identifying domain general entities in input text, replacing or modifying domain general entities with domain specialty identifier.
With respect to claim 39, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the recording comprises media having audio and visual components (Chi in [0039] discloses content including textual and/or visual content, content can be an audio and/or video content; Gopalakrishnan in [0027] and [0042] discloses receiving an audio file including metadata of a conversation, transcribing text from audio or video sources).
With respect to claim 40, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the recording is displayed using at least one of a browser or a video hosting site (Chi in [0019] and [0026] discloses presenting videos in an interactive player in a user interface; Chi in [0026] and [0030] and Figure 1A discloses the user interface providing side panels for support videos and questions and answers similar to a web document or website).
With respect to claim 41, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the user-based text or voice prompt comprises at least one of audio input or text input (Chi in [0037] discloses a prompt including an instruction to summarize one or more sets of textual content, to explain, to generate pairs of questions and answers, and/or other instructions; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model).
With respect to claim 42, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the user interaction is received at a button (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, user can hover over user-selectable chips in the interface that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video; Chi in [0033] discloses viewers or users can select an information button and the interface can generate a response).
With respect to claim 43, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, the operations further comprising displaying a user interface, wherein the user interaction is received at the user interface (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, user can hover over user-selectable chips in the interface that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video; Chi in [0033] discloses viewers or users can select an information button and the interface can generate a response).
With respect to claim 44, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the machine learning model is configured for generative artificial intelligence (Chi in [0019] and [0022] discloses processing textual content with a generative sequence processing model, such as a large language model, to generate an output, prompt a LLM to generate relevant content, such as summaries, explanations, and/or questions and answers; Chi in [0070] and [0087] discloses training machine learning models using various training or learning techniques, a central intelligence layer including a number of machine learning models; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0015] and [0017] discloses train machine learning models to accept and apply domain-specific information as part of input to perform text analysis tasks, tuning a pre-trained large language model using different domain specialties).
With respect to claim 45, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the machine learning model is a large language model (LLM) (Chi in [0019] and [0022] discloses processing textual content with a generative sequence processing model, such as a large language model, to generate an output, prompt a LLM to generate relevant content, such as summaries, explanations, and/or questions and answers; Chi in [0070] and [0087] discloses training machine learning models using various training or learning techniques, a central intelligence layer including a number of machine learning models; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0015] and [0017] discloses train machine learning models to accept and apply domain-specific information as part of input to perform text analysis tasks, tuning a pre-trained large language model using different domain specialties).
With respect to claim 47, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the second system has access to at least one second data domain with a data scope differing from the first data domain (Chi in [0031] and [0041] discloses retrieving external materials such as URL link to a web tutorial; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript or a URL link, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; Gopalakrishnan in [0001] and [0013] discloses summaries created using a special class of machine learning models such as generative large language models that are tuned to follow natural language instructions describing any task; Gopalakrishnan in [0018] and [0049] discloses performing text analysis tasks on different domain specialties or across multiple domains with multiple domain specialties, model store used to maintain different fine-tuned models for different text analysis tasks or domains; Gopalakrishnan in [0021] and [0026] discloses a provided network implementing natural language processing service that implements domain specialty instruction generation for performing text analysis tasks, a provider network may be a private or closed system or may be accessible via the internet, clients may convey network-based services requests).
With respect to claim 48, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 47, wherein the at least one second data domain includes information available on the internet (Chi in [0031] and [0041] discloses retrieving external materials such as URL link to a web tutorial; Chi in [0049] and [0050] discloses each prompt containing a prefix to provide a clear ask to LLM, a target input such as a video transcript or a URL link, and a suffix to specify output format, prompts can be combined with video metadata and provided to an LLM or other model; Gopalakrishnan in [0018] and [0049] discloses performing text analysis tasks on different domain specialties or across multiple domains with multiple domain specialties, model store used to maintain different fine-tuned models for different text analysis tasks or domains; Gopalakrishnan in [0021] and [0026] discloses a provided network implementing natural language processing service that implements domain specialty instruction generation for performing text analysis tasks, a provider network may be a private or closed system or may be accessible via the internet, clients may convey network-based services requests).
With respect to claim 49, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the input engine is presented in response to receiving the user interaction (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces enabling users to interact, user can hover over user-selectable chips in the interface that correspond to a question, when user clicks on a chip the interface pauses the main video and plays a selected support video; Chi in [0033] discloses viewers or users can select an information button and the interface can generate a response).
With respect to claim 50, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, wherein the first data domain comprises the data from a partial timeframe of the recording, the text from a transcript of the recording, and the timestamp associated with the recording (Chi in [0040] discloses associating support videos with one or more timestamps of source video, one or more timestamps correspond to one or more sets of textual content; Chi in [0043] discloses acquiring timecoded sentences, each with a start and end time mapped to the source video; Chi in [0044] discloses analyze video frames to annotate relevant information, annotations can be time-coded; Gopalakrishnan in [0012], [0017], and [0018] discloses generating text from audio or video transcripts and labeling with domain specialty, inserting domain specialty information into prompts to perform domain text analysis tasks).
With respect to claim 51, Chi in view of Gopalakrishnan discloses the method of claim 21, further comprising presenting the input engine on the display (Chi in [0025]-[0027] and in Figures 1A-B discloses graphical user interfaces presenting various input options enabling users to interact by clicking or selecting a chip comprising textual description to display an answer data).
With respect to claim 52, Chi in view of Gopalakrishnan discloses the method of claim 21, wherein the user-based text or voice prompt comprises a natural language input (Chi in [0077] and [0078] discloses input to machine-learned model(s) can be text or natural language data which can be processed to generate an output, input can be speech data which can be processed to generate an output; Chi in [0080] discloses input can include audio data representing a spoken utterance and an output comprises a text output mapped to the spoken utterance; Gopalakrishnan in [0001] and [0040] discloses performing question answering in response to questions expressed in natural language).
Claim(s) 46 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chi (US Pub 2025/0095690, provisional application filing date 09/14/23) in view of Gopalakrishnan (US Pub 2025/0029603) and in further view of Arat (US Pub 2025/0126329, provisional application filing date 10/15/23).
With respect to claim 46, Chi in view of Gopalakrishnan discloses the non-transitory computer readable medium of claim 37, however, Chi and Gopalakrishnan do not explicitly disclose:
the operations further comprising displaying a feedback interface for user feedback to transmit to the second system or a third system.
The Arat reference discloses displaying a feedback interface for user feedback to transmit to a second system or a third system (Arat in [0013] and [0255] discloses posing a question while playing an interactive video content and receiving a response within a graphical user interface, collecting feedback from viewers on the quality and/or relevance of responses provided while watching an interactive video, using feedback as another metric to score and/or rank candidate responses in connection with determining a response to a question posed by a viewer during playback on the interactive video; Arat in [0120] and [0249] discloses generating responses to questions posed by viewers, asking a viewer to rate the quality of an answer provided to a posed question, gathering and using viewer feedback on responses to score and/or rank responses).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having the teachings of Chi, Gopalakrishnan, and Arat, to have combined Chi, Gopalakrishnan, and Arat. The motivation to combine Chi, Gopalakrishnan, and Arat would be to score and/rank responses to questions posed by viewers by asking the viewers to rate the quality of answers provided (Arat: [0120]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Remarks
The relevant prior art of record that are not used in claim rejections but are pertinent to the claims or disclosure are:
Wu (US Pub 2024/0362269), which discloses allowing a user to retrieve video sample or a time-indexed portion of a video sample using non-audio search prompt, such as a natural language text prompt.
McDaniel (US Pub 2024/0184812), which discloses a data retrieval system obtaining a natural language input or a text input converted from audio received from a user
Kirazci (US Pub 2018/0190274) discloses receiving a voice input by a user, converting the input into text, and generating a natural language prompt as a reply to the initial voice input.
Aher (US Pub 2023/0169112) discloses a search query being textual input and/or a selection by a user input device or a natural language voice prompt spoken by a user, which is detected as an input and processed to identify words included in the voice prompt.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to REZWANUL MAHMOOD whose telephone number is (571)272-5625. The examiner can normally be reached M-F 9-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ann J. Lo can be reached at 571-272-9767. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/R.M/Examiner, Art Unit 2159 /ANN J LO/Supervisory Patent Examiner, Art Unit 2159