DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement(s) (IDS) submitted on 15 July 2025 is/are being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 14 and 16-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 14, the modifications to the meaning of “performing sentiment analysis” in claim 7, which are incorporated in both the first and second optional embodiments of claim 14, lack clarity. Claim 7, from which claim 14 depends, recites "performing sentiment analysis using the prompt with a language model" where the language model is described in the specification as a modality-specific model (see [0021] which distinguishes between "large language models (LLMs) and/or multi-modal models (MMs)" and the "plurality of modality-specific ML models 510 and ML model 540 each comprises a language model"). Of note, this corresponds to how language models are known in the art, where language models are distinguished, for example, from vision-language models, and large language models are distinguished from large multimodal models. As such, the broadest reasonable interpretation of language model is a machine learning model which is trained to process an input text and output a classification or prediction based on the input text.
Claim 14 further recites "wherein performing the sentiment analysis comprises: either: performing separate sentiment analyses using a plurality of modality-specific machine learning (ML) models; and combining the separate sentiment analyses into the results of performing the sentiment analysis using a first ML model;” referred to as a first optional embodiment “or: performing the sentiment analysis using a second ML model trained for multi-modal sentiment analysis across two or more multi-modal signals simultaneously" referred to as a second optional embodiment.
Regarding the first optional embodiment, claim 7 describes performing the sentiment analysis using the language model. The relationship between the “sentiment analysis” performed in claim 7 and the “separate sentiment analyses” of claim 14 are unclear. Is the applicant asserting that the sentiment analysis is now a part of multiple sentiment analyses? In the alternative, is the modification subdividing the sentiment analysis of the language model into multiple sentiment analyses performed by a plurality of models? In either case, it is unclear what relationship, if any, the sentiment analysis performed with the language model has with the “separate sentiment analyses” using a plurality of modality-specific machine learning (ML) models. Further, the contents of the results are defined by the operation which created them (the “performing sentiment analysis” by the language model of claim 7). Does applicant envision the “results of performing the sentiment analysis” as now including each of the separate sentiment analyses, which would be provided in the first report to the presenter? Does the applicant envision the language model as being the first ML model? There are numerous possible alternative interpretations, none of which would be clear to a person having ordinary skill in the art.
Regarding the second optional embodiment, “performing the sentiment analysis” as described in claim 14 lacks clarity. Claim 7 recites “performing [the] sentiment analysis… with a language model”. Claim 14, in the second optional embodiment, describes “performing the sentiment analysis comprising…performing the sentiment analysis using a second ML model” As the sentiment analysis is already performed “with a language model” in claim 7, and the second ML model is not related in any way to the language model, it is unclear what operation is being performed by the second ML model. Therefore, claim 14 lacks clarity and is rejected.
Regarding claim 16, the phrase "computer storage device" lacks clarity. The preamble of claim 16 recites "A computer storage device having computer-executable instructions stored thereon, which, on execution by a computer, cause the computer to perform operations comprising…" which, on the surface, would appear to be equivalent language to "computer readable media" or "computer storage media," as used by those of ordinary skill in the art. Therefore, one would expect that the computer is executing instructions from a storage component, and the claim would be subject to computer readable media analysis to confirm that signals per se are not incorporated into the claim. However, the specification appears to conflict with this interpretation.
First and foremost, the application separately discloses "computer storage media" as distinct from the "computer storage device," which applicant describes as "implemented in hardware and exclud[ing] carrier waves and propagated signals" and "for purposes of this disclosure" computer storage media "are not signals per se." (Instant Application, [0075]). Given that computer storage device is not used/described with relation to the computer storage media, or vice versa, the "computer storage device" is understood as a distinct component from the "computer storage media".
The distinction between the two appears to be confirmed based on other portions of the specification. The phrase “computer storage device” occurs three (3) times in the specification. At paragraph [0063], the specification recites the language of the preamble in claim 16. At paragraph [0066], the specification explains that "FIG. 12 is a block diagram of an example computing device 1200 (e.g., a computer storage device) for implementing aspects disclosed herein, and is designated generally as computing device 1200." This language indicates that the computer storage device is an example of the computing device 1200. Paragraph [0069], though using the phrase "computer storage device" in the last sentence, devotes the entirety of the remaining discussion to associating "computer storage media" to the memory 1212. Though, unclear, the use of "computer storage device" at paragraph [0069], as read in the context of paragraphs [0069] and [0070], appears to be a typographical error for the intended phrase "computer storage media".
Therefore, in the context of the entire disclosure, the broadest reasonable interpretation of the phrase "computer storage device" is a computing device 1200. However, this interpretation creates a lack of clarity regarding the remaining parts of claim 16. If computer storage device is understood as a computing device 1200, the operations occurring in claim 16 are unclear in that the claim recites two computers and only one of them appears to be performing a function. If the computer storage device is to be understood as a storage device akin to a computer readable media, applicant is advised that the phrase "computer storage device" is not the same component as "computer readable media," which is defined in paragraph [0075] as excluding transitory signals. Therefore, claim 16 lacks clarity and is rejected.
Examiner notes that "computer storage device" under the second possible interpretation covers both transitory and non-transitory signals, in violation of 35 USC 101. However, since the ambiguity prevents a determination of whether the incorporated definition of “computer readable media” is applicable, the rejection under 35 USC 101 is premature at this point.
Regarding claim 19, the limitation “performing the sentiment analysis” lacks clarity. Claim 16 recites “performing [the] sentiment analysis… with a language model”. Claim 19, which depends from claim 16, describes “performing the sentiment analysis comprising…performing the sentiment analysis using a second ML model.” As the sentiment analysis is already performed “with a language model” in claim 16, and the second ML model is not related in any way to the language model, it is unclear what operation is being performed by the second ML model. Therefore, claim 19 lacks clarity and is rejected.
Regarding claims 17-20, claims 17-20 depend from claim 16 and incorporate all limitations therefrom. Therefore, claims 17-20 are rejected for at least the same reasons as described with relation to claim 16.
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claims 14 and 19 are rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends.
Regarding claim 14, the first and second optional embodiments fail to further limit claim 7. Claim 7 describes the sentiment analysis as being performed using a prompt with the language model, which is understood as describing the entire and singular sentiment analysis is performed "using the prompt with a language model." Examiner notes that “a sentiment analysis” refers to a single component. Further, the “sentiment analysis”, as an action, is not “perform[ed]” comprising or including the language model, but "using the prompt with a language model." Thus, claim 7 creates a closed group with respect to performance of the “sentiment analysis”, indicating that “using the prompt with a language model” constitutes “performing” the process of "sentiment analysis." In light of the above, the “performing” of the process of “sentiment analysis” provides the antecedent basis for "results of performing the sentiment analysis" as performance of a process is understood to produce the results of said process.
In the first optional embodiment, as best as can be understood by the Office in light of the lack of clarity issues described above, applicant attempts to redefine the singular sentiment analysis as a plurality of sentiment analyses, and then attempts to separate the "performing sentiment analysis" (performance of the process) from the "results of performing the sentiment analysis" (the results of the performance of said process). As best as can be understood by the Office, applicant is including the language model in the "plurality of modality-specific machine learning (ML) models", resulting in the "sentiment analysis" of claim 7 becoming merely a member of the "separate sentiment analyses." In turn, applicant redefines the "results of performing the sentiment analysis" as being derived from "a first ML model," based on a combination of the “separate sentiment analyses.” In support of this understanding, the first ML model of claim 14 is a different claim part than the language model of claim 7, as they do not share antecedent basis or a descriptive connection. In claim 7, "performing sentiment analysis" and "results of performing the sentiment analysis" both are results of the "using the prompt with a language model". In claim 14, the sentiment analysis appears to now be subdivided into multiple parts, such that only part of the sentiment analysis is performed "using the prompt with a language model", and "results of performing the sentiment analysis" corresponds to a new aggregated "sentiment analysis", which results from "combining the separate sentiment analyses…using a first ML model". By changing the relationship between "performing sentiment analysis" and "results of performing the sentiment analysis" such that the "results of performing the sentiment analysis" is no longer the results of the "performing sentiment analysis," and by changing the relationship of the language model to said sentiment analysis as defined in claim 7, the first optional embodiment of claim 14 fails to further limit claim 7.
In the second optional embodiment, the language model in the sentiment analysis appears to be replaced entirely by "a second ML model trained for multi-modal sentiment analysis". The same step of "performing the sentiment analysis" which is performed "with the language model" in claim 7, is now performed by a separate and mutually exclusive model, the "second ML model trained for multi-modal sentiment analysis". To clarify, in claim 7, the sentiment analysis is performed using a prompt with the language model and that performance produces the "results of performing the sentiment analysis," as explained above. The language model is understood in the art and in the context of the instant application as a modality-specific model. Though presented as "comprising" either of the options in claim 14, the "sentiment analysis…" as originally described in claim 7 is "perform[ed]", not "perform[ed]" comprising, "using the prompt with a language model." Claim 14 now describes the performing the sentiment analysis [using the prompt with the language model] as comprising "performing the sentiment analysis" using a second ML model for multi-modal sentiment analysis, which creates an open group (e.g., “performing… comprising”) from the closed group (e.g., “performing… with”). Further, even considering the comprising language, the same claim part “the sentiment analysis” remains redefined as now being “perform[ed] …using a second ML model.” By changing the model performing the sentiment analysis from the language model to either “a group comprising a language model” or the “second ML model,” the second optional embodiment of claim 14 fails to further limit claim 7. Therefore, claim 14 is further rejected under 112(d).
Regarding claim 19, the language model in the sentiment analysis appears to be replaced entirely by "a second ML model trained for multi-modal sentiment analysis". The same step of "performing the sentiment analysis" which is performed "with the language model" in claim 16, is now performed by a separate and mutually exclusive model, the "second ML model trained for multi-modal sentiment analysis". To clarify, in claim 16, the sentiment analysis is performed using a prompt with the language model and that performance produces the "results of performing the sentiment analysis." The language model, as explained with relation to claim 7, is understood in the art and in the context of the instant application as a modality-specific model. Though presented as "comprising" either of the options in claim 14, the "sentiment analysis…" as originally described in claim 7 is "perform[ed]", not "perform[ed]" comprising, "using the prompt with a language model." Claim 14 now describes the performing the sentiment analysis [using the prompt with the language model] as comprising "performing the sentiment analysis" using a second ML model for multi-modal sentiment analysis, which creates an open group (e.g., “performing… comprising”) from the closed group (e.g., “performing… with”). Further, even considering the comprising language, the same claim part “the sentiment analysis” remains redefined as now being “perform[ed] …using a second ML model.” By changing the model performing the sentiment analysis from the language model to either “a group comprising a language model” or the “second ML model,” the second optional embodiment of claim 14 fails to further limit claim 7. Therefore, claim 14 is further rejected under 112(d).
Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-2, 5, 7-8, 10, 13, and 16-17 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Rivera-Rodriguez (U.S. Pat. No. 12,063,123, hereinafter Rivera-Rodriguez).
Regarding claim 1, Rivera-Rodriguez discloses A system comprising: a processor; and a computer-readable medium storing instructions that are operative upon execution by the processor to (Systems and methods described with reference to the "media processing service 200" as implemented through "various memories... and/or storage unit 1136" which "may store one or more sets of instructions and data structures (e.g., software)" that "when executed by processor(s) 1110, cause various operations to implement the disclosed embodiments."; Rivera-Rodriguez, ¶ Col. 8, lines 19-38; Col. 21, lines 21-28): capture a plurality of multi-modal signals from a first multi-participant interaction session, ("a media processing service 200 of an online meeting service receives and processes various video streams generated during an online meeting"; Rivera-Rodriguez, ¶ Col. 8, lines 19-38) wherein the captured plurality of multi-modal signals comprises an audio feed, a video feed, and image stills of participants in the first multi-participant interaction session ("A first type of video stream is a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service. Generally, this first type of video stream will depict a video image of a meeting participant," where the video streams comprising video data, audio data ("the audio portion of each video stream of the first type"), and image stills (video data in video streams are understood as a time-coordinated sequence of still frames. Video data, audio data, and image stills {multi-modal signals} as present in the "first type of video stream" are of the participants in the first multi-participant (e.g., shown in Fig. 2 as meeting participants 1-4) interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-54; Col. 9, lines 1-24; FIG. 2); correlate timing information across the captured plurality of multi-modal signals ("each video stream... of a meeting participant is received and processed by a first type of pre-trained machine learning model that converts the audio portion of the video stream, representing the spoken message of a meeting participant, to text" and "each video stream is also processed by one or more additional pre-trained machine learning models—specifically, one or more computer vision models" where, in some embodiments, "the output of these models is text that describes or otherwise represents the detected facial expression, emotion, or body language, along with timing data (e.g., a timestamp, or a combination of a beginning time and an ending time)" for each verbal and non-verbal communication (see FIG. 7).; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; FIG. 7); generate a prompt using the captured plurality of multi-modal signals and the correlated timing information, including the audio feed and the image stills ("the meeting analyzer service uses the text-based digital representation of the online meeting" which corresponds to the output from the pre-trained machine learning models above "with a generative language model, such as a large language model (“LLM”), to generate accurate and complete summary descriptions of the online meeting"; Rivera-Rodriguez, ¶ Col. 7, lines 14-34); perform sentiment analysis using the prompt with a language model (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt {using a prompt} that is provided as input to the generative language model {...with a generative language model}," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46); and provide a first report to a presenter indicating results of performing the sentiment analysis (As described with reference to an example "a scenario where a first meeting participant is sharing content via the online meeting collaboration tool (e.g., a screen or app sharing feature)" where the "first meeting participant may ask a question related to the content...[and] one or more other meeting participants may communicate agreement, or disagreement, with the presenter (e.g., the first meeting participant) by making a non-verbal communication, such as a hand or head gesture," and "the verbal and non-verbal communications are captured and represented in the digital representation of the online meeting, the meeting analyzer service can generate answers to various questions that are accurate, based on the textual descriptions and timestamps associated with non-verbal communications" and "output from the model 902" is then "presented to the end-user" {presenter}.; Rivera-Rodriguez, ¶ Col. 16, lines 17-32; Col. 16, line 51-Col. 17, line 20)
Regarding claim 2, Rivera-Rodriguez discloses wherein the instructions are further operative to: generate a timestamped transcript using the audio feed, ("the audio portion of each video stream of the first type is processed using a speech-to-text algorithm or model to derive text representing the spoken message from each meeting participant."; Rivera-Rodriguez, ¶ Col. 9, lines -24) wherein the captured plurality of multi-modal signals further comprises the timestamped transcript, ("The information that precedes the actual text as shown in the transcript 700 includes information indicating the time at which the spoken message was recorded, an identifier for the meeting participant, and a name of the meeting participant."; Rivera-Rodriguez, ¶ Col. 13, lines 25-34) wherein performing sentiment analysis comprises performing sentiment analysis using the timestamped transcript ("the various outputs—for example, annotated textual elements—of the video processor 806" which includes the speech to text component 808 "and the content share processor 816 are provided as inputs to a serializer 826, which temporarily stores the data before generating the final output in the form of an annotated, text-based transcript 828" and "the text-based transcript 828 of an online meeting may be used in generating a prompt, as part of the pre-processing stage."; Rivera-Rodriguez, ¶ Col. 15, lines 53-67; Col. 16, lines 17-38), and wherein correlating the timing information across the captured plurality of multi-modal signals includes correlating timestamps of the timestamped transcript with the timing information of another multi-modal signal of the captured plurality of multi-modal signals ("an annotated text-based transcript may be generated, where the text-based transcript includes all of the textual elements derived from the various sources, arranged chronologically to reflect the time during the meeting at which an event or act occurred, and from which a portion of text was derived," where each text element of the various outputs can include "a first element for the text itself, a second element for a timestamp, and a third element to identify the source of the text," thus the text elements derived from verbal and non-verbal communications are arranged chronologically based on timestamps for each of the verbal components {timestamps of a timestamped transcript} and for the nonverbal components {the timing information of another multi-modal signal of the captured plurality of multi-modal signals}. See also FIG. 7, depicting an example of a transcript which includes the chronologically organized text elements.; Rivera-Rodriguez, ¶ Col. 6, lines 33-67; Col. 13, lines 57-64; FIG. 7).
Regarding claim 5, Rivera-Rodriguez discloses wherein the first multi-participant interaction session comprises a live video teleconference or a previously recorded video teleconference (In the disclosed example, "all four meeting participants are broadcasting a live video stream" which can also be a "full video recording of the online meeting"; Rivera-Rodriguez, ¶ Col. 4, lines 1-10; Col. 10, lines 4-24); and wherein the captured plurality of multi-modal signals further comprises at least one signal selected from the list consisting of: a chat, participant actions, and displayed media (the captured plurality of multi-modal signals can include "a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service" which is participant actions, "a video stream that represents content that a meeting participant is sharing with other meeting participants via the online meeting service," which is displayed media, and "text extracted from a text-based chat that occurred during the online meeting {a chat}"; Rivera-Rodriguez, ¶ Col. 6, lines 33-67; Col. 8, lines 19-43).
Regarding claim 7, Rivera-Rodriguez discloses A computer-implemented method comprising (Systems and methods described with reference to the "media processing service 200" as implemented through "various memories... and/or storage unit 1136" which "may store one or more sets of instructions and data structures (e.g., software)" that "when executed by processor(s) 1110, cause various operations to implement the disclosed embodiments."; Rivera-Rodriguez, ¶ Col. 8, lines 19-38; Col. 21, lines 21-28): capturing a plurality of multi-modal signals from a first multi-participant interaction session, ("a media processing service 200 of an online meeting service receives and processes various video streams generated during an online meeting"; Rivera-Rodriguez, ¶ Col. 8, lines 19-38) wherein the captured plurality of multi-modal signals comprises an audio feed, a video feed, and image stills of participants in the first multi-participant interaction session ("A first type of video stream is a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service. Generally, this first type of video stream will depict a video image of a meeting participant," where the video streams comprising video data, audio data ("the audio portion of each video stream of the first type"), and image stills (video data in video streams are understood as a time-coordinated sequence of still frames. Video data, audio data, and image stills {multi-modal signals} as present in the "first type of video stream" are of the participants in the first multi-participant (e.g., shown in Fig. 2 as meeting participants 1-4) interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-54; Col. 9, lines 1-24; FIG. 2); correlating timing information across the captured plurality of multi-modal signals ("each video stream... of a meeting participant is received and processed by a first type of pre-trained machine learning model that converts the audio portion of the video stream, representing the spoken message of a meeting participant, to text" and "each video stream is also processed by one or more additional pre-trained machine learning models—specifically, one or more computer vision models" where, in some embodiments, "the output of these models is text that describes or otherwise represents the detected facial expression, emotion, or body language, along with timing data (e.g., a timestamp, or a combination of a beginning time and an ending time)" for each verbal and non-verbal communication (see FIG. 7).; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; FIG. 7); generating a prompt using the captured plurality of multi-modal signals and the correlated timing information, including the audio feed and the image stills ("the meeting analyzer service uses the text-based digital representation of the online meeting" which corresponds to the output from the pre-trained machine learning models above "with a generative language model, such as a large language model (“LLM”), to generate accurate and complete summary descriptions of the online meeting"; Rivera-Rodriguez, ¶ Col. 7, lines 14-34); performing sentiment analysis using the prompt with a language model (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt {using a prompt} that is provided as input to the generative language model {...with a generative language model}," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46); and providing a first report to a presenter indicating results of performing the sentiment analysis (As described with reference to an example "a scenario where a first meeting participant is sharing content via the online meeting collaboration tool (e.g., a screen or app sharing feature)" where the "first meeting participant may ask a question related to the content...[and] one or more other meeting participants may communicate agreement, or disagreement, with the presenter (e.g., the first meeting participant) by making a non-verbal communication, such as a hand or head gesture," and "the verbal and non-verbal communications are captured and represented in the digital representation of the online meeting, the meeting analyzer service can generate answers to various questions that are accurate, based on the textual descriptions and timestamps associated with non-verbal communications" and "output from the model 902" is then "presented to the end-user" {presenter}.; Rivera-Rodriguez, ¶ Col. 16, lines 17-32; Col. 16, line 51-Col. 17, line 20).
Regarding claim 8, Rivera-Rodriguez discloses further comprising: generating a timestamped transcript using the audio feed, ("the audio portion of each video stream of the first type is processed using a speech-to-text algorithm or model to derive text representing the spoken message from each meeting participant."; Rivera-Rodriguez, ¶ Col. 9, lines -24) wherein the captured plurality of multi-modal signals further comprises the timestamped transcript, ("The information that precedes the actual text as shown in the transcript 700 includes information indicating the time at which the spoken message was recorded, an identifier for the meeting participant, and a name of the meeting participant."; Rivera-Rodriguez, ¶ Col. 13, lines 25-34) wherein performing sentiment analysis comprises performing sentiment analysis using the timestamped transcript ("the various outputs—for example, annotated textual elements—of the video processor 806" which includes the speech to text component 808 "and the content share processor 816 are provided as inputs to a serializer 826, which temporarily stores the data before generating the final output in the form of an annotated, text-based transcript 828" and "the text-based transcript 828 of an online meeting may be used in generating a prompt, as part of the pre-processing stage."; Rivera-Rodriguez, ¶ Col. 15, lines 53-67; Col. 16, lines 17-38), and wherein correlating the timing information across the captured plurality of multi-modal signals includes correlating timestamps of the timestamped transcript with the timing information of another multi-modal signal of the captured plurality of multi-modal signals ("an annotated text-based transcript may be generated, where the text-based transcript includes all of the textual elements derived from the various sources, arranged chronologically to reflect the time during the meeting at which an event or act occurred, and from which a portion of text was derived," where each text element of the various outputs can include "a first element for the text itself, a second element for a timestamp, and a third element to identify the source of the text," thus the text elements derived from verbal and non-verbal communications are arranged chronologically based on timestamps for each of the verbal components {timestamps of a timestamped transcript} and for the nonverbal components {the timing information of another multi-modal signal of the captured plurality of multi-modal signals}. See also FIG. 7, depicting an example of a transcript which includes the chronologically organized text elements.; Rivera-Rodriguez, ¶ Col. 6, lines 33-67; Col. 13, lines 57-64; FIG. 7).
Regarding claim 10, Rivera-Rodriguez discloses wherein providing the first report to the presenter comprises: providing the first report to the presenter in near real time during the first multi-participant interaction session; or providing the first report to the presenter after conclusion of the first multi-participant interaction session (As described with reference to an example "a scenario where a first meeting participant is sharing content via the online meeting collaboration tool (e.g., a screen or app sharing feature)" where the "first meeting participant may ask a question related to the content...[and] one or more other meeting participants may communicate agreement, or disagreement, with the presenter (e.g., the first meeting participant) by making a non-verbal communication, such as a hand or head gesture," and "the verbal and non-verbal communications are captured and represented in the digital representation of the online meeting, the meeting analyzer service can generate answers to various questions that are accurate, based on the textual descriptions and timestamps associated with non-verbal communications" and "output from the model 902" is then "presented to the end-user" {presenter}, at least, after the meeting.; Rivera-Rodriguez, ¶ Col. 16, lines 17-32; Col. 16, line 51-Col. 17, line 20).
Regarding claim 13, Rivera-Rodriguez discloses wherein the first multi-participant interaction session comprises a live video teleconference or a previously recorded video teleconference (In the disclosed example, "all four meeting participants are broadcasting a live video stream" which can also be a "full video recording of the online meeting"; Rivera-Rodriguez, ¶ Col. 4, lines 1-10; Col. 10, lines 4-24); and wherein the captured plurality of multi-modal signals further comprises at least one signal selected from the list consisting of: a chat, participant actions, and displayed media (the captured plurality of multi-modal signals can include "a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service" which is participant actions, "a video stream that represents content that a meeting participant is sharing with other meeting participants via the online meeting service," which is displayed media, and "text extracted from a text-based chat that occurred during the online meeting {a chat}"; Rivera-Rodriguez, ¶ Col. 6, lines 33-67; Col. 8, lines 19-43).
Regarding claim 16, Rivera-Rodriguez discloses A computer storage device having computer-executable instructions stored thereon, which, on execution by a computer, cause the computer to perform operations comprising (Systems and methods described with reference to the "media processing service 200" as implemented through "various memories... and/or storage unit 1136" which "may store one or more sets of instructions and data structures (e.g., software)" that "when executed by processor(s) 1110, cause various operations to implement the disclosed embodiments."; Rivera-Rodriguez, ¶ Col. 8, lines 19-38; Col. 21, lines 21-28): capturing a plurality of multi-modal signals from a first multi-participant interaction session, ("a media processing service 200 of an online meeting service receives and processes various video streams generated during an online meeting"; Rivera-Rodriguez, ¶ Col. 8, lines 19-38) wherein the captured plurality of multi-modal signals comprises an audio feed, a video feed, and image stills of participants in the first multi-participant interaction session ("A first type of video stream is a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service. Generally, this first type of video stream will depict a video image of a meeting participant," where the video streams comprising video data, audio data ("the audio portion of each video stream of the first type"), and image stills (video data in video streams are understood as a time-coordinated sequence of still frames. Video data, audio data, and image stills {multi-modal signals} as present in the "first type of video stream" are of the participants in the first multi-participant (e.g., shown in Fig. 2 as meeting participants 1-4) interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-54; Col. 9, lines 1-24; FIG. 2); correlating timing information across the captured plurality of multi-modal signals ("each video stream... of a meeting participant is received and processed by a first type of pre-trained machine learning model that converts the audio portion of the video stream, representing the spoken message of a meeting participant, to text" and "each video stream is also processed by one or more additional pre-trained machine learning models—specifically, one or more computer vision models" where, in some embodiments, "the output of these models is text that describes or otherwise represents the detected facial expression, emotion, or body language, along with timing data (e.g., a timestamp, or a combination of a beginning time and an ending time)" for each verbal and non-verbal communication (see FIG. 7).; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; FIG. 7); generating a prompt using the captured plurality of multi-modal signals and the correlated timing information, including the audio feed and the image stills ("the meeting analyzer service uses the text-based digital representation of the online meeting" which corresponds to the output from the pre-trained machine learning models above "with a generative language model, such as a large language model (“LLM”), to generate accurate and complete summary descriptions of the online meeting"; Rivera-Rodriguez, ¶ Col. 7, lines 14-34); performing sentiment analysis using the prompt with a language model (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt {using a prompt} that is provided as input to the generative language model {...with a generative language model}," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46); and providing a first report to a presenter indicating results of performing the sentiment analysis (As described with reference to an example "a scenario where a first meeting participant is sharing content via the online meeting collaboration tool (e.g., a screen or app sharing feature)" where the "first meeting participant may ask a question related to the content...[and] one or more other meeting participants may communicate agreement, or disagreement, with the presenter (e.g., the first meeting participant) by making a non-verbal communication, such as a hand or head gesture," and "the verbal and non-verbal communications are captured and represented in the digital representation of the online meeting, the meeting analyzer service can generate answers to various questions that are accurate, based on the textual descriptions and timestamps associated with non-verbal communications" and "output from the model 902" is then "presented to the end-user" {presenter}.; Rivera-Rodriguez, ¶ Col. 16, lines 17-32; Col. 16, line 51-Col. 17, line 20).
Regarding claim 17, Rivera-Rodriguez discloses wherein the operations further comprise: generating a timestamped transcript using the audio feed, ("the audio portion of each video stream of the first type is processed using a speech-to-text algorithm or model to derive text representing the spoken message from each meeting participant."; Rivera-Rodriguez, ¶ Col. 9, lines -24) wherein the captured plurality of multi-modal signals further comprises the timestamped transcript, ("The information that precedes the actual text as shown in the transcript 700 includes information indicating the time at which the spoken message was recorded, an identifier for the meeting participant, and a name of the meeting participant."; Rivera-Rodriguez, ¶ Col. 13, lines 25-34) wherein performing sentiment analysis comprises performing sentiment analysis using the timestamped transcript ("the various outputs—for example, annotated textual elements—of the video processor 806" which includes the speech to text component 808 "and the content share processor 816 are provided as inputs to a serializer 826, which temporarily stores the data before generating the final output in the form of an annotated, text-based transcript 828" and "the text-based transcript 828 of an online meeting may be used in generating a prompt, as part of the pre-processing stage."; Rivera-Rodriguez, ¶ Col. 15, lines 53-67; Col. 16, lines 17-38), and wherein correlating the timing information across the captured plurality of multi-modal signals includes correlating timestamps of the timestamped transcript with the timing information of another multi-modal signal of the captured plurality of multi-modal signals ("an annotated text-based transcript may be generated, where the text-based transcript includes all of the textual elements derived from the various sources, arranged chronologically to reflect the time during the meeting at which an event or act occurred, and from which a portion of text was derived," where each text element of the various outputs can include "a first element for the text itself, a second element for a timestamp, and a third element to identify the source of the text," thus the text elements derived from verbal and non-verbal communications are arranged chronologically based on timestamps for each of the verbal components {timestamps of a timestamped transcript} and for the nonverbal components {the timing information of another multi-modal signal of the captured plurality of multi-modal signals}. See also FIG. 7, depicting an example of a transcript which includes the chronologically organized text elements.; Rivera-Rodriguez, ¶ Col. 6, lines 33-67; Col. 13, lines 57-64; FIG. 7).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3 and 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Rivera-Rodriguez as applied to claim 1 and 8 above, and further in view of Cunico (U.S. PG Pub. No. 2017/0177928, hereinafter Cunico).
Regarding claim 3, the rejection of claim 1 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. However, Rivera-Rodriguez fail(s) to expressly recite wherein providing the first report to the presenter comprises: providing the first report to the presenter in near real time during the first multi-participant interaction session.
Cunico teaches systems and methods of “analyzing sentiment of attendees in a video conference.” (Cunico, ¶ [0001]). Regarding claim 3, Cunico teaches wherein providing the first report to the presenter comprises: providing the first report to the presenter in near real time during the first multi-participant interaction session ("In addition to using facial recognition of an attendee’s facial expressions and facial movement, sentiment analysis program 160 further refines the sentiment analysis using NLP and sentiment analysis of the attendee’s spoken words in the video conference discussions. Sentiment analysis program 160 may determine an aggregate, real-time sentiment analysis representing the average or group sentiment of all of the attendees in the video conference analysis using sentiment analysis program 160" where "a moderator viewing the video conference and results from sentiment analysis program 160 on moderator display 125 may include a moderator, one or more other identified presenters, or one or more reviewers of the video conference or the sentiment analysis results."; Cunico, ¶ [0017], [0025]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Cunico to include wherein providing the first report to the presenter comprises: providing the first report to the presenter in near real time during the first multi-participant interaction session. The sentiment analysis described in Cunico allows for real-time analysis which can be focused on individual attendees in a video conference, which provides the known benefit of considering the individual differences in attendees while providing time sensitive analysis, allowing a more appropriate and more timely response from the presenter to address detected problems and/or to improve meeting outcomes generally, as recognized in light of Cunico. (Cunico, ¶ [0018], [0029]).
Regarding claim 9, the rejection of claim 8 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. Rivera-Rodriguez further discloses further comprising: performing participant-specific sentiment analysis for a selected participant, ("One or more prompts may be constructed to identify people who communicated specific knowledge or subject matter, and so forth."; Rivera-Rodriguez, ¶ Col. 16, lines 39-50). However, Rivera-Rodriguez fail(s) to expressly recite wherein the first report further comprises results of performing the participant-specific sentiment analysis attributed to the selected participant.
The relevance of Cunico is described above with relation to claim 3. Regarding claim 9, Cunico teaches wherein the first report further comprises results of performing the participant-specific sentiment analysis attributed to the selected participant ("sentiment analysis program 160 has the capability to learn characteristic expressions of the attendee in video conferences" and "sentiment analysis program 160" can provide "sentiment analysis for an individual attendee" such as based on historical or personal characteristics. In the context of the selected participant of Rivera-Rodriguez, is a participant specific sentiment analysis. As all sentiment analyses are included in the report generated in Rivera-Rodriguez, the sentiment analysis report includes a participant specific sentiment analysis.; Cunico, ¶ [0029]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Cunico to include wherein the first report further comprises results of performing the participant-specific sentiment analysis attributed to the selected participant. The sentiment analysis described in Cunico allows for real-time analysis which can be focused on individual attendees in a video conference, which provides the known benefit of considering the individual differences in attendees while providing time sensitive analysis, allowing a more appropriate and more timely response from the presenter to address detected problems and/or to improve meeting outcomes generally, as recognized in light of Cunico. (Cunico, ¶ [0018], [0029]).
Claims 4, 11-12, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Rivera-Rodriguez as applied to claims 1, 7, and 16 above, and further in view of Chau (U.S. PG. Pub. No. 2025/0315198, hereinafter Chau).
Regarding claim 4, the rejection of claim 1 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. Rivera-Rodriguez further discloses wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session (Discloses "one or more pre-trained machine learning models may be used to detect or identify emotions" and, in one example, the presented "may ask a question related to the content that is being shared, while one or more other meeting participants may communicate agreement {a positive sentiment}, or disagreement {a negative sentiment}, with the presenter (e.g., the first meeting participant) by making a non-verbal communication, such as a hand or head gesture," where the identified emotion as related to communicated agreement or disagreement, is sentiment analysis indicating a positive or negative sentiment. Further "Each video stream is independently analyzed using various pre-trained machine learning models of the media processing service to generate text {for each of the separately-analyzed portions of the first multi-participant interaction session}."; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 16, line 51 - col. 17, line 20). However, Rivera-Rodriguez fails to expressly recite wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session and wherein the first report provides suggestions based on at least the positive and/or negative sentiment.
Chau teaches systems and methods for “automatically generating feedback about content shared during a videoconference.” (Chau, ¶ [0002]). Regarding claim 4, Chau teaches wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session ("The video data may depict the participants in the videoconference" which can be "analyzed using a trained machine-learning model to identify sentiment queues corresponding to different portions of the visual content" and "each participant’s sentiment may change throughout the presentation of the visual content" where "some portions (e.g., slides) of the visual content may elicit a positive sentiment while other portions may elicit a negative sentiment"; Chau, ¶ [0018]), and wherein the first report provides suggestions based on at least the positive and/or negative sentiment ("The system can identify these sentiment queues and generate feedback based on the sentiment queues" where, in one example, "the system can determine that a particular portion of the visual content elicited a negative sentiment from one or more participants in the videoconference and provide that information as feedback to the editor."; Chau, ¶ [0018]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Chau to include wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session and wherein the first report provides suggestions based on at least the positive and/or negative sentiment. Chau discloses the detection and use metadata about the meeting to determine feedback from the meeting itself, including the use of sentiment queues which “correspond to different portions of the visual content,” as well as providing feedback, such that the system can determine portions which elicited changes in engagement or sentiment (e.g., negative or positive), to address known difficulties in “analyz[ing] each individual participants' video stream for engagement queues,” especially when dealing with “a large number of participants in the videoconference,” and feedback providing the known benefit of allowing the “editor” to “make adjustments to the visual content,” which can “improve the effectiveness of the visual content the next time it is presented,” as recognized by Chau. (Chau, ¶ [0017]-[0018]).
Regarding claim 11, the rejection of claim 7 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. Rivera-Rodriguez further discloses wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session (Discloses "one or more pre-trained machine learning models may be used to detect or identify emotions" and, in one example, the presented "may ask a question related to the content that is being shared, while one or more other meeting participants may communicate agreement {a positive sentiment}, or disagreement {a negative sentiment}, with the presenter (e.g., the first meeting participant) by making a non-verbal communication, such as a hand or head gesture," where the identified emotion as related to communicated agreement or disagreement, is sentiment analysis indicating a positive or negative sentiment. Further "Each video stream is independently analyzed using various pre-trained machine learning models of the media processing service to generate text {for each of the separately-analyzed portions of the first multi-participant interaction session}."; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 16, line 51 - col. 17, line 20). However, Rivera-Rodriguez fails to expressly recite wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session, and wherein the first report provides suggestions based on at least the positive and/or negative sentiment.
The relevance of Chau is described above with relation to claim 4. Regarding claim 11, Chau teaches wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session ("The video data may depict the participants in the videoconference" which can be "analyzed using a trained machine-learning model to identify sentiment queues corresponding to different portions of the visual content" and "each participant’s sentiment may change throughout the presentation of the visual content" where "some portions (e.g., slides) of the visual content may elicit a positive sentiment while other portions may elicit a negative sentiment"; Chau, ¶ [0018]), and wherein the first report provides suggestions based on at least the positive and/or negative sentiment ("The system can identify these sentiment queues and generate feedback based on the sentiment queues" where, in one example, "the system can determine that a particular portion of the visual content elicited a negative sentiment from one or more participants in the videoconference and provide that information as feedback to the editor."; Chau, ¶ [0018]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Chau to include wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of separately-analyzed portions of the first multi-participant interaction session, and wherein the first report provides suggestions based on at least the positive and/or negative sentiment. Chau discloses the detection and use metadata about the meeting to determine feedback from the meeting itself, including the use of sentiment queues which “correspond to different portions of the visual content,” as well as providing feedback, such that the system can determine portions which elicited changes in engagement or sentiment (e.g., negative or positive), to address known difficulties in “analyz[ing] each individual participants' video stream for engagement queues,” especially when dealing with “a large number of participants in the videoconference,” and feedback providing the known benefit of allowing the “editor” to “make adjustments to the visual content,” which can “improve the effectiveness of the visual content the next time it is presented,” as recognized by Chau. (Chau, ¶ [0017]-[0018]).
Regarding claim 12, the rejection of claim 11 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. However, Rivera-Rodriguez fails to expressly recite further comprising: detecting triggers within the plurality of multi-modal signals for partitioning the first multi-participant interaction session into the separately-analyzed portions, wherein the first report correlates the triggers with the results of performing the sentiment analysis for each of the separately-analyzed portions of the first multi-participant interaction session.
The relevance of Chau is described above with relation to claim 4. Regarding claim 12, Chau teaches further comprising: detecting triggers within the plurality of multi-modal signals for partitioning the first multi-participant interaction session into the separately-analyzed portions, ("Some portions (e.g., slides) of the visual content may elicit a positive sentiment while other portions may elicit a negative sentiment" where, in one example, "the system can determine {detecting...}that a particular portion of the visual content {...within the plurality of multi-modal signals} elicited a negative sentiment {triggers...} from one or more participants in the videoconference" wherein said detected portions, also referred to as sentiment queues, are identified for separate analysis {partitioning the first multi-participant interaction session into the separately-analyzed portions}; Chau, ¶ [0018]) wherein the first report correlates the triggers with the results of performing the sentiment analysis for each of the separately-analyzed portions of the first multi-participant interaction session ("The system can identify these sentiment queues and generate feedback based on the sentiment queues," based on the elicited positive or negative sentiment which is associated with the sentiment queue.; Chau, ¶ [0018]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Chau to include further comprising: detecting triggers within the plurality of multi-modal signals for partitioning the first multi-participant interaction session into the separately-analyzed portions, wherein the first report correlates the triggers with the results of performing the sentiment analysis for each of the separately-analyzed portions of the first multi-participant interaction session. Chau discloses the detection and use metadata about the meeting to determine feedback from the meeting itself, including the use of sentiment queues which “correspond to different portions of the visual content,” as well as providing feedback, such that the system can determine portions which elicited changes in engagement or sentiment (e.g., negative or positive), to address known difficulties in “analyz[ing] each individual participants' video stream for engagement queues,” especially when dealing with “a large number of participants in the videoconference,” and feedback providing the known benefit of allowing the “editor” to “make adjustments to the visual content,” which can “improve the effectiveness of the visual content the next time it is presented,” as recognized by Chau. (Chau, ¶ [0017]-[0018]).
Regarding claim 18, the rejection of claim 16 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. However, Rivera-Rodriguez fails to expressly recite wherein the operations further comprise: detecting triggers within the plurality of multi-modal signals for partitioning the first multi-participant interaction session into separately-analyzed portions, wherein the first report correlates the triggers with the results of performing the sentiment analysis for each of the separately-analyzed portions of the first multi-participant interaction session, and wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of the separately-analyzed portions of the first multi-participant interaction session.
The relevance of Chau is described above with relation to claim 4. Regarding claim 18, Chau teaches wherein the operations further comprise: detecting triggers within the plurality of multi-modal signals for partitioning the first multi-participant interaction session into separately-analyzed portions, ("Some portions (e.g., slides) of the visual content may elicit a positive sentiment while other portions may elicit a negative sentiment" where, in one example, "the system can determine {detecting...}that a particular portion of the visual content {...within the plurality of multi-modal signals} elicited a negative sentiment {triggers...} from one or more participants in the videoconference" wherein said detected portions, also referred to as sentiment queues, are identified for separate analysis {partitioning the first multi-participant interaction session into the separately-analyzed portions}; Chau, ¶ [0018]) wherein the first report correlates the triggers with the results of performing the sentiment analysis for each of the separately-analyzed portions of the first multi-participant interaction session ("The system can identify these sentiment queues and generate feedback based on the sentiment queues," based on the elicited positive or negative sentiment which is associated with the sentiment queue.; Chau, ¶ [0018]), and wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of the separately-analyzed portions of the first multi-participant interaction session. ("The video data may depict the participants in the videoconference" which can be "analyzed using a trained machine-learning model to identify sentiment queues corresponding to different portions of the visual content" and "each participant’s sentiment may change throughout the presentation of the visual content" where "some portions (e.g., slides) of the visual content may elicit a positive sentiment while other portions may elicit a negative sentiment"; Chau, ¶ [0018]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Chau to include wherein the operations further comprise: detecting triggers within the plurality of multi-modal signals for partitioning the first multi-participant interaction session into separately-analyzed portions, wherein the first report correlates the triggers with the results of performing the sentiment analysis for each of the separately-analyzed portions of the first multi-participant interaction session, and wherein the results of performing the sentiment analysis indicate positive and/or negative sentiment for each of the separately-analyzed portions of the first multi-participant interaction session. Chau discloses the detection and use metadata about the meeting to determine feedback from the meeting itself, including the use of sentiment queues which “correspond to different portions of the visual content,” as well as providing feedback, such that the system can determine portions which elicited changes in engagement or sentiment (e.g., negative or positive), to address known difficulties in “analyz[ing] each individual participants' video stream for engagement queues,” especially when dealing with “a large number of participants in the videoconference,” and feedback providing the known benefit of allowing the “editor” to “make adjustments to the visual content,” which can “improve the effectiveness of the visual content the next time it is presented,” as recognized by Chau. (Chau, ¶ [0017]-[0018]).
Claims 6, 15, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Rivera-Rodriguez as applied to claim 1 and 7 above, and further in view of Peters (U.S. PG. Pub. No. 2024/0195939, hereinafter Peters).
Regarding claim 6, the rejection of claim 1 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. Rivera-Rodriguez further discloses wherein the instructions are further operative to: capture a second plurality of multi-modal signals from a second multi-participant interaction session, ("a media processing service 200 of an online meeting service receives and processes various video streams generated during an online meeting" where the second multi-participant interaction session is mere duplication of the same steps performed for the first multi-participant interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-38) wherein the second captured plurality of multi-modal signals comprises an audio feed, a video feed, and image stills of participants in the second multi-participant interaction session ("A first type of video stream is a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service. Generally, this first type of video stream will depict a video image of a meeting participant," where the video streams comprising video data, audio data ("the audio portion of each video stream of the first type"), and image stills (video data in video streams are understood as a time-coordinated sequence of still frames. Video data, audio data, and image stills {multi-modal signals} as present in the "first type of video stream" are of the participants in the second multi-participant (e.g., shown in Fig. 2 as meeting participants 1-4) interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-54; Col. 9, lines 1-24; FIG. 2); correlate timing information across the second captured plurality of multi-modal signals ("each video stream... of a meeting participant is received and processed by a first type of pre-trained machine learning model that converts the audio portion of the video stream, representing the spoken message of a meeting participant, to text" and "each video stream is also processed by one or more additional pre-trained machine learning models—specifically, one or more computer vision models" where, in some embodiments, "the output of these models is text that describes or otherwise represents the detected facial expression, emotion, or body language, along with timing data (e.g., a timestamp, or a combination of a beginning time and an ending time)" for each verbal and non-verbal communication (see FIG. 7).; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; FIG. 7); perform a further sentiment analysis using the second captured plurality of multi-modal signals and the correlated timing information across the second captured plurality of multi-modal signals (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt that is provided as input to the generative language model," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46); generate a second report indicating results of performing the further sentiment analysis (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt {using a prompt} that is provided as input to the generative language model {...with a generative language model}," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46). However, Rivera-Rodriguez fails to expressly recite compile the first report and the second report into an aggregate report.
Peters teaches systems and methods for managing concurrent conference sessions. (Peters, ¶ [0002] ). Regarding claim 6, Peters teaches compile the first report and the second report into an aggregate report (Discloses the "system provides information about the collaborative and emotional status of participants in different breakout sessions from a communication session" where "the system can report to the host a summary of the emotional and collaborative status of each breakout session allowing the host to know which breakout sessions to direct their attention to" where the summary can be derived from the "average or aggregate measures of sentiment" for each breakout room; Peters, ¶ [0584], [0597]-[0598]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Peters to include compile the first report and the second report into an aggregate report. The breakout session monitoring systems of Peters “can report to the host a summary of the emotional and collaborative status of each breakout session allowing the host to know which breakout sessions to direct their attention to,” which allows a presenter or host “to see how project teams or small-group discussions are progressing,” which in the context of Rivera-Rodriguez provides the known benefit of monitoring the sentiment and emotional condition of multiple smaller groups, such that problems can be addressed quickly, as recognized in light of Peters. (Peters, ¶ [0584]-[0585]).
Regarding claim 15, the rejection of claim 7 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. Rivera-Rodriguez further discloses further comprising: capturing a second plurality of multi-modal signals from a second multi-participant interaction session, ("a media processing service 200 of an online meeting service receives and processes various video streams generated during an online meeting" where the second multi-participant interaction session is mere duplication of the same steps performed for the first multi-participant interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-38) wherein the second captured plurality of multi-modal signals comprises an audio feed, a video feed, and image stills of participants in the second multi-participant interaction session ("A first type of video stream is a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service. Generally, this first type of video stream will depict a video image of a meeting participant," where the video streams comprising video data, audio data ("the audio portion of each video stream of the first type"), and image stills (video data in video streams are understood as a time-coordinated sequence of still frames. Video data, audio data, and image stills {multi-modal signals} as present in the "first type of video stream" are of the participants in the second multi-participant (e.g., shown in Fig. 2 as meeting participants 1-4) interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-54; Col. 9, lines 1-24; FIG. 2); correlating timing information across the second captured plurality of multi-modal signals ("each video stream... of a meeting participant is received and processed by a first type of pre-trained machine learning model that converts the audio portion of the video stream, representing the spoken message of a meeting participant, to text" and "each video stream is also processed by one or more additional pre-trained machine learning models—specifically, one or more computer vision models" where, in some embodiments, "the output of these models is text that describes or otherwise represents the detected facial expression, emotion, or body language, along with timing data (e.g., a timestamp, or a combination of a beginning time and an ending time)" for each verbal and non-verbal communication (see FIG. 7).; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; FIG. 7); performing a further sentiment analysis using the second captured plurality of multi-modal signals and the correlated timing information across the second captured plurality of multi-modal signals (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt that is provided as input to the generative language model," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46); generating a second report indicating results of performing the further sentiment analysis (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt {using a prompt} that is provided as input to the generative language model {...with a generative language model}," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46). However, Rivera-Rodriguez fails to expressly recite compiling the first report and the second report into an aggregate report.
The relevance of Peters is described above with relation to claim 6. Regarding claim 15, Peters teaches compiling the first report and the second report into an aggregate report (Discloses the "system provides information about the collaborative and emotional status of participants in different breakout sessions from a communication session" where "the system can report to the host a summary of the emotional and collaborative status of each breakout session allowing the host to know which breakout sessions to direct their attention to" where the summary can be derived from the "average or aggregate measures of sentiment" for each breakout room; Peters, ¶ [0584], [0597]-[0598]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Peters to include compiling the first report and the second report into an aggregate report. The breakout session monitoring systems of Peters “can report to the host a summary of the emotional and collaborative status of each breakout session allowing the host to know which breakout sessions to direct their attention to,” which allows a presenter or host “to see how project teams or small-group discussions are progressing,” which in the context of Rivera-Rodriguez provides the known benefit of monitoring the sentiment and emotional condition of multiple smaller groups, such that problems can be addressed quickly, as recognized in light of Peters. (Peters, ¶ [0584]-[0585]).
Regarding claim 20, the rejection of claim 16 is incorporated. Rivera-Rodriguez discloses all of the elements of the current invention as stated above. Rivera-Rodriguez further discloses wherein the operations further comprise: capturing a second plurality of multi-modal signals from a second multi-participant interaction session, ("a media processing service 200 of an online meeting service receives and processes various video streams generated during an online meeting" where the second multi-participant interaction session is mere duplication of the same steps performed for the first multi-participant interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-38) wherein the second captured plurality of multi-modal signals comprises an audio feed, a video feed, and image stills of participants in the second multi-participant interaction session ("A first type of video stream is a video stream generated using a video camera device of a client computing device executing the client software application for the online meeting service. Generally, this first type of video stream will depict a video image of a meeting participant," where the video streams comprising video data, audio data ("the audio portion of each video stream of the first type"), and image stills (video data in video streams are understood as a time-coordinated sequence of still frames. Video data, audio data, and image stills {multi-modal signals} as present in the "first type of video stream" are of the participants in the second multi-participant (e.g., shown in Fig. 2 as meeting participants 1-4) interaction session.; Rivera-Rodriguez, ¶ Col. 8, lines 19-54; Col. 9, lines 1-24; FIG. 2); correlating timing information across the second captured plurality of multi-modal signals ("each video stream... of a meeting participant is received and processed by a first type of pre-trained machine learning model that converts the audio portion of the video stream, representing the spoken message of a meeting participant, to text" and "each video stream is also processed by one or more additional pre-trained machine learning models—specifically, one or more computer vision models" where, in some embodiments, "the output of these models is text that describes or otherwise represents the detected facial expression, emotion, or body language, along with timing data (e.g., a timestamp, or a combination of a beginning time and an ending time)" for each verbal and non-verbal communication (see FIG. 7).; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; FIG. 7); performing a further sentiment analysis using the second captured plurality of multi-modal signals and the correlated timing information across the second captured plurality of multi-modal signals (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt that is provided as input to the generative language model," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46); generating a second report indicating results of performing the further sentiment analysis (The text based digital representation provided to the language model includes "text that describes or otherwise represents the detected facial expression, emotion, or body language" and this text "may be selected for inclusion in a prompt {using a prompt} that is provided as input to the generative language model {...with a generative language model}," and which is used by the generative language model to "generate accurate and complete summary descriptions of the online meeting," where an accurate and complete summarization of "detected facial expression, emotion, or body language" is sentiment analysis; Rivera-Rodriguez, ¶ Col. 5, lines 22-46; Col. 7, lines 14-46). However, Rivera-Rodriguez fails to expressly recite compiling the first report and the second report into an aggregate report.
The relevance of Peters is described above with relation to claim 6. Regarding claim 20, Peters teaches compiling the first report and the second report into an aggregate report (Discloses the "system provides information about the collaborative and emotional status of participants in different breakout sessions from a communication session" where "the system can report to the host a summary of the emotional and collaborative status of each breakout session allowing the host to know which breakout sessions to direct their attention to" where the summary can be derived from the "average or aggregate measures of sentiment" for each breakout room; Peters, ¶ [0584], [0597]-[0598]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the meeting analyzer service of Rivera-Rodriguez to incorporate the teachings of Peters to include compiling the first report and the second report into an aggregate report. The breakout session monitoring systems of Peters “can report to the host a summary of the emotional and collaborative status of each breakout session allowing the host to know which breakout sessions to direct their attention to,” which allows a presenter or host “to see how project teams or small-group discussions are progressing,” which in the context of Rivera-Rodriguez provides the known benefit of monitoring the sentiment and emotional condition of multiple smaller groups, such that problems can be addressed quickly, as recognized in light of Peters. (Peters, ¶ [0584]-[0585]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Reece (U.S. PG. Pub. No. 2022/0343911) discloses a machine learning system for determining conversation analysis indicators using acoustic, video, and text data of a multiparty conversation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sean E. Serraguard whose telephone number is (313)446-6627. The examiner can normally be reached 07:00-17:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C. Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Sean E Serraguard/Primary Examiner, Art Unit 2657