DETAILED ACTION
Response to Arguments
Applicant's arguments filed with respect to claims 9-21 have been fully considered but are moot in view of the new ground(s) of rejection. The rejections are necessitated due to claim amendments.
Allowable Subject Matter
Claims 22-24 are allowed.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 9-12, 16, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Quennesson (Pub. No. US 2018/0025078) in view of Shen et al. (Pub. No. US 2019/0228380).
Regarding claim 9, Quennesson teaches a method comprising [Para. 43]: identifying an audio-video (AV) entity (video broadcast stream) [Para. 62 “The video highlight creator 380 may obtain a video broadcast stream 364, and may receive one or more social media metrics indicating a volume of social media engagements 312 associated with the video broadcast stream 364 from a social media platform 304”]; using audio (audio component) from the AV entity (video broadcast stream), identifying a plural of first candidate segments (video segments) of the AV entity for establishing a summary of the entity [Para. 71 “The audio analyzer 314 may be configured to analyze the audio component of the video broadcast stream 364 to obtain information helpful for automatically creating the video segments of the video highlights 381. “For example, the audio analyzer 314 may analyze the audio component of the video broadcast stream 364 to obtain information that can assist with identifying the video broadcast stream 364, the nature of the video broadcast stream 364, the underlying events or objects (including persons) captured by the video broadcast stream 364, and/or the starting and ending points for one or more video segments of the video highlights 381”]; using video (video component) from the AV entity, identifying a plurality of second candidate segments (video segments) of the AV entity for establishing a summary of the entity [Para. 68 “The video analyzer 322 may be configured to analyze the video component of the video broadcast stream 364 to obtain information helpful for automatically creating the video highlights 181.”]; identifying at least one parameter (sentiment) associated with chat (social media engagement) related to the AV entity (video broadcast stream) [Para. 62 “The social media engagements 312 may be information exchanged on the social media platform that relates to the video broadcast stream 364 such as messages, posts, comments, signals of appreciation about the video broadcast stream 364”; Para. 73 “if the video highlight creator 380 determines that the sentiment associated with the social media engagements 312 is negative (or beyond a certain threshold), the video highlight creator 380 may decide not to include that video segments within the video highlights 381. On the other hand, if the video highlight creator 380
determines that the sentiment associated with the segment's social media engagements 312 is relatively positive, the video highlight creator 380 may decide to include those video segments in the video highlights 381”]; selecting at least some of the plurality of first and second candidate segments (video segments) based at least in part on the parameter (sentiment) [Para. 73]; and using the at least some of the plurality of first and second candidate segments (selected video segments), generating a video summary (video highlights) of the AV entity that is shorter than the AV entity [Para. 3, and 4. It is clear that the summary/highlights are shorter].
Quennesson also teaches wherein identifying the plurality of first candidate segments (video segments) comprises converting speech of each individual voice track to words (text) and identifying the plurality of first candidate segments based on the words (text) from the individual voice tracks [Para. 71 and “The audio analyzer 314 may be configured to analyze the closed captioned data for keywords. In some examples, the audio analyzer 314 may include a speech-to-text unit 315 configured to perform speech recognition on the audio component of the video broadcast stream 364, and analyze the text of the audio component for keywords. The audio analyzer 314 may use any known types of natural language processing to detect keywords from the speech of the audio. The audio analyzer 314 may determine which point in the audio or video component the one or more keywords were spoken so that the video highlight creator 380 may have a more accurate point regarding the detection of that relevant moment and/or the start and end of a video segment of the video highlights 381”].
However, Quennesson doesn’t explicitly teach separating the audio from the AV entity into individual voice tracks.
Shen teaches wherein identifying the plurality of first candidate segments (time sections of particular streams) comprises separating the audio from the AV entity (meeting) into individual voice tracks (individual voice streams), converting speech of each individual voice track to words (text), and identifying the plurality of first candidate segments (time sections of particular streams) based on the words from the individual voice tracks [Para. 60, Para. 67 “The individual voice streams may be automatically transcribed, if desired. Transcription of the individual voice streams may be automatically accomplished, for example, via an online (e.g., cloud-based) automatic speech recognition service, an on-device deep-learning method, an offline speech recognition software, or another similar technology”; Para. 73 “Transcripts of the different voice streams may be collected and used to construct a database that can later be used to search, retrieve, and playback particular (e.g., user-selected) aspects of the meeting. For example, an indexing engine may receive the voice streams and map words and/or phrases of the associated text to stream elements (e.g., to time sections of particular streams, to the identification of the attendee assigned to the stream, etc.). In addition, the transcribed text may be distilled into summaries and classified into any number of different topics”; Para. 79 “It should be noted that, when multiple discussions occurring at different times are associated with the same discussion topic, the topic list may include multiple time spans for a give topic due to the coalescing discussed above”; Para. 92, 100, fig. 3 and related description].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify quennesson’s video-highlight creation system so that it performs Shen’s complete teaching of the limitations namely treating the AV entity (meeting) as audio-video content, separating its mixed speech into individual voice tracks (individual voice steams), converting each track to words (text), and identifying first candidate segments (time sections of particular streams) from the words (text). This modification improves Quennesson by preserving speaker-specific speech provenance and aligning each spear’s recognized text with candidate time sections, thereby improving highlight accuracy for multi-speaker content.
Regarding claim 10, Quennesson teaches presenting the video summary/highlight on a display [Para. 51 and 90].
Regarding claim 11, Quennesson teaches wherein using video from the AV entity for identifying plural second candidate segments of the AV entity comprises identifying scene changes in the AV entity [Para. 70].
Regarding claim 12, Quennesson teaches wherein using video from the AV entity for identifying plural second candidate segments of the AV entity comprises identifying text (closed captioned data) in the video of the AV entity [Para. 71].
Regarding claim 16, Quennesson teaches wherein using audio from the AV entity for identifying plural first candidate segments of the AV entity comprises identifying words (keywords) in speech in the audio [Para. 71].
Regarding 17, Quennesson teaches wherein identifying the parameter associated with chat (social media engagement) related to the AV entity comprises identifying sentiment of the chat [Para. 73].
Regarding 19. Quennesson teaches wherein identifying the parameter associated with chat related to the AV entity comprises identifying topic of the chat [Para. 72].
Claims 13-15 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Quennesson (Pub. No. US 2018/0025078) in view of Shen et al. (Pub. No. US 2019/0228380) further in view of Dwyer et al. (Pub. 2015/0195406).
Regarding claim 13, Quennesson in view of Shen doesn’t explicitly teach the claim limitation.
However, Dwyer teaches wherein using audio from the AV entity for identifying plural first candidate segments (call snippets) of the AV entity comprises identifying acoustic events (acoustic events) in the audio [Para. 159 and 154].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Quennesson in view of Shen to teach the claim limitation, feature as taught by Dwyer; because the modification enables the system to real-time automated monitoring systems for monitoring and improving live communications, including by providing feedback on communications performance.
Regarding claim 14, Quennesson in view of Shen doesn’t explicitly teach the claim limitation.
However, Dwyer teaches wherein using audio from the AV entity for identifying plural first candidate segments of the AV entity comprises identifying pitch and/or amplitude (volume) of at least one voice in the audio [Para. 154].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Quennesson in view of Shen to teach the claim limitation, feature as taught by Dwyer; because the modification enables the system to real-time automated monitoring systems for monitoring and improving live communications, including by providing feedback on communications performance.
Regarding claim 15, Quennesson in view of Shen doesn’t explicitly teach the claim limitation.
However, Dwyer teaches wherein using audio from the AV entity for identifying plural first candidate segments of the AV entity comprises identifying emotion (emotion score) in the audio [Para. 120].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Quennesson in view of Shen to teach the claim limitation, feature as taught by Dwyer; because the modification enables the system to real-time automated monitoring systems for monitoring and improving live communications, including by providing feedback on communications performance.
Regarding 18, Quennesson in view of Shen doesn’t explicitly teach the claim limitation.
However, Dwyer teaches wherein identifying the parameter associated with chat related to the AV entity comprises identifying emotion (emotional score) of the chat [Para. 120].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Quennesson in view of Shen to teach the claim limitation, feature as taught by Dwyer; because the modification enables the system to real-time automated monitoring systems for monitoring and improving live communications, including by providing feedback on communications performance.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Quennesson (Pub. No. US 2018/0025078) in view of Shen et al. (Pub. No. US 2019/0228380) further in view of IYER (Pub. No. US 2021/0201045).
Regarding 20, Quennesson in view of Shen doesn’t explicitly teach the claim limitation.
However, IYER teaches wherein identifying the parameter associated with chat related to the AV entity comprises identifying at least one grammatical category of at least one word in the chat [Para. 36].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Quennesson in view of Shen to teach the claim limitation, feature as taught by IYER; because the modification enables the system to receive real-time user’s reaction in order to improve user’s experience.
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Quennesson (Pub. No. US 2018/0025078) in view of Shen et al. (Pub. No. US 2019/0228380) further in view of Gunawardena (Pub. No. US 2021/0110166).
Regarding 21, Quennesson in view of Shen doesn’t explicitly teach the claim limitation.
However, Gunawardena teaches wherein identifying the parameter associated with chat related to the AV entity comprises identifying a summary of the chat [Para. 48, 49, 124 and 126].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Quennesson in view of Shen to teach the claim limitation, feature as taught by Gunawardena; because the modification enables the system automatically update all relevant information automatically for system improvement.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOLOMON G BEZUAYEHU whose telephone number is (571)270-7452. The examiner can normally be reached on Monday-Friday 10 AM-8 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 888-786-0101 (IN USA OR CANADA) or 571-272-4000.
/SOLOMON G BEZUAYEHU/
Primary Examiner, Art Unit 2666