DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1-3, 6-8, 10 and 12-22 are pending for examination.
Claims 1, 17 and 20 are independent Claims.
Claims 1-3, 6-8, 10 and 12-22 are rejected under 35 U.S.C. §103.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 6-8 and 15-22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blinnikka (U.S. 2008/0126387 hereinafter Blinnikka) in view of Mody et al. (U.S. 2020/0272693 hereinafter Mody) in further view of Faulkner et al. (U.S. 2018/0331842 hereinafter Faulkner).
As Claim 1, Blinnikka teaches a method for multimedia processing, comprising:
displaying first text content, wherein the first text content corresponds to a first multimedia content (Blinnikka (¶0038 line 1-6, fig. 3 item 306), text content is displayed in area 306. The text is corresponding to the media stream)
playing second multimedia content in response to a triggering operation for the second multimedia content (Blinnikka (¶0037 line 4-8), user can play and watch the media stream),
Blinnikka does not explicitly disclose:
and is obtained by automatic speech recognition of the first multimedia content; and
wherein the second multimedia content comprises a segment of the first multimedia content that is associated with at least one second text content, and the at least one second text content is extracted from the first text content, and
wherein the second multimedia content is a summarized multimedia segment comprising at least two multimedia sub-segments, each multimedia sub-segment corresponds to a summary type selected from a plurality of summary types, the plurality of summary types comprising content-based semantic categories of the first multimedia content, and the at least one second text content comprises a text summary of at least one of the plurality of summary types; and
wherein the method further comprises:
during the playing of the second multimedia content, displaying an indicator of the at least one summary type corresponding to the text summary in association with a play timeline of the second multimedia content.
Mody teaches:
and is obtained by automatic speech recognition of the first multimedia content (Mody (¶0058 line 1-7, fig. 4 item 404), “the text conversion module 204 generates a text file from the recording of the teleconference meeting. The text conversion module 204 may use any known text conversion technique to generate the text file. The resulting text file may include written text of the conversations captured between meeting participants during the videoconference meeting.”); and
wherein the second multimedia content comprises a segment of the first multimedia content (Mody (¶0072 line 1-4), “As shown, each topic is presented along with a user interface element 610, 612, 614 that is selectable to cause playback of a portion of the recorded meeting that relates to its corresponding topic.”) that is associated with at least one second text content (Mody (¶0074 last 6 lines), “the summary section 608 may be updated based on the user's selection of the expansion button 616 corresponding to topic 1. For example, the summary section 606 may be updated to include a summary that is focused on topic 1 and including statements that relate to topic 1.”), and the at least one second text content is extracted from the first text content (Mody (¶0059 line 4-5), “identifying topics from the text file, and ultimately generating a meeting summary”), and
wherein the second multimedia content is a summarized multimedia segment comprising at least two multimedia sub-segments, each multimedia sub-segment corresponds to a summary type selected from a plurality of summary types (Mody (¶0074 line 3-5), “selection of the expansion button 616 causes two statements corresponding to topic 1 to be listed in the topic listing section 606”), the plurality of summary types comprising content-based semantic categories of the first multimedia content (Mody (¶0074 last 6 lines), “the summary section 608 may be updated based on the user's selection of the expansion button 616 corresponding to topic 1. For example, the summary section 606 may be updated to include a summary that is focused on topic 1 and including statements that relate to topic 1.”), and the at least one second text content comprises a text summary of at least one of the plurality of summary types (Mody (¶0074 last 6 lines), “the summary section 608 may be updated based on the user's selection of the expansion button 616 corresponding to topic 1. For example, the summary section 606 may be updated to include a summary that is focused on topic 1 and including statements that relate to topic 1.”); and
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify close caption of Blinnikka instead be summarized videos taught by Mody, with a reasonable expectation of success. The motivation would be conveniently solve a problem that “A meeting may last multiple hours, making it difficult to find relevant portions. To alleviate this issue, notes, meeting minutes, and/or a meeting summary may be generated to provide some insight into the contents of the meeting.” (Mody (¶0002 line 7-11)).
Blinnikka in view of Mody may not explicitly disclose:
wherein the method further comprises:
during the playing of the second multimedia content, displaying an indicator of the at least one summary type corresponding to the text summary in association with a play timeline of the second multimedia content.
Faulkner teaches:
during the playing of the second multimedia content, displaying an indicator of the at least one summary type corresponding to the text summary in association with a play timeline of the second multimedia content (Faulkner (¶0075, fig. 7), “An individual marker 708(1) through 708(1) or 716 can be displayed with a portion ( e.g., snippet) of surrounding text that captures what was being said at the time the notable event(s) occur.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify text summary of Blinnikka in view of Mody instead be summary on timeline taught by Faulkner, with a reasonable expectation of success. The motivation would be conveniently solve a problem that “the user can quickly access a corresponding playback position of content (e.g., recorded content) to gain an understanding of reasons the increased activity occurred ( e.g., what was said that caused an increased number of people to submit a like reaction, why did an increased number of people leave the conference session, why did an increased number of people submit a comment in the chat conversation, etc.)” (Faulkner (¶0071 last 9 lines)).
As Claim 2, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner teaches wherein the at least one second text content comprises at least two consecutive text segments extracted from the first text content (Blinnikka (¶0043 line 1-8, ¶0052 line 1-4), text includes whole line of text or multiple words).
As Claim 3, besides Claim 2, Blinnikka in view of Mody in further view of Faulkner teaches wherein the playing second multimedia content in response to a triggering operation for the second content comprises:
in response to the triggering operation for the second multimedia content, according to an order of associated time periods of the at least two consecutive text segments of the at least one target second content in the first multimedia content (Mody (¶0058 line 1-7, fig. 4 item 404), “the text conversion module 204 generates a text file from the recording of the teleconference meeting. The text conversion module 204 may use any known text conversion technique to generate the text file. The resulting text file may include written text of the conversations captured between meeting participants during the videoconference meeting.”),
playing by jumping across multimedia segments in the first multimedia content, each of the multimedia segments corresponding to an associated time period of one of the consecutive text segments (Faulkner (¶0071 line 10-13), “Moreover, a user is able to interact with individual representations on the interactive timeline 604 to access and/or view information associated with the activity ( e.g., so the user can better understand the context of the activity”).
As Claim 6, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner teaches wherein the generating the second multimedia content based on the associated time period of the at least one second text content comprises:
in response to that there are a plurality of associated time periods, generating the second multimedia content by joining a plurality of multimedia segments corresponding to the plurality of associated time periods according to an order of the plurality of associated time periods in the first multimedia content (Mody (¶0058 line 1-7, fig. 4 item 404), “the text conversion module 204 generates a text file from the recording of the teleconference meeting. The text conversion module 204 may use any known text conversion technique to generate the text file. The resulting text file may include written text of the conversations captured between meeting participants during the videoconference meeting.”).
As Claim 7, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner teaches wherein the generating the second multimedia content based on the associated time period of the at least one first text content comprises:
adjusting the associated time period of the at least one second text content based on a sentence integrity of an associated text of the at least one second text content (Blinnikka (¶0029 line 1-10), user can overwrite the start /end point of the portion); and
generating the second multimedia content based on the adjusted associated time period (Blinnikka (¶0029 line 1-10), user can overwrite the start /end point of the portion).
As Claim 8, besides Claim 7, Blinnikka in view of Mody in further view of Faulkner teaches wherein the second text is a text corresponding to the at least one first text content in the initial text content (Blinnikka (¶0051 line 1-4, ¶0053 line 1-5), user selects a starting point and ending point of text segment within the text body).
As Claim 15, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner teaches further comprising:
receiving a download operation for the second multimedia content from a user, downloading and storing the second multimedia content (Blinnikka (¶0034 last 5 lines), display device 204 accesses content over a network).
As Claim 16, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner teaches wherein the displaying first text content comprises: displaying the first text content in response to a triggering operation for a list page, wherein the list page displays abstract information of a plurality of multimedia contents (Blinnikka (¶0038 line 1-5), text displays multiple rows of text).
As Claim 17, Blinnikka teaches an electronic device, comprising:
a processor (Blinnikka (¶0031 line 4-5), processor and memory); and
a memory configured to store instructions that are executable by the processor (Blinnikka (¶0031 line 4-5), processor and memory);
the processor being configured to read the instructions from the memory and execute the instructions to implement a method for multimedia processing (Blinnikka (¶0031 line 4-5), processor and memory) comprising:
The rest of the limitation(s) are rejected for the same reasons as Claim 1.
As Claim 18, the Claim is rejected for the same reasons as Claim 2.
As Claim 19, the Claim is rejected for the same reasons as Claim 3.
As Claim 20, the Claim is rejected for the same reasons as Claim 1.
As Claim 21, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner teaches determining the second multimedia content based on the at least one second text content (Mody (¶0058 line 1-7, fig. 4 item 404), “the text conversion module 204 generates a text file from the recording of the teleconference meeting. The text conversion module 204 may use any known text conversion technique to generate the text file. The resulting text file may include written text of the conversations captured between meeting participants during the videoconference meeting.”).
As Claim 22, besides Claim 21, Blinnikka in view of Mody in further view of Faulkner teaches wherein the determining the second multimedia content based on the at least one second text content comprises:
generating the second multimedia content based on the associated time period of the at least one second text content, wherein the associated time period of the at least one second text content is used to characterize a time period of speech information corresponding to the at least one second text content in the second multimedia content (Mody (¶0058 line 1-7, fig. 4 item 404), “the text conversion module 204 generates a text file from the recording of the teleconference meeting. The text conversion module 204 may use any known text conversion technique to generate the text file. The resulting text file may include written text of the conversations captured between meeting participants during the videoconference meeting.”).
Claim(s) 10 and 12-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blinnikka and Mody in view of Faulkner in further view of Gilson (U.S. 2012/0059954 hereinafter Gilson).
As Claim 10, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner does not explicitly disclose:
wherein the playing second multimedia content in response to a triggering operation for the second multimedia content comprises:
determining a type of the summary corresponding to the triggering operation in response to the triggering operation for the second multimedia content; and
obtaining a target multimedia sub-segment corresponding to the type of the summary and playing the target multimedia sub-segment; or
obtaining the summarized multimedia segment and playing the summarized multimedia segment based on a time period of the type of the summary in the summarized multimedia segment.
Gilson teaches wherein the playing an associated multimedia content in response to a triggering operation for the associated multimedia content comprises:
determining a type of the summary corresponding to the triggering operation in response to the triggering operation for the second multimedia content; and (Gilson (¶0074 line 17-26), user selects a caption stream for display); and
obtaining a target multimedia sub-segment corresponding to the type of the summary and playing the target multimedia sub-segment; or
obtaining the summarized multimedia segment and playing the summarized multimedia segment based on a time period of the type of the summary in the summarized multimedia segment.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify caption content of Blinnikka in view of Mody in further view of Faulkner instead be a multiple caption system taught by Gilson, with a reasonable expectation of success. The motivation would be to provide more convenient, usable and/or advanced captioning functionalities (Gilson (¶0002).
As Claim 12, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner does not explicitly disclose:
wherein the displaying an identification of the type of the summary corresponding to the text summary in association on the play timeline of the second multimedia content comprises:
displaying the identification of the type of the summary corresponding to the target text summary at an associated time point corresponding to the text summary on the play timeline of the second multimedia content
Gilson teaches wherein the displaying an identification of the type of the summary corresponding to the text summary in association on the play timeline of the second multimedia content comprises:
displaying the identification of the type of the summary corresponding to the target text summary at an associated time point corresponding to the text summary on the play timeline of the second multimedia content (Gilson (¶0030 line 1-5, fig. 4A), caption type is indicated with playback speed. Caption is displayed next to the scene).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify caption content of Blinnikka in view of Din in further view of Lim instead be a multiple caption system taught by Gilson, with a reasonable expectation of success. The motivation would be to provide more convenient, usable and/or advanced captioning functionalities (Gilson (¶0002).
As Claim 13, besides Claim 12, Blinnikka and Mody in view of Faulkner in further view of Gilson teaches wherein the associated time point corresponding to the text summary is a time point in the associated time period of the text summary (Gilson (¶0030 line 1-5, fig. 4A), caption type is indicated with playback speed. Caption is displayed in according to the scene).
As Claim 14, besides Claim 1, Blinnikka in view of Mody in further view of Faulkner does not explicitly disclose:
further comprising:
during the playing of second multimedia content, prominently displaying at least one second text content corresponding to a playing progress of the second multimedia content in sequence.
Gilson teaches:
during the playing of second multimedia content, prominently displaying at least one second text content corresponding to a playing progress of the second multimedia content in sequence (Gilson (¶0030 line 1-5, fig. 4A), caption type is indicated with playback speed. Caption is displayed in according to the scene).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify caption content of , Blinnikka in view of Din in further view of Bishop instead be a multiple caption system taught by Gilson, with a reasonable expectation of success. The motivation would be to provide more convenient, usable and/or advanced captioning functionalities (Gilson (¶0002).
Response to Arguments
Rejections under 35 U.S.C. §§ 102 and 103:
As Claim 1, Applicants argue that cited references do not disclose “semantic categories of the first multimedia content” (first paragraph under section A of page 10 in the remarks).
Applicants’ arguments are moot because new reference Mody teaches the limitation(s).
As Claim 1, Applicants argue that cited references do not disclose “a summarized multimedia segment” during a playback (first paragraph under section B of page 11 in the remarks).
Applicants’ arguments are moot because new reference Faulkner teaches the limitation(s).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Mahmoud (U.S. 2019/0384813) disclose a system/method for generating and insights from meeting recording.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NHAT HUY T NGUYEN whose telephone number is (571)270-7333. The examiner can normally be reached M-F: 12:00-8:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NHAT HUY T NGUYEN/Primary Examiner, Art Unit 2147