DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/12/2026 has been entered.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Response to Amendment
This communication is responsive to the applicant’s amendment dated 06/12/2026. The applicant(s) amended claims 1, 12, and 13 and added claims 14-19.
Response to Arguments
Applicant's arguments with respect to claims 1, 12, and 13 have been considered but are moot in view of the new ground(s) of rejection because the arguments pertain to the newly amended limitations.
Claim Rejections - 35 USC § 103
Claim(s) 1-19 are rejected under 35 U.S.C. 103 as being unpatentable over Donaldson (US 20150348538 A1) in view of Huang et al. (US 20230035155 A1).
Regarding claims 1, 12, and 13, Donaldson teaches:
“A summarizing system comprising processing circuitry” (par. 0018; ‘Various embodiments or examples may be implemented in numerous ways, including as a system, a process, an apparatus, a user interface, or a series of program instructions on a computer readable medium such as a computer readable storage medium or a computer network where the program instructions are sent over optical, electronic, or wireless communication links.’) configured to:
“acquire speech” (par. 0020; ‘The audio signal may include speech.’);
“convert the speech into a plurality of texts” (par. 0026; ‘Speech recognizer 221 may translate or convert spoken words into text.’);
“generate a summary of the plurality of texts when the plurality of texts satisfy a summarizing execution condition” (par. 0029; ‘Summary generator 213 may be configured to generate a summary of the speech.’; par. 0030; ‘A content summary may be determined based on the words, vocal fingerprints, speakers, acoustic properties, or other parameters determined by speech analyzer 212. For example, based on word counts, and a comparison to the frequency that the words are used in the general English language, one or more keywords may be identified.’);
“output the summary to a user” (par. 0052; ‘Before connecting the caller to the conference, a speech summary manager may present the summary or keyword to the caller. The summary or keyword may be presented using a loudspeaker local to the caller, which may be remote from the speech summary manager.’).
However, Donaldson does not expressly teach:
“and in response to receiving a change in at least a part of the plurality of texts caused by editing of text information included in the plurality of texts, update the summary based on the change.”
In a similar field of endeavor (generating summaries of text), Huang teaches:
“and in response to receiving a change in at least a part of the plurality of texts caused by editing of text information included in the plurality of texts, update the summary based on the change” (par. 0096; ‘In some implementations, a user may choose to modify the transcript highlights (e.g., extend, contract, add, or remove highlights), or leave the auto generated highlights of the text summary alone. For example, the technique 800 of FIG. 8 may be implemented to present the highlighted transcript to a user and receive user modifications of the highlighting that can be used to update the text summary of the transcript.’).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Donaldson’s summary generator by incorporating Huang’s summary updating method in order to update a summary based on modified transcript highlights. The combination would provide a way of automatically extracting a summary from a conference recording transcript using natural language processing techniques. (Huang: par. 0042)
Regarding claim 2 (dep. on claim 1), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry determines that the summarizing execution condition is satisfied when an amount of the plurality of texts that are converted reached a first threshold value, the amount of the plurality of texts being calculated by counting a number of characters included in the plurality of texts respectively” (Donaldson: par. 0047; ‘FIG. 7A illustrates examples of a flowchart for determining keywords based on one or more speech parameters, such as word count, vocal fingerprint, acoustic properties, and the like, according to some examples.’; par. 0052; ‘A keyword may be determined by assigning weights words referenced in the speech based on word counts, vocal fingerprint durations, acoustic properties, and other parameters, and determining a significance of a word. The significance of a vocal fingerprint may also be determined based on word counts and other parameters, and the significance of a vocal fingerprint may in turn affect the significance of a word.’).
Regarding claim 3 (dep. on claim 1), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry: generates chunk data of the plurality of texts” (Huang: par. 0043; ‘The system may also consider the duration of speaker segments (i.e., sections of uninterrupted speaking by one speaker) when selecting highlights to ensure the most relevant strings from the longest speaker segments are included as highlights.’);
“vectorizes the chunk data” (Huang: par. 0090; ‘FIG. 13 is a flowchart of an example of a technique 1300 for generating a summary of a video recording of a conference based on highlighting of a transcript of the conference, which is determined based on comparison of sentence vectors for strings of the transcript.’);
“calculates a similarity based on the vectorized chunk data” (Huang: par. 0090; ‘FIG. 13 is a flowchart of an example of a technique 1300 for generating a summary of a video recording of a conference based on highlighting of a transcript of the conference, which is determined based on comparison of sentence vectors for strings of the transcript.’);
“determines that the summarizing execution condition is satisfied when the similarity is lower than a second threshold value” (Donaldson: par. 0046; ‘A match may be determined if there is substantial similarity or a match within a tolerance, or may be determined based on statistical analysis, machine learning, neural networks, natural language processing, and the like.’).
Regarding claim 4 (dep. on claim 1), the combination of Donaldson in view of Huang further teaches:
“a memory that stores the summary and the plurality of texts in association with each other” (Donaldson: par. 0051; ‘A speech summary manager may cause data representing an event to be stored in an electronic calendar, which may be stored in a local or remote memory.’; Huang: par. 0088; ‘For example, the file server 414 may store and or control access to recordings of conferences and transcripts of conferences.’).
Regarding claim 5 (dep. on claim 1), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry distributes the summary and the plurality of texts used for generating the summary to a terminal apparatus used by the user, and display, on a display of the terminal apparatus, a display screen that displays at least the summary” (Donaldson: par. 0023; ‘Speech summary manager 110 may be implemented on media device 101 (as shown), mobile device 102, a server, or another device, or distributed across any combination of devices.’ ‘Speech summary manager 110 may also present summary 160 on a display, or via another user interface.’).
Regarding claim 6 (dep. on claim 5), the combination of Donaldson in view of Huang further teaches:
“wherein the display screen displays the plurality of texts used for generating the summary in association with the summary in a manner that is editable by the user” (Donaldson: par. 0033; ‘Further, user interface 234 may be used to configure speech summary manager 210, such as adding a user profile to user profile database 241, modifying rules for creating action items, correcting a word that is repeatedly misrecognized by speech recognizer 221, and the like.’).
Regarding claim 7 (dep. on claim 5), the combination of Donaldson in view of Huang further teaches:
“wherein the display screen displays the summary in a manner that allows the user to extract, copy, or edit the summary” (Donaldson: par. 0033; ‘Further, user interface 234 may be used to configure speech summary manager 210, such as adding a user profile to user profile database 241, modifying rules for creating action items, correcting a word that is repeatedly misrecognized by speech recognizer 221, and the like.’).
Regarding claim 8 (dep. on claim 5), the combination of Donaldson in view of Huang further teaches:
“wherein the speech is collected during a conference in which content data is transmitted or received between a plurality of terminal apparatuses, and the processing circuitry is configured to display the display screen superimposed on a conference screen provided for the conference” (Donaldson: par. 0022; ‘A speech session may be associated with a variety of purposes, such as, delivering an address to an audience, giving a lecture or presentation, having a discussion, meeting, debate, chat, brainstorming session, and the like.’; par. 0023; ‘Speech summary manager 110 may be implemented on media device 101 (as shown), mobile device 102, a server, or another device, or distributed across any combination of devices.’ ‘Speech summary manager 110 may also present summary 160 on a display, or via another user interface.’).
Regarding claim 9 (dep. on claim 4), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry is configured to update the summary, after a first time period has elapsed since at least a part of the plurality of texts was changed for the first time or after a second time period has elapsed since at least a part of the plurality of texts was changed for the last time” (Donaldson: par. 0033; ‘Further, user interface 234 may be used to configure speech summary manager 210, such as adding a user profile to user profile database 241, modifying rules for creating action items, correcting a word that is repeatedly misrecognized by speech recognizer 221, and the like. Still, user interface 234 may be used for other purposes.’).
Regarding claim 10 (dep. on claim 1), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry stops updating the summary, in a case where the summary has been changed by a user input” (Donaldson: par. 0023; ‘Before connecting him to the conference call, speech summary manager 110 may provide the tardy user with an option to listen to a summary 160 of what has been discussed in the conference call thus far.’).
Regarding claim 11 (dep. on claim 4), the combination of Donaldson in view of Huang further teaches:
“wherein in response to reception of a predetermined operation by the user, the processing circuitry is configured to update the summary” (Donaldson: par. 0033; ‘Further, user interface 234 may be used to configure speech summary manager 210, such as adding a user profile to user profile database 241, modifying rules for creating action items, correcting a word that is repeatedly misrecognized by speech recognizer 221, and the like. Still, user interface 234 may be used for other purposes.’).
Regarding claim 14 (dep. on claim 1), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry is further configured to cause a display device to display a single screen including a first display area displaying the summary and a second display area displaying an action item” (Huang: par. 0096; ‘At 508, the technique 500 includes displaying the text summary of the transcript as a highlighted transcript. For example, the highlighted transcript may be presented in a webpage of a conference recording website (e.g., hosted by the web server 420). For example, the highlighted transcript may be presented by transmission to a user device (e.g., the mobile device 308) for display in a client application. At 510, the technique 500 includes receiving user modifications of the highlighting for the transcript. In some implementations, a user may choose to modify the transcript highlights (e.g., extend, contract, add, or remove highlights), or leave the auto generated highlights of the text summary alone.’).
Regarding claim 15 (dep. on claim 14), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry is further configured to: select, from among the plurality of texts, a text to be used for updating the action item; and update the action item based on the selected text” (Huang: par. 0096; ‘At 508, the technique 500 includes displaying the text summary of the transcript as a highlighted transcript. For example, the highlighted transcript may be presented in a webpage of a conference recording website (e.g., hosted by the web server 420). For example, the highlighted transcript may be presented by transmission to a user device (e.g., the mobile device 308) for display in a client application. At 510, the technique 500 includes receiving user modifications of the highlighting for the transcript. In some implementations, a user may choose to modify the transcript highlights (e.g., extend, contract, add, or remove highlights), or leave the auto generated highlights of the text summary alone.’).
Regarding claim 16 (dep. on claim 14), the combination of Donaldson in view of Huang further teaches:
“wherein the processing circuitry is further configured to: select, from among the summary obtained by summarizing the plurality of texts, a summary sentence to be used for updating the action item; and update the action item based on the selected summary sentence” (Huang: par. 0136; ‘At 1232, the technique 1200 includes selecting sentences for inclusion in a summary 1234 based on the sentence ranking 1230. In some implementations, sentences with the highest rankings in the entire transcript 1210 are selected. In some implementations, sentences with the highest rankings within a selected speaker segment are selected for inclusion in the summary 1234.’).
Regarding claim 17 (dep. on claim 14), the combination of Donaldson in view of Huang further teaches:
“wherein the action item includes a jump button configured to proceed to a converted text or a summary sentence corresponding to the action item” (Huang: par. 0074; ‘An input interface may, for example, be a positional input device, such as a mouse, touchpad, touchscreen, or the like; a keyboard; or another suitable human or machine interface device.’; par. 0089; ‘Seventh, the web server 420 may present 470 the transcript summary as highlighted text on a recording web page user interface (UI) and allow a user to use the user device 422 to modify 472 the highlighting that may be used to generate summary video clips of the conference.’; A jump button is just another type of input. Given the teaching of input interface and user modified highlighting, it would be obvious to include a jump button to proceed to a converted text or sentence, e.g., highlights.).
Regarding claim 18 (dep. on claim 14), the combination of Donaldson in view of Huang further teaches:
“wherein in response to the action item edited, the processing circuitry reflects an edited content of the edited action item into a converted text or a summary sentence corresponding to the edited action item” (Huang: par. 0089; ‘Seventh, the web server 420 may present 470 the transcript summary as highlighted text on a recording web page user interface (UI) and allow a user to use the user device 422 to modify 472 the highlighting that may be used to generate summary video clips of the conference.’).
Regarding claim 19 (dep. on claim 1), the combination of Donaldson in view of Huang further teaches:
“wherein the editing includes modifying a first text included in the converted text to a second text that is different from the first text” (Huang: par. 0089; ‘Seventh, the web server 420 may present 470 the transcript summary as highlighted text on a recording web page user interface (UI) and allow a user to use the user device 422 to modify 472 the highlighting that may be used to generate summary video clips of the conference.’).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed towhose telephone number is (571)270-3191. The examiner can normally be reached 10 am - 6pm EST Monday through Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MARK . VILLENA
Examiner
Art Unit 2658
/MARK VILLENA/Examiner, Art Unit 2658