DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 5/20/2026 has been entered.
Claim Status
Claims 1-4, 6-8, 10-12 and 14-22 are currently pending for examination.
Response to Arguments
Applicant's arguments filed 5/20/2026 have been fully considered.
In response to the arguments made under “Claim Rejections - 35 USC § 102”, the amendment “use the audio thread ID to compare the audio data to a name of a user of the electronic device to identify a match; and generate a notification for the user of the electronic device in response to the match.” to claim 1, and corresponding claims 6 and 11, fails to overcome the teaching of Youel. The basis for the rejection is found in paragraphs [0019] and [0025] of Youel, which are reproduced below.
Paragraph [0019] recites:
“Alternatively, the generate function 114 and/or monitor function 112 can also detect a responsible party for the action item based on the audio content before and/or after the audio trigger. For example, in response to detecting the audio trigger "draft" in the phrase "David will draft a memo on the licensing deal," the monitor function 112 and/or generate function 114 may also determine that David is the responsible party for the action item.” and
paragraph [0025] recites:
“In some embodiments, simply speaking something like "capture action, assign to Bob Smith, action due Monday, end action" will trigger the creation of an email, meeting minutes entry, or entry in an action tracking application that includes the audio file segment of the action statement data extracted from the audio, such as date due, person assigned to, context of where action came from, etc. The assigned person can be checked against the list of participants in the conversation if desired to ensure actions are issued in the context of the conversation.”.
The cited paragraphs clearly teach that the conference controller is configured to identify a participant’s name spoken in a phrase during an audio conference and to generate an action item for that participant in response to the identification. In one embodiment, Youel discloses the controller recognizes when a first participant mentions the name “David” to perform an assigned task of drafting a memo. The controller then generates the action item based on the context of the first participant’s phrase and displays the action item to “David” via a dialog box 176.
Therefore, the rejection is sustained in view of the teachings of Youel, which clearly demonstrates the recognition of the participant’s name, thereby meeting the claimed limitations.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-4, 6-8, 10-12 and 14-22 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Youel (Pub. No.: US 2014/0219434 A1).
Regarding claim 1, Youel teaches a non-transitory machine-readable medium storing instructions which, when executed by a processor of an electronic device (abstract, Fig, 1, Fig. 6, conference controller 102 includes memories and processors to perform audio detection and action item generation.), cause the processor to:
identify a process identifier (ID) assigned to a voice conferencing session executed on the electronic device (Fig. 4, Fig. 5 and para [0040], “In this embodiment, the first defined area 164 is associated with a conference or other conversation. A first plurality of participant identifiers 168A.sub.1-168F (generally, participant identifier 168 or participant identifiers 168) is displayed in association with the first defined area 164. In one embodiment, the conference may have been initiated by, for example, the controlling participant 168A.sub.1 by clicking on a "New Conference" button 170, which may cause the first defined area 164 to be depicted. The controlling participant 168A.sub.1 may then "drag and drop" the desired participant identifiers 168 from a contacts list 172 into the first defined area 164.”. The system assigns a participant identifier (e.g., 168A -168F) to each participant upon joining the conference call. The participant identifiers are represented as computer readable “process identifiers”, such as in the form of hash strings.), the process ID being a unique identifier assigned by the electronic device to distinguish multiple processes being executed by the processor (Fig. 5, each participant identifier is unique (e.g., identifier 168A or “John” is distinguishable from identifier 168D or “David”).);
identify an audio thread using the process ID, the audio thread including audio data for the voice conferencing session (Fig. 4, Fig. 5, para [0036], “In the embodiment in FIG. 4, Participant 1/Moderator 140 speaks (block 146), at which point the audio is received by the conference bridge controller 110 and relayed to the other participants. In block 148, Participant 2 142 speaks next, and the associated audio is likewise relayed to the other participants by the conference bridge controller 110. In block 150, Participant 3 144 speaks and utters an audio trigger, which is relayed to the other participants by the conference bridge controller 110.”. The system identifies the audio threads associated with each participant. For example, the system identifies the moderator 168A has made a speech at step 146 of Fig. 4.);
identify an audio thread ID of the audio thread (Fig. 5, the system identifies a part of an audio of a participant that includes a keyword. For example, a phrase within a speech that includes a keyword);
use the audio thread ID to compare the audio data to a name of a user the electronic device to identify a match (Fig. 3, step 126, para [0019], “Alternatively, the generate function 114 and/or monitor function 112 can also detect a responsible party for the action item based on the audio content before and/or after the audio trigger. For example, in response to detecting the audio trigger "draft" in the phrase "David will draft a memo on the licensing deal," the monitor function 112 and/or generate function 114 may also determine that David is the responsible party for the action item.” and para [0025], “In some embodiments, simply speaking something like "capture action, assign to Bob Smith, action due Monday, end action" will trigger the creation of an email, meeting minutes entry, or entry in an action tracking application that includes the audio file segment of the action statement data extracted from the audio, such as date due, person assigned to, context of where action came from, etc. The assigned person can be checked against the list of participants in the conversation if desired to ensure actions are issued in the context of the conversation.”. The system determines a participant has mentioned the name of another participant (e.g., David or Bob Smith) in the phrase); and
generate a notification for the user of the electronic device in response to the match (Fig. 3 step 130, and Fig. 5, generate action item associated with the identified participant (e.g., David or Bob Smith) to perform the assigned task.).
Regarding claim 2, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the instructions, when executed by the processor, cause the processor to:
transcribe the audio data into text (Fig. 5, para [0022], “The monitor function 112 may monitor the audio content of the conversation in a number of ways. For example, one or more microphones associated with each physical meeting location and/or participant may be configured to capture each participant's speech. The captured audio can then be transmitted to a central location where it can be centrally stored for later playback. The audio can also be analyzed by the monitor function 112 and other functions. The monitor function 112 may be capable of performing speech-to-text conversion of the audio content, such that the monitor function 112 is able to interpret the audio content. In other embodiments, the monitor function 112 can use phonetic speech recognition methods to convert the audio content into a format that the monitor function 112 can analyze.”); and
compare the text to the name of the user to identify the match (para [0025], “In some embodiments, simply speaking something like "capture action, assign to Bob Smith, action due Monday, end action" will trigger the creation of an email, meeting minutes entry, or entry in an action tracking application that includes the audio file segment of the action statement data extracted from the audio, such as date due, person assigned to, context of where action came from, etc. The assigned person can be checked against the list of participants in the conversation if desired to ensure actions are issued in the context of the conversation.”).
Regarding claim 3, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the instructions, when executed by the processor, cause the processor to:
sample data from a plurality of threads associated with the process ID (Fig. 4 – Fig. 5, the system samples or captures a plurality of audio speeches from each participant);
determine that a thread of the plurality of threads comprises the audio data based on the sampling (Fig. 5, the system determines whether any participant has made an audio trigger.); and
identify the thread of the plurality of threads as the audio thread based on the determination (Fig. 4, the system determines participant 2’s speech include an audio trigger at step 152 and displays the speech as an action item in Fig. 5).
Regarding claim 4, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the notification comprises a visual notification presented on a display of the electronic device (Fig. 5 displays the action item 174).
Regarding claim 6, Youel teaches a method, comprising:
identifying process identifiers (IDs) assigned by an electronic device for multiple voice conferencing sessions being executed simultaneously on the electronic device (Fig. 4, Fig. 5, and para [0040], “In this embodiment, the first defined area 164 is associated with a conference or other conversation. A first plurality of participant identifiers 168A.sub.1-168F (generally, participant identifier 168 or participant identifiers 168) is displayed in association with the first defined area 164. In one embodiment, the conference may have been initiated by, for example, the controlling participant 168A.sub.1 by clicking on a "New Conference" button 170, which may cause the first defined area 164 to be depicted. The controlling participant 168A.sub.1 may then "drag and drop" the desired participant identifiers 168 from a contacts list 172 into the first defined area 164.”. The system assigns a participant identifier (e.g., 168A -168F) to each participant upon joining the conference call. The participant identifiers are represented as computer readable “process identifiers”, such as in the form of hash strings. The conference call includes multiple voice conference sessions, wherein each voice conference session is associated with each participant’s involvement and/or speeches made from the time of that participant joins to the time of that participant leaves.), the process IDs being unique identifiers assigned by the electronic device to distinguish multiple processes being executed on the electronic device (Fig. 5, participant identifier is unique (e.g., identifier 168A or “John” is distinguishable from identifier 168D or “David”).);
sampling data from threads associated with the process IDs (Fig. 4 – Fig. 5, the system samples or captures a plurality of audio speeches from each participant);
determining that the threads include audio data for the voice conferencing sessions based on the sampling;
identifying audio thread IDs for the threads (Fig. 3, step 124, Fig. 4 -Fig. 5, para [0036], “In the embodiment in FIG. 4, Participant 1/Moderator 140 speaks (block 146), at which point the audio is received by the conference bridge controller 110 and relayed to the other participants. In block 148, Participant 2 142 speaks next, and the associated audio is likewise relayed to the other participants by the conference bridge controller 110. In block 150, Participant 3 144 speaks and utters an audio trigger, which is relayed to the other participants by the conference bridge controller 110.”. The system samples each participant’s audio to detect predefined audio triggers. The audio thread IDs are labeled as speeches 146, 148, 150 as shown in Fig. 4);
using the audio thread IDs to detect a name of a user of the electronic device in the audio data streams; and
cuing the user of the electronic device to provide input to one of the voice conferencing sessions based on the detection (Fig. 3, Fig. 5, action item 174, para [0019], “Alternatively, the generate function 114 and/or monitor function 112 can also detect a responsible party for the action item based on the audio content before and/or after the audio trigger. For example, in response to detecting the audio trigger "draft" in the phrase "David will draft a memo on the licensing deal," the monitor function 112 and/or generate function 114 may also determine that David is the responsible party for the action item.” and para [0025], “In some embodiments, simply speaking something like "capture action, assign to Bob Smith, action due Monday, end action" will trigger the creation of an email, meeting minutes entry, or entry in an action tracking application that includes the audio file segment of the action statement data extracted from the audio, such as date due, person assigned to, context of where action came from, etc. The assigned person can be checked against the list of participants in the conversation if desired to ensure actions are issued in the context of the conversation.”. The system determines a participant has mentioned the name of another participant (e.g., David or Bob Smith) in the phrase and creates an action item to the mentioned participant to confirm or cancel.).
Regarding claim 7, Youel teaches a method of claim 6, wherein a first process ID of the process IDs is associated with a first voice conferencing application executed by the electronic device, and a second process ID of the process IDs is associated with a second voice conferencing application executed by the electronic device, wherein the second voice conferencing application is different from the first voice conferencing application (Fig. 1, Fig. 3, Fig. 5, each participant has a computer 106 that run its own conferencing application.).
Regarding claim 8, recites a limitation that is similar to claim 2. Therefore, it is rejection for the same reason.
Regarding claim 10, recites a limitation that is similar to claim 4. Therefore, it is rejection for the same reason.
Regarding claim 11, Youel teaches an electronic device, comprising:
a memory; and
a processor communicatively coupled to the memory, wherein the processor is to:
identify a process identifier (ID) assigned by the electronic device to a voice conferencing session executed on the electronic device, the process ID being a unique identifier assigned by the electronic device to distinguish multiple processes being executed by the processor;
identify a thread associated with the process ID that includes audio data for the voice conferencing session;
determine an audio thread ID for the thread and use the audio thread ID to access the audio data;
transcribe the audio data into text;
compare the text to a keyword stored in the memory;
determine that a user of the electronic device is being cued in the voice conferencing session based on the comparison; and
generate a notification for the user that input is requested in the voice conferencing session (The rejection is similar to the combination of claims 1 and 2.).
Regarding claim 12, recites a limitation that is similar to claim 3. Therefore, it is rejection for the same reason.
Regarding claim 14, recites a limitation that is similar to claim 4. Therefore, it is rejection for the same reason.
Regarding claim 15, Youel teaches the electronic device of claim 11, comprising a speaker coupled to the processor, wherein the notification comprises an audio feed of the voice conferencing session that is emitted from the speaker (para [0032], “With continuing reference to FIG. 3, notification of the participant (block 130), moderator or other authorized person of the generation of the action item (block 128) may comprise an audio notification in some embodiments. In some embodiments, the audio notification can include an alert sound, or can include a recorded or generated speech prompt informing one or more participants or other authorized persons that the action item has been generated.”. The system generates audio notification to inform the participant about the action item.).
Regarding claim 16, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the voice conferencing session comprises at least one of a voice call or a video conference (abstract “Embodiments include methods, apparatuses, and systems for generating an action item in response to a detected audio trigger during a conversation. Embodiments relate to generation of one or more action items in response to detection of an audio trigger, such as a spoken command, keyword, audio tone or other indicator, which is detected during a conversation, such as an audio or video conference or peer-to-peer conversation.”).
Regarding claim 17, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the user of the electronic device is a participant in the voice conferencing session (Fig. 5, para [0039], “FIG. 5 illustrates an exemplary software user interface 162 for generating an action item according to one embodiment and that can be employed in the conference system 100 of FIG. 1 and the peer-to-peer system 117 of FIG. 2.”. Each participant is a user of the conference controller 102 to participate in the conference.).
Regarding claim 18, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the voice conferencing session is one of a plurality of voice conferencing sessions executed simultaneously on the electronic device (Fig. 4 – Fig. 5, multiple participates are simultaneously attending the conference call.).
Regarding claim 19, Youel teaches the method of claim 6, wherein cuing the user comprises audibly cuing the user via a speaker of the electron device (para [0032], “With continuing reference to FIG. 3, notification of the participant (block 130), moderator or other authorized person of the generation of the action item (block 128) may comprise an audio notification in some embodiments. In some embodiments, the audio notification can include an alert sound, or can include a recorded or generated speech prompt informing one or more participants or other authorized persons that the action item has been generated.”. The system generates audio notification to inform the participant about the action item.).
Regarding claim 20, Youel teaches the electronic device of claim 11, wherein the electronic device is at least one of a desktop computer, a laptop computer, an all-in-one computer, or a smartphone (Fig. 1, computer 106), and the voice conferencing session is at least one of a voice call or a video conference (abstract, “Embodiments include methods, apparatuses, and systems for generating an action item in response to a detected audio trigger during a conversation. Embodiments relate to generation of one or more action items in response to detection of an audio trigger, such as a spoken command, keyword, audio tone or other indicator, which is detected during a conversation, such as an audio or video conference or peer-to-peer conversation.”).
Regarding claim 21, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the notification comprises a pop-up window presented on a display of the electronic device that minimizes other content presented on the display (Fig. 5, the pop-up window 174 minimizes the background content. For example, the pop-up window reduces the area of the white background.).
Regarding claim 22, Youel teaches the non-transitory machine-readable medium of claim 1, wherein the multiple processes being executed by the processor are processes generated by execution of multiple voice conferencing applications on the electronic device (Fig. 1 and Fig. 4, shows the conference controller 102 processes multiple audio data from a plurality of computers 106 joining the conference).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZHEN Y WU whose telephone number is (571)272-5711. The examiner can normally be reached Monday-Friday, 10AM-6PM, EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Quan-Zhen Wang can be reached at 571-272-3114. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZHEN Y WU/Primary Examiner, Art Unit 2685