DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because:
reference character “240” has been used to designate both “origination sources” and “information datastores” in Figure 2
reference character “252” has been used to designate both “locations” in Figure 2 and “environment” in Figure 3.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference characters not mentioned in the description:
“238” in Figure 2
“410” in Figure 4
“700” in Figure 7.
Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference characters in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
In paragraph 0050, lines 3-4, “signal data 172” should read “signal data 174”.
In paragraph 0091, line 7, “analytical model 22” should read “analytical model 222”.
In paragraph 0097, line 3, “analytical model 22” should read “analytical model 222”.
In paragraph 0117, line 1, “an alterative root node” should read “an alternative root node”.
In paragraph 0127, line 4, “proceed to step 814” should read “proceed to step 812” to be consistent with the drawings.
In paragraph 0127, line 4, “proceed back to step 812” should read “proceed back to step 808” to be consistent with the drawings.
Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1 – 2, 4 and 9 – 11 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Cartwright et al. (US Patent No. 10,522,151), hereinafter Cartwright.
Regarding claim 1, Cartwright discloses a system, comprising:
at least one processor (Column 13, lines 27-33, "The control system may include at least one of a general purpose single- or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components.");
and at least one memory that stores executable instructions that, when executed by the at least one processor, facilitates performance of operations (Column 20, lines 25-27, "Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on non-transitory media."), comprising:
based on sensed signals corresponding to a conversation, wherein the sensed signals at least partially correspond to voice soundwaves comprised in the conversation, recording the sensed signals as signal data (Column 36, lines 8-16, "The present disclosure includes methods and devices for recording, analyzing and playing back teleconference audio data such that the teleconference audio data presented during playback may be substantially different from what would have been heard by conference participants during the original teleconference and/or what would have been recorded during the original teleconference by a recording device such as the teleconference recording device 2 shown in FIG. 1A."; Column 49, lines 7-11, "In this implementation, each of the cables 712a-712d convey an individual stream of audio data from a corresponding one of the microphones 715a-715d to a recording device 720, which is located under the conference participant table 705 in this instance."; Recording teleconference audio data with microphones reads on recording sensed voice soundwaves in a conversation as signal data.);
analyzing the signal data based on classification data representative of at least one specification determined to be applicable to the conversation (Column 2, lines 65-67, "According to some examples, analyzing the audio data may involve identifying speech corresponding to individual conference participants."; Column 38, line 66 - Column 39, line 4, "In this example, the teleconferencing apparatus 200 also sends conference metadata 210 to the conference recording database 3. The conference metadata 210 may, for example, include data regarding individual conference participants, such as conference participant name, conference participant location, etc."; Column 45, lines 54-61, "In the example shown in FIG. 5, the joint analysis module includes an overview module 520. In this implementation, the overview module 520 receives the conference metadata 210 as well as data from the conference database 308. The conference metadata 210 may include data regarding individual conference participants, such as conference participant name and conference participant location, data indicating the time and date of a conference, etc."); Analyzing audio data to identify speech corresponding to individual conference participants reads on analyzing the signal data based on classification data, and conference metadata including data regarding individual conference participants, such as conference participant name and conference participant location, reads on a specification determined to be applicable to the conversation.);
and based on the analyzing, mapping the signal data to an origination source representative of an origin of the sensed signals (Column 2, lines 65-67, "According to some examples, analyzing the audio data may involve identifying speech corresponding to individual conference participants."; Identifying speech corresponding to individual conference participants reads on mapping the signal data to an origination source representative of an origin of the sensed signals.).
Regarding claim 2, Cartwright discloses the system as claimed in claim 1.
Cartwright further discloses:
wherein the origination source is a first origination source (Column 2, lines 65-67, "According to some examples, analyzing the audio data may involve identifying speech corresponding to individual conference participants."; Identifying speech corresponding to individual conference participants reads on mapping the signal data to an origination source representative of an origin of the sensed signals, where the origination source is a first origination source.),
and wherein the operations further comprise: identifying a pair of locations, associated with a time period applicable to the conversation, relative to one another, wherein the pair of locations correspond to the first origination source and a second origination source (Column 48, lines 13-28, "In the example shown in FIG. 6, each row 615 of the graphical user interface 606 corresponds to a particular conference participant. In this implementation, the graphical user interface 606 indicates conference participant information 620, which may include a conference participant name, conference participant location, conference participant photograph, etc. In this example, waveforms 625, corresponding to instances of the speech of each conference participant, are also shown the graphical user interface 606. The display device 610 may, for example, display the waveforms 625 according to instructions from playback control module 605. Such instructions may, for example be based on visualization data 410D-403D that is included in the analysis results 301C-303C. In some examples, a user may be able to change the scale of the graphical user interface 606, according to a desired time interval of the conference to be represented."; A graphical user interface indicating conference participant information including a conference participant name and conference participant location, where each row of the graphical user interface corresponds to a particular conference participant and the user can change the scale of the graphical user interface to represent a desired time interval of the conference, reads on identifying a pair of locations, associated with a time period applicable to the conversation, relative to one another, where the pair of locations correspond to the first origination source and a second origination source.).
Regarding claim 4, Cartwright discloses the system as claimed in claim 1.
Cartwright further discloses:
wherein the operations further comprise: determining that the sensed signals involved in the conversation are from multiple distinct environments, and wherein the recording of the sensed signals as the signal data comprises recording the sensed signals with metadata representing that the sensed signals are from the multiple distinct environments involved in the conversation (Column 45, lines 54-61, "In the example shown in FIG. 5, the joint analysis module includes an overview module 520. In this implementation, the overview module 520 receives the conference metadata 210 as well as data from the conference database 308. The conference metadata 210 may include data regarding individual conference participants, such as conference participant name and conference participant location, data indicating the time and date of a conference, etc."; Column 48, lines 13-28, "In the example shown in FIG. 6, each row 615 of the graphical user interface 606 corresponds to a particular conference participant. In this implementation, the graphical user interface 606 indicates conference participant information 620, which may include a conference participant name, conference participant location, conference participant photograph, etc. In this example, waveforms 625, corresponding to instances of the speech of each conference participant, are also shown the graphical user interface 606. The display device 610 may, for example, display the waveforms 625 according to instructions from playback control module 605. Such instructions may, for example be based on visualization data 410D-403D that is included in the analysis results 301C-303C. In some examples, a user may be able to change the scale of the graphical user interface 606, according to a desired time interval of the conference to be represented."; A graphical user interface indicating conference participant information including a conference participant name and conference participant location, where each row of the graphical user interface corresponds to a particular conference participant, reads on determining that the sensed signals involved in the conversation are from multiple distinct environments, and wherein the recording of the sensed signals as the signal data comprises recording the sensed signals with metadata representing that the sensed signals are from the multiple distinct environments involved in the conversation, where different participant locations read on distinct environments involved in the conversation.).
Regarding claim 9, Cartwright discloses a method, comprising:
identifying, by a system comprising at least one processor, recorded signal data corresponding to voice soundwaves of a conversation (Column 13, lines 27-33, "The control system may include at least one of a general purpose single- or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components."; Column 36, lines 8-16, "The present disclosure includes methods and devices for recording, analyzing and playing back teleconference audio data such that the teleconference audio data presented during playback may be substantially different from what would have been heard by conference participants during the original teleconference and/or what would have been recorded during the original teleconference by a recording device such as the teleconference recording device 2 shown in FIG. 1A."; Column 49, lines 7-11, "In this implementation, each of the cables 712a-712d convey an individual stream of audio data from a corresponding one of the microphones 715a-715d to a recording device 720, which is located under the conference participant table 705 in this instance."; Recording teleconference audio data with microphones reads on identifying recorded signal data corresponding to voice soundwaves of a conversation.);
assigning, by the system, respective origination sources corresponding to the voice soundwaves (Column 2, lines 65-67, "According to some examples, analyzing the audio data may involve identifying speech corresponding to individual conference participants."; Identifying speech corresponding to individual conference participants reads on assigning respective origination sources corresponding to the voice soundwaves.);
and generating, by the system, environmental map data comprising location data representative of respective locations of the respective origination sources relative to one another (Column 48, lines 13-28, "In the example shown in FIG. 6, each row 615 of the graphical user interface 606 corresponds to a particular conference participant. In this implementation, the graphical user interface 606 indicates conference participant information 620, which may include a conference participant name, conference participant location, conference participant photograph, etc. In this example, waveforms 625, corresponding to instances of the speech of each conference participant, are also shown the graphical user interface 606. The display device 610 may, for example, display the waveforms 625 according to instructions from playback control module 605. Such instructions may, for example be based on visualization data 410D-403D that is included in the analysis results 301C-303C. In some examples, a user may be able to change the scale of the graphical user interface 606, according to a desired time interval of the conference to be represented."; A graphical user interface indicating conference participant information including a conference participant name and conference participant location, where each row of the graphical user interface corresponds to a particular conference participant, reads on generating environmental map data comprising location data representative of respective locations of the respective origination sources relative to one another.).
Regarding claim 10, Cartwright discloses the method as claimed in claim 9.
Cartwright further discloses:
wherein the environmental map data further comprises environment data representative at least a pair of environments involved in the conversation (Column 48, lines 13-28, "In the example shown in FIG. 6, each row 615 of the graphical user interface 606 corresponds to a particular conference participant. In this implementation, the graphical user interface 606 indicates conference participant information 620, which may include a conference participant name, conference participant location, conference participant photograph, etc. In this example, waveforms 625, corresponding to instances of the speech of each conference participant, are also shown the graphical user interface 606. The display device 610 may, for example, display the waveforms 625 according to instructions from playback control module 605. Such instructions may, for example be based on visualization data 410D-403D that is included in the analysis results 301C-303C. In some examples, a user may be able to change the scale of the graphical user interface 606, according to a desired time interval of the conference to be represented."; A graphical user interface indicating conference participant information including a conference participant name and conference participant location, where each row of the graphical user interface corresponds to a particular conference participant, reads on the environmental map data comprising environment data representative at least a pair of environments involved in the conversation.),
wherein a first environment of the pair of environments comprises a first device that is connected, via a network, to a second device, wherein a second environment of the pair of environments comprises the second device that recorded the recorded signal data, and wherein the first environment is distinct from the second environment (Column 4, line 58 - Column 5, line 6, "In another aspect the present invention provides a system for sharing data streams between meeting devices, the system comprising at least two meeting devices, at least two base units, at least one server, a bus, hosted by the at least one server, a meeting identifier, at least two local networks separated from a global network, the at least two meeting devices being each connected to a different base unit over the at least one local network, the at least two base units being connected to the bus over the global network and, wherein base units belonging to the same meeting share one meeting identifier."; At least two meeting devices connected over a network reads on a first device that is connected, via a network, to a second device, wherein a second environment of the pair of environments comprises the second device that recorded the recorded signal data, and wherein the first environment is distinct from the second environment.).
Regarding claim 11, Cartwright discloses the method as claimed in claim 9.
Cartwright further discloses:
further comprising: obtaining, by the system, classification data defining specifications corresponding to the conversation, the specifications comprising participant data representative of participants in the conversation (Column 2, lines 65-67, "According to some examples, analyzing the audio data may involve identifying speech corresponding to individual conference participants."; Column 38, line 66 - Column 39, line 4, "In this example, the teleconferencing apparatus 200 also sends conference metadata 210 to the conference recording database 3. The conference metadata 210 may, for example, include data regarding individual conference participants, such as conference participant name, conference participant location, etc."; Column 45, lines 54-61, "In the example shown in FIG. 5, the joint analysis module includes an overview module 520. In this implementation, the overview module 520 receives the conference metadata 210 as well as data from the conference database 308. The conference metadata 210 may include data regarding individual conference participants, such as conference participant name and conference participant location, data indicating the time and date of a conference, etc."; Receiving conference metadata including data regarding individual conference participants, such as conference participant name and conference participant location, reads on obtaining classification data defining specifications corresponding to the conversation, the specifications comprising participant data representative of participants in the conversation.);
and assigning, by the system, the respective origination sources based on the classification data (Column 2, lines 65-67, "According to some examples, analyzing the audio data may involve identifying speech corresponding to individual conference participants."; Identifying speech corresponding to individual conference participants reads on assigning the respective origination sources based on the classification data).
Claims 15 and 19 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Koum et al. (US Patent No. 9,998,602), hereinafter Koum.
Regarding claim 15, Koum discloses a method, comprising:
detecting, by a computing device comprising at least one processor, voice soundwaves of a conversation, using a sensor at an external surface of the computing device (Column 3, lines 30-32, "FIG. 1 is a block diagram of a system for facilitating recorded voice communications with real-time status notifications, according to some embodiments of the invention."; Column 5, lines 56-59, "Device 202 includes a touch-screen display, one or more microphones and one or more speakers, and may include other components not referenced herein."; Column 13, lines 36-37, "Device 602 comprises one or more processing units or processors"; Recording voice communications with a device that includes a microphone reads on detecting voice soundwaves of a conversation using a sensor at an external surface of the computing device.);
recording, by the computing device, signal data based on the voice soundwaves, using a microphone at the external surface or another external surface of the computing device (Column 3, lines 30-32, "FIG. 1 is a block diagram of a system for facilitating recorded voice communications with real-time status notifications, according to some embodiments of the invention."; Column 5, lines 56-59, "Device 202 includes a touch-screen display, one or more microphones and one or more speakers, and may include other components not referenced herein."; Column 13, lines 36-37, "Device 602 comprises one or more processing units or processors"; Recording voice communications with a device that includes a microphone reads on recording signal data based on the voice soundwaves, using a microphone at the external surface or another external surface of the computing device.);
and in response to the recording, generating, by the computing device, a notification that the recording has begun or is in progress (Column 3, lines 4-11, "In different embodiments, one or more of multiple complementary features are implemented, such as one-touch voice recording, dynamic real-time notification to a communication partner of the commencement of a voice recording, reliable delivery of the recording to the partner, real-time notification of playback of the recording by the communication partner and automatic selection of a output device for playing an audio recording."; Providing a dynamic real-time notification of the commencement of a voice recording reads on generating a notification that the recording has begun or is in progress in response to the recording.).
Regarding claim 19, Koum discloses the method as claimed in claim 15.
Koum further discloses:
further comprising: in response to manual contact of the sensor or another sensor at the computing device while the computing device is in a dormant state, triggering, by the computing device, starting of the recording (Column 7, lines 18-20, “Specifically, by pressing and holding control 322, a microphone in device 302 is activated and begins recording."; Activating a microphone in a device to begin recording by pressing a control reads on triggering, by the computing device, starting of the recording in response to manual contact of the sensor or another sensor at the computing device while the computing device is in a dormant state.).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3, 6 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Cartwright in view of Kim et al. (US Patent No. 10,438,595), hereinafter Kim.
Regarding claim 3, Cartwright discloses the system as claimed in claim 1, but does not specifically disclose: wherein the operations further comprise: determining that the voice soundwaves comprise a previously unidentified voiceprint for which a corresponding voiceprint has not previously been stored in the classification data.
Kim teaches:
determining that the voice soundwaves comprise a previously unidentified voiceprint for which a corresponding voiceprint has not previously been stored in the classification data (Column 10, line 62 - Column 11, line 5, "For example, block 508 can include comparing the audio input received at block 302 with all alternate speaker profiles to determine if the speaker of the audio input matches an existing speaker profile. If it is determined at block 508 that the speaker of the audio input matches one of the existing alternate speaker profiles, the audio input can be added to that alternate speaker profile. If it is instead determined at block 508 that the speaker of the audio input does not match one of the existing alternate speaker profiles, a new alternate speaker profile can be generated using the audio input."; Determining that a speaker of audio input does not match one of the existing speaker profiles reads on determining that the voice soundwaves comprise a previously unidentified voiceprint for which a corresponding voiceprint has not previously been stored in the classification data.).
Kim is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cartwright to incorporate the teachings of Kim to determine that a speaker of audio input does not match one of the existing speaker profiles. Doing so would allow for not triggering the operation of a virtual assistant in response to determining that the speaker of the user speech is not a predetermined user (Kim; Column 2, lines 35-49).
Regarding claim 6, Cartwright discloses the system as claimed in claim 1, but does not specifically disclose: wherein the operations further comprise: determining, from the classification data, different voiceprints matching different origination sources, comprising the origination source, during the recording of the signal data.
Kim teaches:
determining, from the classification data, different voiceprints matching different origination sources, comprising the origination source, during the recording of the signal data (Column 7, lines 40-43, "Speaker recognition can be performed using the voice prints of a speaker profile by comparing an audio input containing user speech with the voice prints in the speaker profile."; Performing speaker recognition by comparing an audio input containing user speech with the voice prints in a speaker profile reads on determining different voiceprints matching different origination sources during the recording of the signal data.).
Kim is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cartwright to incorporate the teachings of Kim to perform speaker recognition by comparing an audio input containing user speech with the voice prints in a speaker profile. Doing so would allow for not triggering the operation of a virtual assistant in response to determining that the speaker of the user speech is not a predetermined user (Kim; Column 2, lines 35-49).
Regarding claim 8, Cartwright discloses the system as claimed in claim 1, but does not specifically disclose: wherein the mapping comprises: in response to analyzing the signal data, and based on the classification data, associating previously identified voiceprint data, representative of at least one voiceprint previously stored in the classification data, with voiceprint data determined from the analyzing of the signal data, and tagging the signal data to correspond to the voiceprint previously stored in the classification data.
Kim teaches:
in response to analyzing the signal data, and based on the classification data, associating previously identified voiceprint data, representative of at least one voiceprint previously stored in the classification data, with voiceprint data determined from the analyzing of the signal data, and tagging the signal data to correspond to the voiceprint previously stored in the classification data (Column 7, lines 31-46, "At block 305, the user device can generate a speaker profile, selectively perform speaker recognition using the speaker profile, and selectively activate the virtual assistant in response to positively identifying the speaker using speaker recognition. In some examples, the speaker profile can generally include one or more voice prints generated from an audio recording of a speaker's voice. The voice prints can be generated using any desired speech recognition technique, such as by generating i-vectors to represent speaker utterances. Speaker recognition can be performed using the voice prints of a speaker profile by comparing an audio input containing user speech with the voice prints in the speaker profile. As discussed in greater detail below, block 305 can include blocks 306, 308, 310, and 312 for allowing the user device to operate in multiple modes of operation based on a status of the speaker profile."; Performing speaker recognition by comparing an audio input containing user speech with the voice prints in a speaker profile, where a virtual assistant is selectively activated in response to positively identifying the speaker using speaker recognition, reads on associating previously identified voiceprint data, representative of at least one voiceprint previously stored in the classification data, with voiceprint data determined from the analyzing of the signal data in response to analyzing the signal data and based on the classification data, and tagging the signal data to correspond to the voiceprint previously stored in the classification data.).
Kim is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cartwright to incorporate the teachings of Kim to perform speaker recognition by comparing an audio input containing user speech with the voice prints in a speaker profile, where a virtual assistant is selectively activated in response to positively identifying the speaker using speaker recognition. Doing so would allow for not triggering the operation of a virtual assistant in response to determining that the speaker of the user speech is not a predetermined user (Kim; Column 2, lines 35-49).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Cartwright in view of Treglia (US Patent Application Publication No. 2012/0245936).
Regarding claim 5, Cartwright discloses the system as claimed in claim 1, but does not specifically disclose: wherein the operations further comprise: analyzing the signal data, resulting in a determination that the sensed signals are from multiple recording devices that recorded the conversation.
Treglia teaches:
analyzing the signal data, resulting in a determination that the sensed signals are from multiple recording devices that recorded the conversation (Paragraph 0108, lines 1-6, "In one embodiment, the audio recordings of the same conversation from separate devices is matched by using acoustic fingerprinting technology, such as for example SoundPrint or similar technology. Acoustic fingerprinting technology is capable of quickly matching different recordings of the same conversation by using an algorithm."; Paragraph 0109, lines 1-8, "In one embodiment, the identification of two or more devices recording the same conversation, using one of the techniques described above or other technology capable of making such an identification, is performed in real time or near real time (i.e., while the conversation is being recorded) by communication with a coordinating device, such as one of the devices or another device or server, using any wired or wireless technology known in the art."; Using acoustic fingerprinting technology to identify that two or more devices are recording the same conversation reads on analyzing the signal data resulting in a determination that the sensed signals are from multiple recording devices that recorded the conversation.).
Treglia is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cartwright to incorporate the teachings of Treglia to use acoustic fingerprinting technology to identify that two or more devices are recording the same conversation. Doing so would allow for increasing the accuracy of the system to differentiate and locate each speaker (Treglia; Paragraph 0102, lines 1-5).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Cartwright in view of Smith et al. (US Patent No. 11,120,794), hereinafter Smith.
Regarding claim 7, Cartwright discloses the system as claimed in claim 1, but does not specifically disclose: wherein the operations further comprise: in response to sensing a signal via an external sensor of a computing device in a dormant state, starting the recording of the signal data.
Smith teaches:
in response to sensing a signal via an external sensor of a computing device in a dormant state, starting the recording of the signal data (Column 4, lines 43-52, "In some embodiments, for example, a user may speak a wake word and a voice utterance (e.g., a command) in the vicinity of multiple NMDs. Two or more of the NMDs may detect sound based on the user's speech and identify the wake word therein. Each of these NMDs may then transition from an inactive state to an active state. In the inactive state, the NMD listens for a wake word in detected sound but does not transmit any data based on the detected sound. Once transitioned to the active state, the NMD is readied to capture sound data corresponding to the detected sound."; A network microphone device transitioning from an inactive state to an active state in response to detecting sound based on a user's speech and identifying a wake word, where the network microphone device captures sound data in the active state, reads on starting the recording of the signal data in response to sensing a signal via an external sensor of a computing device in a dormant state.).
Smith is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cartwright to incorporate the teachings of Smith to implement a device transitioning from an inactive state to an active state in response to detecting sound based on a user's speech and identifying a wake word, where the device captures sound data in the active state. Doing so would allow for multiple devices coordinating responsibility for voice control interactions to deliver an improved user experience (Smith; Column 4, line 43 - Column 5, line 6).
Claims 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Cartwright in view of Wengrovitz et al. (US Patent Application Publication No. 2007/0133437), hereinafter Wengrovitz
Regarding claim 13, Cartwright discloses the method as claimed in claim 9, but does not specifically disclose: further comprising: receiving, by the system, query data comprising a query requesting an origination source, of the respective origination sources, corresponding to a specified time period corresponding to occurrence of the conversation; and based on the assigning of the respective origination sources, generating, by the system, a response to the query.
Wengrovitz teaches:
receiving, by the system, query data comprising a query requesting an origination source, of the respective origination sources, corresponding to a specified time period corresponding to occurrence of the conversation (Paragraph 0032, lines 1-4, "FIG. 3 is an architectural overview of a communications network 300 where multi-party conferencing and use of who is speaking data is supported according to an embodiment of the present invention."; Paragraph 0094, lines 1-8, 'FIG. 8 is a process flow chart 800 illustrating steps for preparing and submitting an information search of conference archives for WIS-related information according to an embodiment of the present invention. At step 801, a user invokes a search engine interface adapted to search conference archives using any one or a combination of keywords, phrasing, temporal data, WIS data, and presence information."; Paragraph 0096, lines 1-3, "At step 803, the user may specify conference event parameters such as conference titles, conference dates, and time windows."; A user submitting an information search of conference archives for who-is-speaking information to invoke a search engine interface adapted to search conference archives using temporal data and who-is-speaking data reads on receiving query data comprising a query requesting an origination source, of the respective origination sources, corresponding to a specified time period corresponding to occurrence of the conversation.);
and based on the assigning of the respective origination sources, generating, by the system, a response to the query (Paragraph 0094, lines 1-8, 'FIG. 8 is a process flow chart 800 illustrating steps for preparing and submitting an information search of conference archives for WIS-related information according to an embodiment of the present invention. At step 801, a user invokes a search engine interface adapted to search conference archives using any one or a combination of keywords, phrasing, temporal data, WIS data, and presence information."; Paragraph 0097, lines 10-13, "At step 807, the user may submit the query to the third-party node hosting the search. Results returned may vary according to the goal of the information search."; A third-party node returning results to a query for searching conference archives for who-is-speaking information reads on generating a response to the query based on the assigning of the respective origination sources.).
Wengrovitz is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cartwright to incorporate the teachings of Wengrovitz to provide for a user submitting an information search of conference archives for who-is-speaking information to invoke a search engine interface adapted to search conference archives using temporal data and who-is-speaking data and a third-party node returning results to a query for searching conference archives for who-is-speaking information. Doing so would allow for real-time identification of who is speaking information and presence information resulting from a multi-party conference session practiced over a network (Wengrovitz; Paragraph 0031, lines 1-7).
Regarding claim 14, Cartwright in view of Wengrovitz discloses the method as claimed in claim 13.
Wengrovitz further teaches:
wherein the generating of the response to the query comprises generating the response based on assignment data from a previous assignment of one or more origination sources corresponding to a prior conversation from a time prior to the conversation (Paragraph 0017, lines 1-17, "According to another aspect of the present invention, an audio content transcription and annotation system is provided for rendering annotated text transcription of live or recorded speech from a multiparty conference session enabled by a conference bridging switch, software, or a combination thereof having multiple conference input channels and for annotating the transcribed text files with who-is-speaking data. The system includes an input port for receiving the audio content, a time synchronization module for recording temporal offsets of changes in a channel activity signal relevant to conference session run time, a channel to speaker association module, and a text annotation engine. In a preferred embodiment, the transcribed text files are annotated according to indication of signal changes over time with relevance to audible words, phrases or segments of the content found within the scope of time periods existing in between the signal changes."; Rendering annotated text transcription of live or recorded speech from a multiparty conference session, where the transcribed text files is annotated with who-is-speaking data, reads on assignment data from a previous assignment of one or more origination sources corresponding to a prior conversation from a time prior to the conversation.).
Wengrovitz is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cartwright in view of Wengrovitz to further incorporate the teachings of Wengrovitz to render annotated text transcription of live or recorded speech from a multiparty conference session, where the transcribed text files is annotated with who-is-speaking data. Doing so would allow for real-time identification of who is speaking information and presence information resulting from a multi-party conference session practiced over a network (Wengrovitz; Paragraph 0031, lines 1-7).
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Koum in view of Fransson et al. (US Patent No. 12,088,909), hereinafter Fransson.
Regarding claim 16, Koum discloses the method as claimed in claim 15, but does not specifically disclose: wherein the generating of the notification comprises activating a light at the computing device.
Fransson teaches:
wherein the generating of the notification comprises activating a light at the computing device (Column 1, lines 32-35, "In a general aspect, a device may include a recording device and a warning light configured to be activated in conjunction with recording operations of the recording device."; A warning light configured to be activated in conjunction with recording operations of a recording device reads on activating a light at the computing device for a recording notification.).
Fransson is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Koum to incorporate the teachings of Fransson to implement a warning light configured to be activated in conjunction with recording operations of a recording device. Doing so would allow for alerting bystanders that an audio recording is actively being captured when the corresponding recording device may not be visible or otherwise detectable to the bystanders (Fransson; Column 2, lines 42-49).
Claims 17 – 18 are rejected under 35 U.S.C. 103 as being unpatentable over Koum in view of Smith.
Regarding claim 17, Koum discloses the method as claimed in claim 15, but does not specifically disclose: wherein the detecting and the recording are performed while the computing device is in a dormant state.
Smith teaches:
wherein the detecting and the recording are performed while the computing device is in a dormant state (Column 4, lines 43-52, "In some embodiments, for example, a user may speak a wake word and a voice utterance (e.g., a command) in the vicinity of multiple NMDs. Two or more of the NMDs may detect sound based on the user's speech and identify the wake word therein. Each of these NMDs may then transition from an inactive state to an active state. In the inactive state, the NMD listens for a wake word in detected sound but does not transmit any data based on the detected sound. Once transitioned to the active state, the NMD is readied to capture sound data corresponding to the detected sound."; A network microphone device transitioning from an inactive state to an active state in response to detecting sound based on a user's speech and identifying a wake word, where the network microphone device captures sound data in the active state, reads on the detecting and the recording being performed while the computing device is in a dormant state.).
Smith is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Koum to incorporate the teachings of Smith to implement a device transitioning from an inactive state to an active state in response to detecting sound based on a user's speech and identifying a wake word, where the device captures sound data in the active state. Doing so would allow for multiple devices coordinating responsibility for voice control interactions to deliver an improved user experience (Smith; Column 4, line 43 - Column 5, line 6).
Regarding claim 18, Koum in view of Smith discloses the method as claimed in claim 17.
Smith further teaches:
wherein the dormant state of the computing device corresponds to a deactivated state or a sleep state of the computing device (Column 4, lines 43-52, "In some embodiments, for example, a user may speak a wake word and a voice utterance (e.g., a command) in the vicinity of multiple NMDs. Two or more of the NMDs may detect sound based on the user's speech and identify the wake word therein. Each of these NMDs may then transition from an inactive state to an active state. In the inactive state, the NMD listens for a wake word in detected sound but does not transmit any data based on the detected sound. Once transitioned to the active state, the NMD is readied to capture sound data corresponding to the detected sound."; A network microphone device transitioning from an inactive state to an active state in response to detecting sound based on a user's speech and identifying a wake word, where the device does not transmit any data based on the detected sound in the inactive state, reads on the dormant state of the computing device corresponds to a deactivated state of the computing device.).
Smith is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Koum in view of Smith to further incorporate the teachings of Smith to implement a device transitioning from an inactive state to an active state in response to detecting sound based on a user's speech and identifying a wake word, where the device does not transmit any data based on the detected sound in the inactive state. Doing so would allow for multiple devices coordinating responsibility for voice control interactions to deliver an improved user experience (Smith; Column 4, line 43 - Column 5, line 6).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Koum in view of Perotti (US Patent No. 11,056,117).
Regarding claim 20, Koum discloses the method as claimed in claim 15, but does not specifically disclose: further comprising: in response to the sensing the voice soundwaves corresponding to at least one of a known voiceprint or a known timestamp, initiating, by the computing device, the recording, wherein at least one of the known voiceprint or the known timestamp correspond to classification data that defines specifications corresponding to the conversation.
Perotti teaches:
in response to the sensing the voice soundwaves corresponding to at least one of a known voiceprint or a known timestamp, initiating, by the computing device, the recording, wherein at least one of the known voiceprint or the known timestamp correspond to classification data that defines specifications corresponding to the conversation (Column 7, lines 41-50, 'As described herein, the hardware processor 112 processes data, including the execution of applications stored in the memory 116. In particular, and as described below, the hardware processor 112 executes applications for performing keyword matching and voiceprint matching operations on the speech of a user, received as input via the microphone 120. Moreover, in response to the successful authentication of a user by way of the keyword matching and voiceprint matching operations, the processor may retrieve and present data in accordance with various commands from the user."; Responding to commands from the user in response to the successful authentication of a user using voiceprint matching reads on initiating the recording in response to sensing voice soundwaves corresponding to a known voiceprint).
Perotti is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Koum to incorporate the teachings of Perotti to respond to commands from the user in response to the successful authentication of a user using voiceprint matching. Doing so would allow for confirming the user's identity in response to voice commands (Perotti; Column 4, lines 3-18).
Allowable Subject Matter
Claim 12 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
The primary reason claim 12 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims is the inclusion, in all the claims, of the limitation “based on the recorded signal data, generating, by the system, graph data representative of a graph comprising nodes corresponding to the respective origination sources and edges corresponding to the respective locations of the respective origination sources relative to one another” in combination with the limitations to identify recorded signal data corresponding to voice soundwaves of a conversation, assign respective origination sources corresponding to the voice soundwaves, and generate environmental map data comprising location data representative of respective locations of the respective origination sources relative to one another.
Cartwright discloses the method as claimed in claim 9.
Doganata et al. (US Patent Application Publication No. 2013/0132138), hereinafter Doganata, discloses:
based on the recorded signal data, generating, by the system, graph data representative of a graph comprising nodes corresponding to the respective origination sources (Paragraph 0039, lines 1-9, "It is realized herein that a meeting can be modeled as a set of activities executed by various actors such as a process where textual and audio visual data are consumed or produced at different steps. In effect, a meeting is a process with a start event and an end event and a sequence of other events in between. Hence, provenance graphs may be generated for meetings applications. In a meeting provenance graph, meeting activities, data and participants are represented as nodes and causal relations are represented as edges."; Generating provenance graphs for meetings applications, where participants are represented as nodes, reads on generating graph data representative of a graph comprising nodes corresponding to the respective origination sources based on the recorded signal data.).
However, Cartwright and Doganata, individually or in combination, do not disclose the limitation “based on the recorded signal data, generating, by the system, graph data representative of a graph comprising nodes corresponding to the respective origination sources and edges corresponding to the respective locations of the respective origination sources relative to one another”.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Bradley et al. (US Patent No. 12,020,708) teaches a method for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed transcript.
Zhang et al. (US Patent No. 11,222,640) teaches a method that utilize a joint speaker location/speaker identification neural network to enhance both speaker ID and speaker location results.
Beaurepaire et al. (US Patent Application Publication No. 2022/0100796) teaches a method for processing audio data and metadata associated with the audio data to determine location data.
Jackson (US Patent Application Publication No. 2017/0287482) teaches a method for generating a transcription of multi-party communication and differentiating individual speakers of the plurality of speakers in a final transcript.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Boggs whose telephone number is (571)272-2968. The examiner can normally be reached M-F 8:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAMES BOGGS/Examiner, Art Unit 2657