DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed with the application.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/13/2025 was filed in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
1. Claims 1-10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claims 1, 9, and 10, “A voice processing system”, “a voice processing method”, and “A non-transitory computer-readable medium” are recited, which are each directed to one of the four statutory categories of invention (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
determine a degree of similarity among a respective plurality of voices acquired…: a person listens to voices in different audio signals and determines a degree of similarity (e.g. person determines they hear a 1st person in each of the signals)
output a specific first voice from among the plurality of voices…in a case where the degree of similarity among the plurality of voices is equal to or greater than a threshold value: a person selects a voice signal if they determine the signals are similar enough to each other
Claims 1, 9, and 10 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only additional elements are “A voice processing system comprising one or more processors, wherein the one or more processors are configured to” (claim 1), “A voice processing method executed by one or plurality of processors” (claim 9), “A non-transitory computer-readable medium in which a voice processing program is recorded, the voice processing program causing one or more processor to” (claim 10), “acquire voices uttered by users and input to respective microphones of a plurality of audio devices arranged in the same space” (claims 1, 9, 10), “determine a degree of similarity…from the plurality of audio devices”, and “output…to a voice processing unit…”. These limitations are recited at a high level of generality and amount to mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination with the claim as a whole, mere instructions to implement the judicial exception using a generic computer do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract ideas. Therefore, claims 1, 9, and 10 are directed to abstract ideas.
Claims 1, 9, and 10 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination with the claim as a whole, mere instructions to implement the judicial exception using a generic computer do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 1, 9, and 10 are not patent eligible.
Regarding claims 2-8, “The voice processing system” is recited, which is directed to one of the four statutory categories of invention (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
Claim 2:
wherein the one or more processors output the first voice having a highest sound pressure from among the plurality of voices…in a case where the degree of similarity among the plurality of voices is equal to or greater than the threshold value: a person decides the voices are similar enough to each other and decides to select the audio with the highest sound pressure (loudest)
Claim 2 contains the additional limitation “output…to the voice processing unit”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 3:
wherein the one or more processors output the first voice having a shortest delay time from among the plurality of voices…in a case where the degree of similarity among the plurality of voices is equal to or greater than the threshold value: a person decides the voices are similar enough to each other and decides to select the audio with the shortest delay (which audio has the voice start soonest)
Claim 3 contains the additional limitation “output…to the voice processing unit”, which amounts to mere instructions to implement the judicial exception using a generic computer.
` Claim 4:
wherein the one or more processors output the plurality of voices …in a case where the degree of similarity among the plurality of voices is less than the threshold value: a person selects the plurality of voices if the voices are not similar enough to each other
Claim 4 contains the additional limitation “output…to the voice processing unit”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 5:
wherein the one or more processors compare waveforms of the respective plurality of voices and determine a degree of similarity: a person can compare waveforms (listen to both to determine how similar they sound to each other)
Claim 6:
wherein the one or more processors execute at least one of voice conversion processing of converting the voice into text information or voice synthesis processing of synthesizing the voice: a person can listen to audio and either a) transcribe what they hear into text information using pen and paper, or b) can reproduce speech of what they hear
Claim 7:
converts the first voice into text information in a case where the degree of similarity among the plurality of voices is equal to or greater than the threshold value and converts each of the plurality of voices into text information in a case where the degree of similarity among the plurality of voices is less than the threshold value: a person transcribes and writes what they hear in a first voice or a plurality of voices based on how similar each audio signal is to each other
Claim 8:
wherein in a case where the degree of similarity among the plurality of voices is less than the threshold value, the one or more processors output a predetermined number of voices from among the plurality of voices…executes the voice conversion processing and output the plurality of voices …executes the voice synthesis processing: a person determines how similar signals are to each other, and in response to not being similar enough, select a certain number of voice signals to transcribe using pen and paper, and uses the plurality of signal to reproduce speech that they hear
Claim 8 contains the additional limitations “output…to a voice processing unit”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 2-8 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the only additional limitations amount to mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination with the claims as a whole, mere instructions to implement the judicial exception using a generic computer do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract ideas. Therefore, claims 2-8 are directed to abstract ideas.
Claims 2-8 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination with the claims as a whole, mere instructions to implement the judicial exception using a generic computer do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 2-8 are not patent eligible.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
2. Claims 1, 4-7, and 9-10 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Leblang (US 11,862,168 B1).
Regarding claim 1, Leblang discloses A voice processing system comprising one or more processors, wherein the one or more processors are configured to (Fig. 11, “CPU 1104”; Col. 28 Lines 41-67): acquire voices uttered by users and input to respective microphones of a plurality of audio devices arranged in the same space (Fig. 8; 802 and 804, receiving first and second audio signals via first and second microphones within a same environment; see Fig. 2, microphones 200(1)-(4)); determine a degree of similarity among a respective plurality of voice acquired from the plurality of audio devices (Fig. 8, 806; Col. 25 Lines 44-52 “At 806, the process 800 may compare the first audio signal and/or the second audio signal. For example, audio processing components of the transcription service 110 may compare the first audio signal and the second audio signal to identify similarities and/or differences therebetween. In some instances, comparing the first audio signal and the second audio signal may include comparing frequencies, amplitudes, pitch, and/or other audio characteristics to identify the similarities and/or differences.”); and output a specific first voice from among the plurality of voices to a voice processing unit in a case where the degree of similarity among the plurality of voices is equal to or greater (Fig. 8, “yes” branch; Col. 25 Lines 53-67 and Col. 26 Lines 1-14 “At 808, the process 800 may determine whether there is a similarity and/or a difference between the first audio signal and the second audio signal. For example, the transcription service 110, based on comparing the first audio signal and the second audio signal, may determine a portion of the first audio signal that corresponds to a portion of the second audio signal, vice versa, that represents the same speech or sound…if at 808 the process 800 determines that there are similarities and/or differences, the process 800 may follow the “YES” route and proceed to 812, whereby the process 800 may associate the similarity and/or difference with a participant. At 814, the process may filter the similarity and/or the difference from the first audio signal and/or the second audio signal …”; filtered audio of specific first voice used for further voice processing (transcription): Col. 26 Lines 32-41 “… Additionally, each of these similarities and/or differences, or the portions of the audio signals that are filtered out, may be used for generating a transcription of the meeting and/or associating microphones with participants. Furthermore, participants may be associated with virtual microphones, or the combination of audio signals across microphones, to determine a speech signal used to generate corresponding audio and/or data for the participant.”).
Regarding claim 4, Leblang discloses wherein the one or more processors output the plurality of voices to the voice processing unit in a case where the degree of similarity among the plurality of voices is less than the threshold value (Fig. 8, “no” branch; Col. 25 Lines 65-67 and Col. 26 Lines 1-3 “If at 808 the process 800 determines that there is not a similarity between the first audio signal and the second audio signal, then the process 800 may follow the “NO” route and proceed to 810 whereby the process 800 may determine a number of participants within the environment 102 based on the number of similarities and/or differences.”; Fig. 3, processed audio from first and second microphones are sent for further voice processing (transcription); Col. 20 Lines 26-45).
Regarding claim 5, Leblang discloses wherein the one or more processors compare waveforms of the respective plurality of voices and determine a degree of similarity (Col. 25 Lines 44-52 “At 806, the process 800 may compare the first audio signal and/or the second audio signal. For example, audio processing components of the transcription service 110 may compare the first audio signal and the second audio signal to identify similarities and/or differences therebetween. In some instances, comparing the first audio signal and the second audio signal may include comparing frequencies, amplitudes, pitch, and/or other audio characteristics to identify the similarities and/or differences.”).
Regarding claim 6, Leblang discloses wherein the one or more processors execute at least one of voice conversion processing of converting the voice into text information or voice synthesis processing of synthesizing the voice (Col. 26 Lines 32-41 “… Additionally, each of these similarities and/or differences, or the portions of the audio signals that are filtered out, may be used for generating a transcription of the meeting and/or associating microphones with participants. Furthermore, participants may be associated with virtual microphones, or the combination of audio signals across microphones, to determine a speech signal used to generate corresponding audio and/or data for the participant.”).
Regarding claim 7, Leblang discloses wherein the one or more processors: converts the first voice into text information in a case where the degree of similarity among the plurality of voices is equal to or greater than the threshold value (Col. 26 Lines 4-20 “…if at 808 the process 800 determines that there are similarities and/or differences, the process 800 may follow the “YES” route and proceed to 812, whereby the process 800 may associate the similarity and/or difference with a participant. At 814, the process may filter the similarity and/or the difference from the first audio signal and/or the second audio signal …”; Col. 26 Lines 32-41 “… Additionally, each of these similarities and/or differences, or the portions of the audio signals that are filtered out, may be used for generating a transcription of the meeting and/or associating microphones with participants. Furthermore, participants may be associated with virtual microphones, or the combination of audio signals across microphones, to determine a speech signal used to generate corresponding audio and/or data for the participant.”) and converts each of the plurality of voices into text information in a case where the degree of similarity among the plurality of voices is less than a threshold value (Col. 25 Lines 65-67 and Col. 26 Lines 1-3 “If at 808 the process 800 determines that there is not a similarity between the first audio signal and the second audio signal, then the process 800 may follow the “NO” route and proceed to 810 whereby the process 800 may determine a number of participants within the environment 102 based on the number of similarities and/or differences.”; Fig. 3, processed audio from first and second microphones are sent for further voice processing (transcription); Col. 20 Lines 26-45).
Regarding claim 9, claim 9 is a method claim with limitations similar to those recited in system claim 1, and thus is rejected under similar rationale.
Regarding claim 10, claim 10 is a non-transitory CRM claim with limitations similar to those recited in system claim 1, and thus is rejected under similar rationale.
Additionally, Leblang discloses A non-transitory computer-readable medium in which a voice processing program is recorded, the voice processing program causing one or more processors to (Col. 30, Lines 6-17).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
3. Claims 2 and 3 are rejected under 35 U.S.C. 103 as being unpatentable over Leblang in view of Leblang et al. (US 2020/0153646 A1, hereinafter Leblang 2).
Regarding claim 2, Leblang discloses the step of determining a case where the degree of similarity from among the plurality of voices is equal to or greater than the threshold value (see above claim mapping in claim 1), but does not specifically disclose that the one or more processors output the first voice having a highest sound pressure from among the plurality of voices to the voice processing unit [in a case where the degree of similarity among the plurality of voices is equal to or greater than the threshold value.]
Leblang 2 teaches the one or more processors output the first voice having a highest sound pressure from among the plurality of voices to the voice processing unit (plurality of voice signals received via several devices: para. 0075 “At block 322, multiple audio signals can be received from different devices. One or more voice-enabled devices, microphones, or conference devices may transmit multiple audio signals to the voice based system 200…” para. 0075 “Once a voice command is received from a group, the voice based system 200 can listen for commands from the same group and/or can determine if other devices in the same group received the same command.”; voice signal captured by audio device with the highest recorded energy level is selected: para. 0080 “At block 335, a particular device can be determined to be associated with the command. For example, the arbitration service 270 can use the data from the previous blocks to determine that a particular device is associated with the command. A particular device can be selected from among the group of devices that received the same command or instead of another device that received the same command. The arbitration service 270 can identify a particular device from a group of devices that received the same command with the highest energy level; within a particular energy band…”’; this selected voice signal is sent for further voice processing: para. 0083 “At block 340, the command is executed. For example, the execution service 252 may execute the command that was associated with a particular device.”).
Leblang and Leblang 2 are considered to be analogous to the claimed invention as they both are in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Leblang to incorporate the teachings of Leblang 2 in order to select the first voice having a highest sound pressure from among the plurality of voices to the voice processing unit. Doing so would be beneficial, as this would identify a best microphone to select reflecting speech of a current speaker, as higher energy levels indicate the microphone is closer to the user speaking (para. 0099).
Regarding claim 3, Leblang discloses a case where the degree of similarity from among the plurality of voices is equal to or greater than the threshold value (see above claim mapping in claim 1), but does not specifically disclose that the one or more processors output the first voice having a shortest delay time from among the plurality of voices to the voice processing unit [in a case where the degree of similarity among the plurality of voices is equal to or greater than the threshold value.]
Leblang 2 teaches the one or more processors output the first voice having a shortest delay time from among the plurality of voices to the voice processing unit (plurality of voice signals received via several devices: para. 0075 “At block 322, multiple audio signals can be received from different devices. One or more voice-enabled devices, microphones, or conference devices may transmit multiple audio signals to the voice based system 200…” para. 0075 “Once a voice command is received from a group, the voice based system 200 can listen for commands from the same group and/or can determine if other devices in the same group received the same command.”; para. 0119 “At block 720, the particular device associated with the command can be determined based on the time data. The arbitration service 270 can select a first voice command instead of a second voice command based at least in part on a first timestamp of the first voice command being earlier than a second timestamp of the second voice command. For example, a first voice command can be associated with a first timestamp (such as the value 1 millisecond) that indicates a time when the corresponding audio signal was received by voice-enabled device and/or the voice-based system. A second voice command can be associated with a second timestamp (such as the value 1000 milliseconds) that indicates a time when the corresponding audio signal was received by voice-enabled device and/or the voice-based system. Accordingly, the arbitration service 270 can select the first voice command because the first timestamp (with the value 1 millisecond) is earlier or less than the second timestamp (with the value 1000 milliseconds) of the second voice command.”; this selected voice signal is sent for further voice processing: para. 0120 “At block 725, the command is executed. For example, the execution service 252 may execute the command that was determined to be associated with the particular voice-enabled device.”).
Leblang and Leblang 2 are considered to be analogous to the claimed invention as they both are in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Leblang to incorporate the teachings of Leblang 2 in order to select the first voice having a shortest delay time from among the plurality of voices to the voice processing unit. Doing so would be beneficial, as this would identify a best microphone to select reflecting speech of a current speaker, as shorter delay time would indicate the microphone is closer to the user speaking.
4. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Leblang in view of Yoshioka et al. (US 2021/0407516 A1, hereinafter Yoshioka) and further in view of Cutler et al. (US 2024/0406621 A1, hereinafter Cutler).
Regarding claim 8, Leblang discloses determination of a case where the degree of similarity among the plurality of voices is less than the threshold value (Fig. 8, “no” branch; Col. 25 Lines 65-67 and Col. 26 Lines 1-3 “If at 808 the process 800 determines that there is not a similarity between the first audio signal and the second audio signal, then the process 800 may follow the “NO” route and proceed to 810 whereby the process 800 may determine a number of participants within the environment 102 based on the number of similarities and/or differences.”), the one or more processors output a …number of voices from among the plurality of voices to a voice processing unit that executes the voice conversion processing… (Fig. 3, processed audio from first and second microphones are sent for further voice processing (transcription); Col. 20 Lines 26-45).
Leblang does not specifically disclose to [output a ] predetermined number of voices [from among the plurality of voices to a voice processing unit that executes the voice conversion processing…].
Yoshioka teaches to [output a ] predetermined number of voices [from among the plurality of voices to a voice processing unit that executes the voice conversion processing…] (predetermined number of audio signals (two) sent for transcription processing: Col. 21 Lines 33-42 “Method 1700 begins by receiving multiple channels of audio at operation 1710 from three or more microphones detecting speech from a meeting of multiple users. At operation 1720, directions of active speakers are estimated. A speech unmixing model is used to select two channels that may correspond to a primary and a secondary microphone at operation 1730, or may correspond to a fused audio channel. The two selected channels are sent at operation 1740 to a meeting server for generation of an intelligent meeting transcript.”).
Leblang and Yoshioka are considered to be analogous to the claimed invention as they both are in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Leblang to incorporate the teachings of Yoshioka in order to output a predetermined number of voices from among the plurality of voices to the voice processing unit. Doing so would be beneficial, as reducing the amount of data being transmitted would conserve bandwidth without significantly reducing transcription accuracy (Col. 21 Lines 42-45).
Leblang in view of Yoshioka does not specifically disclose and output the plurality of voices to a voice processing unit that executes the voice synthesis processing.
Cutler teaches to output the plurality of voices to a voice processing unit that executes the voice synthesis processing (a plurality of voices (signals 322 and 332) are output to server 150, and have subsequent voice synthesis processing performed (synchronization) to then be output to a remote user in a second room: para. 0046 “As shown in signal mixing and synchronization 360, the server synchronizes signals 322 and 332, received from client devices 120 and 130, respectively, and mixes them to produce playback signal 352. Note, however, that signals 312 and 342 originating from client devices 110 and 140, respectively, are excluded. The server communicates playback signals 352(1), 352(3), and 352(4) to client devices 110, 130, and 140, respectively.”).
Leblang, Yoshioka, and Cutler are considered to be analogous to the claimed invention as they are all in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Leblang in view of Yoshioka to incorporate the teachings of Cutler in order to output the plurality of voices to a voice processing unit that executes the voice synthesis processing. Doing so would be beneficial, as such mixing would mitigate echoes/undesirable artifacts (para. 0048), improving speech quality.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Chun et al. (US 2023/0282224 A1): selecting between different microphones, selecting audio source with loudest volume or greatest clarity (para. 0106), performing of transcription of concurrent conversations (para. 0091)
Shen & Han (US 2019/0341068 A1): selection of audio stream based on having a higher signal-to-noise ratio than that of a predetermined number of other audio streams (Abstract)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CODY DOUGLAS HUTCHESON/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659