Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/17/2026 has been entered.
Response to Amendment
3. In response to the office action mailed on 04/07/2026, applicant filed an amendment on 06/17/2026, amending claims 1, 2, 6, 7, 8, 12, 13, 14, and 20. The pending claims are 1-20.
Response to Arguments
4. Applicant’s arguments with respect to the pending claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Applicant argues that the prior art Novitasari does not teach a processing queue that is separate from a transcription model. Applicant asserts there is no disclosure in Novitasari of a queue data structure, an enqueue operation, or any structural separation between a collection mechanism and a. transcription model. The examiner notes that the limitation “if speech is detected in a particular short segment, then adding a particular short segment to a processing queue that is separate from a transcription model” is broad and not tied to other claimed limitations. Moreover, applicant is referred to Figure 3D and paragraph [0077] of Novitasari wherein the output of VAD model is fed into a fully connected layer exterior to the RNN-T 302 (separate from the transcription model).
Applicant argues that the prior art Novitasari fails to teach claim 1 limitation “to determine a presence or an absence of speech within each short segment” because Novitasari’s VAD model does not operate as a binary gate. Instead, it produces a continuous probability output. The examiner notes that Novitasari teaches an ASR systems that is deployed together with a voice activity detection (VAD) system to run ASR on the voiced acoustic signals. The VAD system separates speech from non-speech segments and removes unnecessary non-speech parts from input audio signals during inference ([0026]). More, paragraph [0027], of Novitasari explicitly disclose “ASR can be paired with a VAD system that extracts actual speech parts from an input audio signal by removing nonspeech parts before a decoding process of ASR starts”. This means Novitasari determines a presence or an absence of speech within each short segment Accordingly, the prior art Novitasari reads on the claim language. Moreover, operating a binary gate is not recited in the rejected claim. Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Furthermore, Kaskari in the same field of endeavor teaches a speech enhancement system that includes a voice activity detector (VAD) associated with a linear filter that determines or predicts whether speech is present (or absent) in a current audio frame (binary value). The speech enhancement system applies multiple VADs in a cascaded fashion with different parameters performing linear filtering operations based on different corresponding noise covariances of the queued frames to produce a more accurate result ([0047], [0007], [0061] -0063], [0089]).
Applicant argues that the prior art Novitasari fails to teach the limitation "wherein each short segment is between 250 milliseconds and 500 milliseconds in duration”. The examiner notes that Hiray (US 20240221721) in the same field of endeavor teaches processing speech segments that are 300 ms (0.3 second), see paragraph [0023]. Therefore, it would have been obvious at the time the application was filed to use the above length feature of Hiray with the system of Novitasari, in order to improve accuracy, make it easier to correct errors, and enhance efficiency.
As per the rest of the claims, and combinations of prior art reference, applicant has no further arguments beside the ones mentioned above. Therefore, all the combinations of prior art reference mentioned above are valid, and all other claims are rejected for the same reasons as set above.
Claim Rejections - 35 USC § 103
5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Novitasari (US 2024/0038221) in view of Hiray (US 20240221721), and further in view of Kaskari (US 20240257827).
As per claims 1 and 13, Novitasari teaches continuously capturing audio data and segment it into short segments ([0026]- [0029], [0071], [0101], Figs. 2A and 2B, capturing audio signals and segmenting the segments into a sequence of frames);
implementing a pre-trained enterprise-grade voice activity detection (“VAD”) system on each of the short segments to determine a presence or an absence of speech within each short segment ([0026], wherein said, the automatic speech recognition (ASR) systems can be deployed together with a voice activity detection (VAD) system to run ASR on the voiced acoustic signals. More, paragraph [0027], of Novitasari explicitly disclose “ASR can be paired with a VAD system that extracts actual speech parts from an input audio signal by removing nonspeech parts before a decoding process of ASR starts”. This means Novitasari determines a presence or an absence of speech within each short segment), wherein, if speech is detected in a particular segment, then adding the particular short segment to a processing queue that is separate from a transcript model, and wherein, if speech is not detected in the particular short segment, immediately discarding the particular short segment and, declining to add the particular segment to the processing queue, reducing unnecessary processing (in addition to [0026]- [0028], Novitasari teaches an ASR systems that is deployed together with a voice activity detection (VAD) system to run ASR on the voiced acoustic signals. The VAD system separates speech from non-speech segments and removes unnecessary non-speech parts from input audio signals during inference ([0026]). More, paragraph [0027], of Novitasari explicitly disclose “ASR can be paired with a VAD system that extracts actual speech parts from an input audio signal by removing nonspeech parts before a decoding process of ASR starts”. This means Novitasari determines a presence or an absence of speech within each short segment. For a processing queue that is separate from a transcript model, see Figure 3D and paragraph [0077] of Novitasari wherein the output of VAD model is fed into a fully connected layer exterior to the RNN-T 302 (separate from the transcription model); and
filtering out non-speech segments to reduce computational waste, focusing resources on relevant audio data and minimizing latency ([0026]- [0028], maintaining performance by extracting/processing necessary speech parts and removing unnecessary non-speech parts and long silence regions from input audio signals during inference and improving robustness of speech recognition in noisy conditions [0030]- [0031] and Fig. 7).
Novitasari may not explicitly disclose wherein each short segment is optimized to fall between 250 millisecond and 500 millisecond in duration. Hiray in the same field of endeavor teaches processing speech segments that are 300 millisecond in duration (0.3 second), see paragraph [0023]. Therefore, it would have been obvious at the time the application was filed to use the above segment duration feature of Hiray with the system of Novitasari, in order to improve accuracy, make it easier to correct errors, and enhance efficiency.
Novitasari in view of Hiray may not explicitly disclose applying the VAD again to queued audio in the processing queue to eliminate any residual noise and silence, refining the audio data further. Kaskari in the same field of endeavor teaches a speech enhancement system that includes a voice activity detector (VAD) associated with a linear filter that determines or predicts whether speech is present (or absent) in a current audio frame (binary value). The speech enhancement system applies multiple VADs in a cascaded fashion with different parameters performing linear filtering operations based on different corresponding noise covariances of the queued frames to produce a more accurate result ([0047], [0007], [0061] -0063], [0089]). Therefore, it would have been obvious at the time the application was filed to use the above feature of Kaskari with the system of Novitasari in view of Hiray in order to remove unwanted noise and improve sound quality.
As to stitching together cleaned audio segments from the processing queue to form a
coherent audio stream without gaps, wherein the coherent audio stream is more representative of natural speech, improving accuracy and effectiveness of subsequent machine learning processes, Kaskari teaches a speech enhancement system that performs filtering or suppressing noise in a received audio signal by segmenting received audio signal, applying voice activity detection to determine where speech is present or absent, determining interframe correlation of the speech component between consecutive frames in the series of frames, filtering the noise component of the frames of the audio signal, and producing a clean/enhanced audio signal ([0087], [0076]). Reproducing a clean, continuous audio signal after filtering necessarily discloses stitching (or overlap-adding) the audio frames, otherwise, the resulting audio will suffer from severe distortion, clicks, and a "buzzy" artifact.
As per claim 2, Novitasari in view of Hiray may not explicitly disclose applying VAD again to queued audio comprises applying the VAD with stricter parameters or enhanced sensitivity settings to identify and remove remnants of noise or insignificant pauses within speech segments. Kaskari in the same field of endeavor teaches a speech enhancement system that includes a voice activity detector (VAD) associated with a linear filter that determines or predicts whether speech is present (or absent) in a current audio frame (binary value). The speech enhancement system applies multiple VADs in a cascaded fashion with different and stricter parameters performing linear filtering operations based on different corresponding noise covariances of the queued frames to produce a more accurate result ([0047], [0007], [0061] -0063], [0089]). Therefore, it would have been obvious at the time the application was filed to use the above feature of Kaskari with the system of Novitasari in view of Hiray in order to remove unwanted noise and improve sound quality.
As per claim 3, Novitasari teaches organizing the coherent audio stream into segments and pad them to uniform lengths to fit the expected input format for the transcription model; and enhancing an efficiency of deep learning models by reducing variability in input data [0071]- [0073], wherein said, the VAD model 204 can predict the sequence of voice activity class v=(v.sub.1, . . . , v.sub.T) of length T from a speech frame sequence x with the same length. The VAD integration system 100 can concatenate features between the VAD output probability p(v|x) and the ASR feature of the corresponding speech frame.
As per claim 4, Novitasari teaches transforming the input data into a transcribed text ([0095] text decoded by standard RNN-T of the speech recognition system). Novitasari may not explicitly disclose automatically detecting the language of the transcribed text, facilitating targeted translation processes; and translating the transcribed text into the desired language as a translated text using a robust language model from open-source libraries, supporting multiple language pairs, wherein the multiple language pair is an identifier that describes a combination of multiple languages as used in the translation process; and converting the translated text back into speech to provide auditory feedback, enhancing accessibility for users who may not be able to read text conveniently. Hiray in the same field of endeavor teaches a multi-language translation system (“MHLTS”) that detects and translates between different spoken languages in real-time, and automatically adapt to users speaking different languages during the presentation, conversation, or conference, and that convert, translate, and/or transcribe the audio from each spoken language to text or audio of a desired target language ([0012]). Converting the translated text back into speech is well known in the and suggested by paragraph [0012], wherein said, transcribing the audio from each spoken language to text or audio of a desired target language, and necessarily disclosed by the process of paragraph [0013], wherein said the MLTS allows the conference participants to speak in different native languages (e.g., different languages they are most comfortable with that other participants may not understand), and translates the spoken dialog to one or more target languages selected by each conference participant. Therefore, it would have been obvious at the time the application was filed to use the above features of Hiray with the system of Novitasari, in order to provide further technological benefit of a multilingual conferencing solution for large numbers of participants speaking different languages ([0016]).
As per claim 5, Novitasari may not explicitly disclose processing audio data without waiting for long recordings to end to enable live translation and responsive voice-activation. Hiray in the same field of endeavor teaches a multi-language translation system for detecting and translating between different spoken languages in real-time as one or more users switch between the different languages during an online or computer-hosted presentation, conversation, or conference ([0012]), providing a real-time transcription of the dialog as it is spoken ([0014]), and [0016], wherein said the system performs a real-time translation and/or transcription of the dialog within each conference. Therefore, it would have been obvious at the time the application was filed to use the above features of Hiray with the system of Novitasari, in order to provide further technological benefit of a multilingual conferencing solution for large numbers of participants speaking different languages ([0016]).
As per claim 6, Novitasari may not explicitly disclose wherein each short segment is adaptively adjusted in duration within the range of 250 milliseconds and 500 milliseconds based on a complexity of an audio environment, and wherein in a noisy setting, shorter segments are used to isolate speech from background noise, and in a clear audio environment, longer segments are used to reduce processing overhead. Hiray in the same field of endeavor teaches parsing the audio signal into snippets of 0-3 sec, 1-4 sec, 300 millisecond in duration ([0023]). Therefore, it would have been obvious at the time the application was filed to use the above segmenting feature of Hiray with the system of Novitasari to adaptively adjust each segment as claimed, in order to improve accuracy, make it easier to correct errors, and enhance efficiency. As to adjusting segment length based on environmental noise, it’s noted that this technique is standard and highly effective practice in fields like data communications, audio processing, and digital signal processing. Therefore, it would have been obvious at the time the application was filed for the combined system of Novitasari in view of Hiray and Kaskari to adaptively adjust the segments duration as claimed. This would maximize throughput and efficiency by balancing overhead against error rates.
As per claims 7-12, system claims 7-12 and method claims 1-6 are related as apparatus and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claims 7-12 are similarly rejected under the same rationale as applied above with respect to method claims 1-6. Furthermore, Novitasari teaches one or more processors; and memory storing thereon instructions, as claimed ([0111]- [0114]).
As per claims 14-20, the claims are similarly rejected under the same rationale as applied above with respect to claims 2-6.
Conclusion
6. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO-892.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDELALI SERROU whose telephone number is (571)272-7638. The examiner can normally be reached M-F 9 Am - 5 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABDELALI SERROU/ Primary Examiner, Art Unit 2659