Prosecution Insights
Last updated: October 02, 2026
Application No. 18/755,051

DETECTING BREAKS IN SPEECH FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS

Final Rejection §101§102§103
Filed
Jun 26, 2024
Examiner
ADESANYA, OLUJIMI A
Art Unit
2658
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
66%
Grant Probability
Favorable
3-4
OA Rounds
1y 2m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 66% — above average
66%
Career Allowance Rate
446 granted / 678 resolved
+3.8% vs TC avg
Strong +27% interview lift
Without
With
+26.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
19 currently pending
Career history
708
Total Applications
across all art units

Statute-Specific Performance

§101
19.6%
-20.4% vs TC avg
§103
44.1%
+4.1% vs TC avg
§102
17.1%
-22.9% vs TC avg
§112
13.0%
-27.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 678 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 7/17/26 have been fully considered but they are not persuasive. Regarding the 35 U.S.C. 101 rejection of the claims, Applicant argues that the instant claims are directed to patent eligible subject matter as the claims recite a specialized machine-learning architecture that determines end of sentence (EOS) and end of utterance (EOU) predictions from speech and that an human cannot perform the steps as claimed as the claims rely on trained machine learning models (Arguments, pg. 15, second para. – pg. 16, second para.). Examiner respectfully disagrees as there is no recitation of utilizing a trained machine learning model, actively training a machine learning model, nor actively training a machine learning model to perform unconventional functions in the claims. Furthermore, a human (including a transcriptionist and a stenographer) can analyze received audio/speech, transcribe such speech on paper, and indicate or assign probability values to portions of the speech/transcription that correspond to end of sentences/sentence boundaries as well as the ending of such speech/utterance, while further providing a response to such speech. Applicant also argues that the use of the unclaimed trained machine learning model as described in the specification provides for improved responsiveness and latency reduction, thereby providing technological improvements to an abstract idea and that the claims are further analogous to court decision Enfish, LLC, McRo, Inc and Thales Visionix Inc, in employing a specific arrangement of machine learning models to predict EOU and EOS (Arguments, pg. 16, third para. – pg. 18, first para.). Examiner respectfully disagrees as there is no recitation of utilizing a trained machine learning model, actively training a machine learning model, nor actively training a machine learning model to perform unconventional functions in the claims, where Applicant’s original specification explicitly describes the claimed generic models as components of memory executed by a processor (pg. 27, para. [0085]). Also, the improved responsiveness and latency reduction referenced by Applicant correspond to improvements (i.e., tying to acomputer) to the abstract idea of analyzing manually/mentally transcribed speech, but do not reflect improvements to the functioning of the computer implementing the steps, nor to other technology. Furthermore, unlike Enfish, LLC, McRo, Inc and Thales Visionix Inc, the claims simply utilize a generic computer or generic computer components to perform steps that would otherwise be performed manually/mentally by a human transcriber/stenographer. Applicant further argues that none of the cited prior art teaches or suggests the architecture reflected by the claimed steps and that the claims are analogous to the Office’s Examples 40 (Network Accounting and Billing) and 46 (Stable Training of a Neural Network), and as such, argues that the claims recite significantly more than the abstract idea (arguments, pg. 18, second para. – pg. 20, first para.). Examiner respectfully disagrees as the novelty or non-obviousness (35 U.S.C. 102/103) of the claims over prior art does not automatically confer patent eligibility in terms of statutory subject matter (35 U.S.C. 101) - "groundbreaking, innovative, or even brilliant discovery does not by itself satisfy the § 101 inquiry." see Ass'n for Molecular Pathology V. Myriad Genetics, Inc., 569 U.S. 576, 591 (2013)., "A novel and non obvious claim directed to a purely abstract idea is, nonetheless, patent ineligible" see Mayo, 566 U.S. at 90, "The 'novelty' of any element or steps in a process, or even of the process itself, is of no relevance in determining whether the subject matter of a claim falls within the § 101 categories of possibly patentable subject matter" see also Diamond V. Diehr, 450 U.S. 175, 188-89 (1981). Nevertheless, the limitations reflected in the claims are taught by the references as presented in the prior art rejection below. Also, it is unclear what examples Applicant is referring to as Examples 40 and 46 of the Office’s Subject Matter Eligibility Examples respectively refer to “Adaptive Monitoring of Network Traffic Data” and “Livestock Management”. The instant claims neither improve the operation of a computer network through a particular distributed architecture nor actively operate/train machine-learning models to constitute practical technological improvements like Applicant argues. Instead, the claims merely tie the abstract idea of analyzing transcribed text data to generic computer components. Regarding the amended claims in light of the previous rejections of the claims with references Rangarajan and Pathak, Applicant argues that Pathak discloses calculating a score that represents the probability of the end of speech at a word boundary but does not teach calculating a probability associated with an end of utterance occurring at the end of the last sentence as required by claim 1 (“generating by a first model and based at least on the text data, output data representing first probabilities indicating whether each token corresponding to the text data is associated with an end of the first sentence within the utterance and second probabilities indicating whether each token is associated with an end of the utterance that occurs at the end of the last sentence”) and claim 19 ("generate, using one or more models and based at least on text data corresponding to an utterance that includes one or more words, output data representing: one or more first probabilities indicating whether the one or more words are associated with an end of a sentence within the utterance; one or more second probabilities indicating whether the one or more words are associated with an end of the utterance; and one or more third probabilities indicating whether the one or more words are not associated with the end of the sentence or the end of the utterance,") and as a result similar claim 7 that does not recite the claimed probabilities presented in claims 1 and 19, and claims dependent therefrom (Arguments, pg. 20, third para.- pg. 26, fifth para.). Examiner respectfully disagrees. Regarding claim 7, Rangarajan discloses its systems receiving stream of speech 212 ("La idea es trabajar junto a la provincial y la nacion en la lucha contra el narcotrafico."), performing automatic speech recognition on the stream of speech and producing via a speech recognizer model, partial speech hypotheses 216, where in addition to recognizing the speech, the model also searches for text indicating a portion of the transcribed text should be segmented and where the model performs such sentence segmentation 226 on the text hypotheses 216 identified up to time 222, resulting in a segment 228/initial sentence "La idea es trabajar junto a la provincial y” (fig. 2, para. [0038]-[0039]) i.e., identifying an initial “end of sentence” in the stream of speech. Rangarajan also discloses its system continuing to output hypotheses including additional sentences as part of the remainder of the speech is "la nacion en la lucha contra el narcotrafico," while identifying that the end of the second sentence “la nacion en la lucha contra el narcotrafico." is also the end of the speech (fig. 2, para. [0041]), i.e., identifying an end of utterance. Therefore, Rangarajan discloses its model identifying an initial end of sentence (at the end of segment 1), as well as a subsequent end of sentence (at the end of segment 2) that also corresponds to an end of utterance, corresponding to limitations “to: generate, using one or more models and based at least on text data associated with an utterance that includes words, first outputs indicating whether the words are associated an end of a sentence and second outputs indicating whether the one or more words are associated with an end of the utterance that occurs at an end of a last word of the words” and “determine, based at least on the first outputs and the second outputs, a first location within the utterance that is associated with the end of the sentence and a second location within the utterance that is associated with the end of the utterance” as required by the argued limitations of claim 7. Likewise, these teachings by Rangarajan corresponds to limitations “generating by a first model and based at least on the text data, output data indicating whether each token corresponding to the text data is associated with an end of the first sentence within the utterance and indicating whether each token is associated with an end of the utterance that occurs at the end of the last sentence” and “determining, based at least on the output data, a first location within the text data that is associated with the end of the first sentence and a second location within the text data that is associated with the end of the utterance” as required by claim 1 and limitations “generate, using one or more models and based at least on text data corresponding to an utterance that includes one or more words, output data” and “determine, based at least on the output data, a first location within the text data that is associated with the end of the sentence and a second location within the text data that is associated with the end of the utterance” as required by claim 19. What Rangarajan does not explicitly disclose includes the use of the first probabilities and second probabilities recited in claims 1 and 19. Regarding claims 1 and 19, as described above, Pathak discloses its system decoding received audio to recognize words present in the audio (fig. 6; para. [0001]) as well as utilizing a language segmentation model in identifying potential segmentation boundaries in the decoded audio that are predicted to correspond to an end of a speech utterance, including calculating segmentation scores corresponding to potential segmentation boundaries in the decoded received audio (para. [0034]; para. [0037]), where when the model identifies the potential segmentation boundary, the model “then” calculates a language segmentation score that represents the probability of an end of speech (EOS) (para. [0060]), corresponding to limitations “generating by a first model and based at least on the text data, output data representing first probabilities” and second probabilities” as required by claim 1, as well as “output data representing: one or more first probabilities and one or more second probabilities” as required by claim 19. Pathak further discloses in addition to determining the scores representing the potential sentence boundaries (i.e., end of sentence) and the end of speech (end of speech), the model determines if the scores are low (para. [0060]-[0061]), where the system continues to process the audio if the scores are low, corresponding to limitation “one or more third probabilities indicating whether the one or more words are not associated with the end of the sentence or the end of the utterance” as additional required by claim 19. Therefore, Examiner maintains that Rangarajan discloses the limitations of claim 7 and the combination of Rangarajan and Pathak discloses the limitations of claims 1 and 19. Regarding dependent claims 3, 4 and 12-15 rejected with additional references Bijwadia, Hassid and Lineback, Applicant argues that Bijwadia, Hassid and Lineback fail to disclose limitations recited in claims 1 and 7 from which the y depend and as such do not teach limitations recited in the dependent claims (Arguments, pg. 27-28). Examiner respectfully disagrees as references Bijwadia, Hassid and Lineback are/were not applied to teach language recited in the independent claims, and absent any argument as to why the references fail to disclose language recited in the dependent claims, Examiner maintains that the rejection of the dependent claims are appropriate. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (text analysis) without significantly more. Claim 1 recites steps of generating, based at least on audio data representative an utterance, text data corresponding to the utterance that includes at least and a first sentence and a last sentence (i.e., a data analysis/evaluation step), generating, using one or more by a first model and based at least on the text data, output data representing first probabilities indicating whether each token corresponding to the text data is associated with an end of the first sentence within the utterance and second probabilities indicating whether each token is associated with an end of the utterance that occurs at an end of the last sentence (i.e., a data analysis/evaluation step), determining, based at least on the output data, a first location within the text data that is associated with the end of the first sentence and a second location within the text data that is associated with the end of the utterance (i.e., a data analysis/evaluation step) processing, using one or more second models and based at least on the first location and the second location, a first portion of the text data corresponding to the first sentence of the utterance prior to followed by processing a second portion of the text data corresponding to a remainder of at least the last sentence of the utterance (i.e., a data analysis/evaluation step) and generating, based at least on the processing the first portion of the text data followed by the second portion of the text data, a response associated with the utterance (i.e., a data analysis/evaluation step). Claim 7 recites generate, using one or more models and based at least on text data associated with an utterance that includes words, first output indicating whether the words are associated an end of a sentence and second outputs indicating whether the one or more words are associated with an end of the utterance that occurs at an end of a last word of the words (i.e., a data analysis/evaluation step), determine, based at least on the first outputs and the second outputs, a first location within the utterance that is associated with the end of the sentence and a second location within the utterance that is associated with the end of the utterance (i.e., a data analysis/evaluation step), and cause, based at least on the first location and the second location, processing of a first portion of the text data associated with the end of the sentence followed by processing of a second portion of the text data associated with the end of the utterance (i.e., a data analysis/evaluation step). Claim 19 recites generate, using one or more models and based at least on text data corresponding to an utterance that includes one or more words, output data representing: one or more first probabilities indicating whether the one or more words are associated with an end of a sentence within the utterance (i.e., a data analysis/evaluation step), one or more second probabilities indicating whether the one or more words are associated with an end of the utterance (i.e., a data analysis/evaluation step), and one or more third probabilities indicating whether the one or more words are not associated with the end of the sentence or the end of the utterance (i.e., a data analysis/evaluation step), determine, based at least on the output data, a first location within the text data that is associated with the end of the sentence and a second location within the text data that is associated with the end of the utterance (i.e., a data analysis/evaluation step) and process, based at least on the first location and the second location, a first portion of the text data corresponding to the end of the sentence of the utterance followed by a second portion of the text data corresponding to a remainder of the utterance (i.e., a data analysis/evaluation step). These claims include steps achievable by a human mentally analyzing audio/speech or manually using a pen and paper in transcribing the speech to obtain text/words, assigning probabilities and locations indicating first and second portions indicating end of sentence and end of audio/utterance based on an analysis of uppercase/lowercase letters of the text/words corresponding to the audio and generating a response (e.g. a summary) based on processing the first and second portion, corresponding to the mental processes category of abstract ideas. This judicial exception is not integrated into a practical application because the claims are directed to an abstract idea with additional generic computer elements, where the generically recited computer elements (system, processor, models, processing circuitry) do not add a meaningful limitation to the abstract idea because they amount to simply implementing the abstract idea on a computer. The models are described in Applicant’s original specification as components of memory executed by a processor (pg. 27, para. [0085]). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because steps “generating, based at least on the processing the first portion of the text data followed by the second portion of the text data, a response associated with the utterance”, “processing of a first portion of the text data associated with the end of the sentence followed by processing of a second portion of the text data associated with the end of the utterance” and “process, based at least on the first location and the second location, a first portion of the text data corresponding to the end of the sentence of the utterance followed by a second portion of the text data corresponding to a remainder of the utterance” correspond to well-understood, routine, conventional computer functions of collecting information and analyzing it, as well as utilizing a generic computer to process data as recognized by the court decisions listed in MPEP § 2106.05, and as provided by cited reference Rangarajan and Pathak (PTO 892 form). The dependent claims also recite mental processes and do not add significantly more than the abstract idea and are as such similarly rejected. Claim Objections Claim 1 is objected to because of the following informalities: “generating, based on audio …that includes at least and a first sentence …” as recited in the claims should be “generating, based on audio …that includes at least a first sentence …”. Appropriate correction is required. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 1. Claims 7-9, 11, 17 and 18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Rangarajan Sridhar et al US 2015/0134320 A1 (“Rangarajan”) Per claim 7, Rangarajan discloses a system comprising: one or more processors to: generate, using one or more models and based at least on text data associated with an utterance that includes words, first outputs indicating whether the words are associated an end of a sentence and second outputs indicating whether the one or more words are associated with an end of the utterance that occurs at an end of a last word of the words (fig. 2; para. [0023]; para. [0029]; The system then performs sentence segmentation 226 on the text hypotheses 216 identified up to time 222, resulting in a segment 228. The segment consisting of the beginning of the sentence to the conjunction, is referred to as segment 1 228…., para. [0039]; The system's segmenter 226, instead of finding a conjunction, identifies the end of the speech as a likely end of a sentence, and segments the text hypothesis "la nacion en la lucha contra el narcotrafico" as a second segment 230 …, para. [0041]; para. [0047]); determine, based at least on the first outputs and the second outputs, a first location within the utterance that is associated with the end of the sentence and a second location within the utterance that is associated with the end of the utterance (fig. 2; The system then performs sentence segmentation 226 on the text hypotheses 216 identified up to time 222, resulting in a segment 228…., para. [0039]; The system's segmenter 226, instead of finding a conjunction, identifies the end of the speech as a likely end of a sentence, and segments the text hypothesis "la nacion en la lucha contra el narcotrafico" as a second segment 230…., para. [0041], end of first segment 228 (segment 1, fig. 2) as end of sentence, end of second segment 230 (segment 2, fig. 2) as endof utterance and end of second sentence); and cause, based at least on the first location and the second location, processing of a first portion of the text data associated with the end of the sentence followed by processing of a second portion of the text data associated with the end of the utterance (fig. 2, elements 232, 246; Immediately upon segmenting the segment 1 (at time 234), the system begins machine translation 232 on segment 1. As illustrated, this process 232 begins at time 234 and ends at time 242 with a machine translation of segment 1 236, which is translated text corresponding to segment 1. Continuing with the example, the machine translation produces "The idea is to work together with the province and" as text corresponding to segment 1. Upon producing the machine translation of segment 1 at time 242, the system immediately begins to output 246 an audio version of the machine translation, para. [0040]). Per claim 8, Rangarajan discloses the system of claim 7, wherein the first location is associated with a first word of the words that is associated with the end of the sentence (fig. 2; para. [0038]-[0041]); and the second location is associated with the last word of the words that is associated with the end of the utterance (fig. 2; para. [0038]-[0041]) Per claim 9, Rangarajan discloses the system of claim 8, wherein the one or more processors are further to: determine the first portion of the text data based at least on the first word being associated with the end of the sentence (para. [0038]-[0039]); and determine the second portion of the text data based at least on the last word being associated with the end of the utterance (para. [0038]-[0041]). Per claim 11, Rangarajan discloses system of claim 7, wherein the first outputs represent first indicators that indicate whether the words are associated with the end of sentence (fig. 2; para. [0038]-[0039]; para. [0041]); and the second outputs represent second indicators that indicate whether the words are associated with the end of the utterance (fig. 2; para. [0039]-[0041]). Per claim 17, Rangarajan discloses the system of claim 7, wherein: the text data represents tokens associated with the words (para. [0038]-[0039]); and the first outputs indicates whether the tokens are associated with the end of the sentence and the second outputs indicate whether the tokens are associated with the end of the utterance (para. [0039]; para. [0041]). Per claim 18, Rangarajan discloses the system of claim 7, wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative Al operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more visual language models (VLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data (para. [0042]; para. [0044]; para. [0047]); a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 2. Claims 1, 2, 6, 10, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Rangarajan in view of Pathak et al - US 2025/0054491 A1 (“Pathak”) Per claim 1, Rangarajan discloses a method comprising: generating, based at least on audio data representative an utterance, text data corresponding to the utterance that includes at least and a first sentence and a last sentence (fig. 2; para. [0017]; As the system receives the stream of speech 212, the system is performing automatic speech recognition 214 on the stream of speech 212. The automatic speech recognition process 214 produces, via a speech recognizer, partial speech hypotheses 216…, para. [0038]; These remainder hypotheses can include words, sentences, paragraphs, or …, para. [0041]); generating by a first model and based at least on the text data, output data indicating whether each token corresponding to the text data is associated with an end of the first sentence within the utterance and indicating whether each token is associated with an end of the utterance that occurs at the end of the last sentence (fig. 2; para. [0023]; para. [0029]; The system then performs sentence segmentation 226 on the text hypotheses 216 identified up to time 222, resulting in a segment 228. The segment consisting of the beginning of the sentence to the conjunction, is referred to as segment 1 228…., para. [0039]; The system's segmenter 226, instead of finding a conjunction, identifies the end of the speech as a likely end of a sentence, and segments the text hypothesis "la nacion en la lucha contra el narcotrafico" as a second segment 230 …, para. [0041]; para. [0047]); determining, based at least on the output data, a first location within the text data that is associated with the end of the first sentence and a second location within the text data that is associated with the end of the utterance (fig. 2; The system then performs sentence segmentation 226 on the text hypotheses 216 identified up to time 222, resulting in a segment 228…., para. [0039]; The system's segmenter 226, instead of finding a conjunction, identifies the end of the speech as a likely end of a sentence, and segments the text hypothesis "la nacion en la lucha contra el narcotrafico" as a second segment 230…., para. [0041], end of first segment 228 (segment 1, fig. 2) as end of sentence, end of second segment 230 (segment 2, fig. 2) as end of utterance and end of second sentence); processing, using one or more second models and based at least on the first location and the second location, a first portion of the text data corresponding to the first sentence of the utterance followed by processing a second portion of the text data corresponding to at least the last sentence of the utterance (fig. 2; para. [0025]; para. [0039]; Immediately upon segmenting the segment 1 (at time 234), the system begins machine translation 232 on segment 1…., para. [0040]; The system, upon identifying the second segment 230 at time 238, immediately begins machine translation 232 of the second segment 241…., para. [0041]; para. [0047]); and generating, based at least on the processing the first portion of the text data followed by the second portion of the text data, a response associated with the utterance (fig. 2, elements 232, 246; Immediately upon segmenting the segment 1 (at time 234), the system begins machine translation 232 on segment 1. As illustrated, this process 232 begins at time 234 and ends at time 242 with a machine translation of segment 1 236, which is translated text corresponding to segment 1. Continuing with the example, the machine translation produces "The idea is to work together with the province and" as text corresponding to segment 1. Upon producing the machine translation of segment 1 at time 242, the system immediately begins to output 246 an audio version of the machine translation, para. [0040]) Rangarajan does not explicitly disclose generating by a first model and based at least on the text data, output data representing first probabilities indicating whether each token corresponding to the text data is associated with an end of the first sentence within the utterance and second probabilities indicating whether each token is associated with an end of the utterance that occurs at the end of the last sentence However, these features are taught by Pathak: generating by a first model and based at least on the text data, output data representing first probabilities indicating whether each token corresponding to the text data is associated with an end of the first sentence within the utterance and second probabilities indicating whether each token is associated with an end of the utterance that occurs at the end of the last sentence (fig. 6; para. [0001]; Potential segmentation boundaries 146 are locations in audio that are predicted to correspond to an end of a speech utterance…. The goal of the disclosed embodiments is to identify potential segmentation boundaries at the ends of speech utterances …, para. [0034]; para. [0054]; The language segmentation model then calculates a language segmentation score that represents the probability of an end of speech (EOS) …, para. [0060]-[0061]; The computing system processes the audio with a decoder to recognize speech utterances included in the audio …The computing system also generates an acoustic segmentation score and a language segmentation score associated with the potential segmentation boundary, para. [0075]; para. [0086]-[0087]; Alternatively, when the particular segment comprises multiple sentences, the computing system recognizes one or more sentences within the particular segment and generates one or more punctuation marks to be places at an end of each of the one or more sentences included in the particular segment, para. [0088]); It would have been obvious to one of ordinary skill in the art to combine the teachings of Pathak with the method of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in improving the quality of speech recognition and machine translation (Pathak, para. [0049]) Per claim 2, Rangarajan in view of Pathak discloses the method of claim 1, Pathak discloses wherein the output data further represents third probabilities indicating whether each token is not associated with the end of the first sentence and the end of the utterance (Abstract; para. [0037]; The language segmentation model then calculates a language segmentation score that represents the probability of an end of speech (EOS) …, para. [0060]; If the LM-EOS (language model-end of sentence) or language segmentation score is low, then the system continues to process the audio …, para. [0061], low scores as implying limitation) Per claim 6, Rangarajan in view of Pathak discloses the method of claim 1 Rangarajan discloses wherein the first portion of the text data is processed using the one or more second models based at least on determining the first location and prior to determining the second location (fig. 2; para. [0025]; para. [0039]; para. [0041]; para. [0047]). Per claim 10, Rangarajan discloses the system of claim 7, Rangarajan does not explicitly disclose wherein the first outputs represent first probabilities indicating whether the words are associated with the end of the sentence or the second outputs represent second probabilities indicating whether the words are associated with the end of the utterance However, these features are taught by Pathak: wherein the first outputs represent first probabilities indicating whether the words are associated with the end of the sentence (para. [0060]-[0061]); and the second outputs represent second probabilities indicating whether the words are associated with the end of the utterance (para. [0060]-[0061]) It would have been obvious to one of ordinary skill in the art to combine the teachings of Pathak with the system of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in improving the quality of speech recognition and machine translation (Pathak, para. [0049]). Per claim 19, Rangarajan discloses one or more processors comprising processing circuitry to: generate, using one or more models and based at least on text data corresponding to an utterance that includes one or more words, output data (fig. 2; para. [0017]; As the system receives the stream of speech 212, the system is performing automatic speech recognition 214 on the stream of speech 212. The automatic speech recognition process 214 produces, via a speech recognizer, partial speech hypotheses 216…, para. [0038]; These remainder hypotheses can include words, sentences, paragraphs, or …, para. [0041]); determine, based at least on the output data, a first location within the text data that is associated with the end of the sentence and a second location within the text data that is associated with the end of the utterance (fig. 2; The system then performs sentence segmentation 226 on the text hypotheses 216 identified up to time 222, resulting in a segment 228…., para. [0039]; The system's segmenter 226, instead of finding a conjunction, identifies the end of the speech as a likely end of a sentence, and segments the text hypothesis "la nacion en la lucha contra el narcotrafico" as a second segment 230…., para. [0041], end of first segment 228 (segment 1, fig. 2) as end of sentence, end of second segment 230 (segment 2, fig. 2) as endof utterance and end of second sentence); and process, based at least on the first location and the second location, a first portion of the text data corresponding to the end of the sentence of the utterance followed by a second portion of the text data corresponding to a remainder of the utterance (fig. 2, elements 232, 246; Immediately upon segmenting the segment 1 (at time 234), the system begins machine translation 232 on segment 1. As illustrated, this process 232 begins at time 234 and ends at time 242 with a machine translation of segment 1 236, which is translated text corresponding to segment 1. Continuing with the example, the machine translation produces "The idea is to work together with the province and" as text corresponding to segment 1. Upon producing the machine translation of segment 1 at time 242, the system immediately begins to output 246 an audio version of the machine translation, para. [0040]) Rangarajan does not explicitly disclose output data representing: one or more first probabilities indicating whether the one or more words are associated with an end of a sentence within the utterance, one or more second probabilities indicating whether the one or more words are associated with an end of the utterance or one or more third probabilities indicating whether the one or more words are not associated with the end of the sentence or the end of the utterance; However, these features are taught by Pathak: output data representing: one or more first probabilities indicating whether the one or more words are associated with an end of a sentence within the utterance (fig. 6; Potential segmentation boundaries 146 are locations in audio that are predicted to correspond to an end of a speech utterance…. The goal of the disclosed embodiments is to identify potential segmentation boundaries at the ends of speech utterances …, para. [0034]; para. [0054]; para. [0060]-[0061]; The computing system also generates an acoustic segmentation score and a language segmentation score associated with the potential segmentation boundary, para. [0075]; para. [0086]-[0087]; Alternatively, when the particular segment comprises multiple sentences, the computing system recognizes one or more sentences within the particular segment and generates one or more punctuation marks to be places at an end of each of the one or more sentences included in the particular segment, para. [0088]); one or more second probabilities indicating whether the one or more words are associated with an end of the utterance (fig. 6; Potential segmentation boundaries 146 are locations in audio that are predicted to correspond to an end of a speech utterance…. The goal of the disclosed embodiments is to identify potential segmentation boundaries at the ends of speech utterances …, para. [0034]; para. [0054]; para. [0060]-[0061]; para. [0075]; para. [0086]-[0088]); and one or more third probabilities indicating whether the one or more words are not associated with the end of the sentence or the end of the utterance (Abstract; para. [0060]-[0061], low scores as implying limitation) It would have been obvious to one of ordinary skill in the art to combine the teachings of Pathak with the processors of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in improving the quality of speech recognition and machine translation (Pathak, para. [0049]) Per claim 20, Rangarajan in view of Pathak discloses the one or more processors of claim 19, Rangarajan discloses wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more visual language models (VLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data (para. [0042]; para. [0044]; para. [0047]); a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. 3. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Rangarajan in view of Pathak as applied to claim 1 above, and further in view of Bijwadia et al - “Text Injection for Capitalization and Turn-Taking Prediction in Speech Models” (“Bijwadia”) Per claim 3, Rangarajan in view of Pathak discloses the method of claim 1, Rangarajan does not explicitly disclose generating, using one or more third models and based at least on the text data, second output data representative of whether each token is associated with a lowercase word or an uppercase word, wherein the determining the first location and the second location is further based at least on the second output data However, this feature is taught by Bijwadia (sec. 3.2; sec. 4.1; The capitalization sequence is defined as follows: each token is either ⟨cap⟩ (capitalized) or ⟨non-cap) (not capitalized), based on the corresponding wordpiece in the ASR transcript …, sec. 5.3) It would have been obvious to one of ordinary skill in the art to combine the teachings of Bijwadia with the method of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in restoring the correct case of noisy text (Bijwadia, sec. 3.1). 4. Claims 12 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Rangarajan in view of Bijwadia Per claim 12, Rangarajan discloses the system of claim 7, Rangarajan does not explicitly disclose wherein the one or more processors are further to: determine, using one or more second models and based at least on the text data, a third output indicating whether the at least one or more words from the words are at least one of lowercase or uppercase, wherein the first location and the second location are further determined based at least on the third output. However, this feature is taught by Bijwadia (sec. 3.2; sec. 4.1; sc. 5.2; The capitalization sequence is defined as follows: each token is either ⟨cap⟩ (capitalized) or ⟨non-cap) (not capitalized), based on the corresponding wordpiece in the ASR transcript …, sec. 5.3) It would have been obvious to one of ordinary skill in the art to combine the teachings of Bijwadia with the system of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in restoring the correct case of noisy text (Bijwadia, sec. 3.1) Per claim 13, Rangarajan in view of Bijwadia discloses the system of claim 12, Bijwadia discloses wherein the third output represents at least: one or more first probabilities indicating whether the one or more words are lowercase (sec. 4.1; sec. 5.2; sec. 5.3); and one or more second probabilities indicating whether the one or more words are uppercase (sec. 4.1; sec. 5.2; sec. 5.3). 5. Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Rangarajan in view of Pathak as applied to claim 1 above, and further in view of Hassid et al - US 2023/0335111 A1 (“Hassid”) Per claim 4, Rangarajan in view of Pathak discloses the method of claim 1, Rangarajan does not explicitly disclose generating, using one or more third models and based at least on the text data, second output data representative of whether each token is associated with one or more types of punctuation marks, wherein the determining the first location and the second location is further based at least on the second output data However, this feature is taught by Hassid (para. [0092]; para. [0097]) It would have been obvious to one of ordinary skill in the art to combine the teachings of Hassid with the method of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in detecting the end of an input string (Hassid, para. [0097]). 6. Claims 14 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Rangarajan in view of Hassid Per claim 14, Rangarajan discloses the system of claim 7, Rangarajan does not explicitly disclose wherein the one or more processors are further to: determine, using one or more second models and based at least on the text data, a third output indicating whether at least the one or more words from the words are associated with one or more types of punctuation marks, wherein the first location and the second location are further determined based at least on the third output However, this feature is taught by Hassid (para. [0092]; para. [0097]) It would have been obvious to one of ordinary skill in the art to combine the teachings of Hassid with the system of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in in detecting the end of an input string (Hassid, para. [0097]). Per claim 15, Rangarajan in view of Hassid discloses the system of claim 14, Hassid discloses wherein the third output represents at least: one or more first probabilities indicating whether the one or more words are associated with one or more first types of punctuation marks (para. [0092]; para. [0097]); and one or more second probabilities indicating whether the one or more words are associated with one or more second types of punctuation marks (para. [0092]; para. [0097]). 7. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Rangarajan in view of Pathak as applied to claim 1 above, and further in view of Lineback et al - US 2025/0330665 a1 (“Lineback”) Per claim 5, Rangarajan discloses method of claim 1, further comprising: Rangarajan does not explicitly disclose generating, using one or more first encoders and based at least on the audio data, one or more first embeddings, generating, using one or more second encoders and based at least on the text data, one or more second embeddings or generating input data based at least on the one or more first embeddings and the one or more second embeddings, wherein the generating the output data uses the first model and is based at least on the input data However, these features are taught by Lineback: generating, using one or more first encoders and based at least on the audio data, one or more first embeddings (fig. 9A; para. [0128]; the neural network 908 can use an audio signal (e.g., audio data) in the one or more media content items 906 to generate an embedding 910B …, para. [0129]); generating, using one or more second encoders and based at least on the text data, one or more second embeddings (fig. 9A; para. [0128]; The neural network 908 can use a text signal (e.g., closed caption data, metadata, etc.) in the one or more media content items 906 to generate an embedding 910N representing and/or encoding information from the text signal …, para. [0129]); and generating input data based at least on the one or more first embeddings and the one or more second embeddings, wherein the generating the output data uses the first model and is based at least on the input data (fig. 9A; para. [0129]). It would have been obvious to one of ordinary skill in the art to combine the teachings of Lineback with the method of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in determining one or more segment categories for media content (Lineback, para. [0135]). 8. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Rangarajan in view of Lineback Per claim 16, Rangarajan discloses system of claim 7, Rangarajan does not explicitly disclose wherein the one or more processors are further to: generate, using one or more encoders and based at least on audio data, one or more first embeddings, generate, using the one or more encoders and based at least on the text data, one or more second embeddings or generate input data based at least on the one or more first embeddings and the one or more second embeddings, wherein the determination of the output is based at least on the input data However, these features are taught by Lineback: wherein the one or more processors are further to: generate, using one or more encoders and based at least on audio data, one or more first embeddings (fig. 9A; para. [0128]; the neural network 908 can use an audio signal (e.g., audio data) in the one or more media content items 906 to generate an embedding 910B …, para. [0129]); generate, using the one or more encoders and based at least on the text data, one or more second embeddings (fig. 9A; para. [0128]; The neural network 908 can use a text signal (e.g., closed caption data, metadata, etc.) in the one or more media content items 906 to generate an embedding 910N representing and/or encoding information from the text signal …, para. [0129]); and generate input data based at least on the one or more first embeddings and the one or more second embeddings, wherein the first output and the second outputs are determined based at least on the input data (fig. 9A; para. [0129]; para. [0335]). It would have been obvious to one of ordinary skill in the art to combine the teachings of Lineback with the system of Rangarajan in arriving at the missing features of Rangarajan, because such combination would have resulted in determining one or more segment categories for media content (Lineback, para. [0135]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO 892 form. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLUJIMI A ADESANYA whose telephone number is (571)270-3307. The examiner can normally be reached Monday-Friday 8:30-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /OLUJIMI A ADESANYA/Primary Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Jun 26, 2024
Application Filed
May 07, 2026
Non-Final Rejection mailed — §101, §102, §103
Jul 16, 2026
Applicant Interview (Telephonic)
Jul 16, 2026
Examiner Interview Summary
Jul 17, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743415
MACHINE LEARNING MODEL PROMPT HYDRATION VIA PROMPT REGISTRY AND CONTEXT STORE
2y 7m to grant Granted Sep 22, 2026
Patent 12724967
MULTI-TASK LEARNING FOR NATURAL LANGUAGE PROCESSING TASKS USING A SHARED PRE-TRAINED LANGUAGE MODEL
2y 6m to grant Granted Sep 01, 2026
Patent 12717825
METHOD AND SYSTEM FOR AUTOMATED INFORMATION MANAGEMENT
2y 11m to grant Granted Aug 25, 2026
Patent 12717831
TRANSCRIPT SEGMENTATION AND SUMMARIZATION
2y 5m to grant Granted Aug 25, 2026
Patent 12694208
DOMAIN-SPECIFICITY PREDICTION FOR NATURAL LANGUAGE PROCESSING
3y 4m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
66%
Grant Probability
93%
With Interview (+26.9%)
3y 5m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 678 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month