Prosecution Insights
Last updated: September 17, 2026
Application No. 19/044,674

SPEAKER MODULES IN AN AUDIO-VISUAL REAL-TIME TRANSCRIPTION SYSTEM

Non-Final OA §101§102§103
Filed
Feb 04, 2025
Examiner
LOWEN, NICHOLAS DANIEL
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Rtc Vision Ltd.
OA Round
1 (Non-Final)
64%
Grant Probability
Moderate
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 64% of resolved cases
64%
Career Allowance Rate
9 granted / 14 resolved
+2.3% vs TC avg
Strong +69% interview lift
Without
With
+68.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
17 currently pending
Career history
37
Total Applications
across all art units

Statute-Specific Performance

§101
35.4%
-4.6% vs TC avg
§103
45.3%
+5.3% vs TC avg
§102
15.1%
-24.9% vs TC avg
§112
3.1%
-36.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION This communication is in response to the Application filed on 2/4/2025. Claims 1-19 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 13, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 11/18/2025 and 6/25/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 2, 4-7, and 16 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US Patent Publication US 11615781 B2 (Braga). Regarding Claims 1 and 16, Braga teaches A method of target speaker recognition and transcription of speech, comprising: (Another aspect of the disclosure provides a method for transcribing speech from audio-visual data.) (Col. 2, Lines 28-29). (Audio-visual (A/V) automated speech recognition (ASR) is able to make conventional ASR more robust by leveraging video data of a face of a speaker in addition to audio data of a spoken from the speaker.) (Col. 3, Lines 50-53). Claim 16 presents the alternative A computer system, comprising one or more microphone devices; one or more camera devices; a processor configured to execute stored executable instructions; and a non-transitory computer readable medium storing executable instructions that, when executed by a processor, cause the computer system to perform a method of target speaker recognition and transcription, the method comprising: (For instance, an audio capture device 116, 116a (e.g., an array of one or more microphones) is configured to capture utterances 14 spoken by the participants 10a-g and convert the captured utterances 14 into audio data that corresponds to the audio portion 210 of the audio-visual data 204. On the other hand, an image capture device 116, 116b (e.g., one or more cameras) is configured to capture image data that corresponds to the video portion 220 of the audio-visual data 204.) (Col. 5, Liens 22-30). (The processor 510 can process instructions for execution within the computing device 500, including instructions stored in the memory 520 or on the storage device 530 to display graphical information for a graphical user interface (GUI) on an external input/output device, such as display 580 coupled to high speed interface 540.) (Col. 10, Lines 61-67). receiving an audio stream from a first input; (Simply put, for multi-speaker A/V ASR tasks, the single A/V ASR model is configured to receive audio-visual inputs with multiple face tracks and an audio track) (Col. 4, Lines 24-27). The model receives audio input. receiving a video stream from a second input; (Simply put, for multi-speaker A/V ASR tasks, the single A/V ASR model is configured to receive audio-visual inputs with multiple face tracks and an audio track) (Col. 4, Lines 24-27). The model receives video input computing a plurality of initial audio features from the audio stream; (For instance, an audio capture device 116, 116a (e.g., an array of one or more microphones) is configured to capture utterances 14 spoken by the participants 10a-g and convert the captured utterances 14 into audio data that corresponds to the audio portion 210 of the audio-visual data 204. On the other hand, an image capture device 116, 116b (e.g., one or more cameras) is configured to capture image data that corresponds to the video portion 220 of the audio-visual data 204.) (Col. 5, Lines 22-30). The input audio is converted to audio data that can be paired with the video data. computing a plurality of initial video features from the video stream; (For instance, an audio capture device 116, 116a (e.g., an array of one or more microphones) is configured to capture utterances 14 spoken by the participants 10a-g and convert the captured utterances 14 into audio data that corresponds to the audio portion 210 of the audio-visual data 204. On the other hand, an image capture device 116, 116b (e.g., one or more cameras) is configured to capture image data that corresponds to the video portion 220 of the audio-visual data 204.) (Col. 5, Lines 22-30).’ The input video is converted to video data that can be paired with the audio. applying a neural network to the plurality of initial video features, thereby generating a plurality of extended video features; (The encoder 260 is associated with an encoder frontend that includes an attention mechanism 270. The attention mechanism 270 may be associated with an attention layer in the encoder portion 260 of the neural network model 200.) (Col. 7, Lines 39-40). (For each video face track 230, the attention mechanism 270 determines a corresponding confidence score indicating a likelihood that the face of the respective person associated with the corresponding video face track 230 includes a speaking face of the audio track 210.) (Col. 7, Lines 56-60). A confidence score is generated by a neural network represented an extended video feature. determining a correlation between the plurality of initial audio features and the plurality of extended video features, thereby generating a plurality of enhanced audio features; (In some implementations, the encoder 260 concatenates the attention-weighted visual feature vector 272 that soft-selects the video face track 230 associated with the active speaking face with the acoustic feature vector to provide a corresponding combined feature vector at each time step. The combined feature vector at each time step indicates an encoding of the audio track 210 and the video face track 230 among the plurality of video face tracks that is associated with the highest confidence score.) (Col. 8, Lines 16-24) The video data, confidence scores, and audio are concatenated to create a combined feature vector which corresponds the audio to the appropriate face in the video. This combined feature vector acts as the enhanced audio feature. processing the plurality of initial audio features, the plurality of initial video features and the plurality of enhanced audio features, thereby generating a plurality of target speaker features; (In some implementations, the encoder 260 concatenates the attention-weighted visual feature vector 272 that soft-selects the video face track 230 associated with the active speaking face with the acoustic feature vector to provide a corresponding combined feature vector at each time step. The combined feature vector at each time step indicates an encoding of the audio track 210 and the video face track 230 among the plurality of video face tracks that is associated with the highest confidence score. Accordingly, at each time step, the decoder portion 280 is configured to decode the combined feature vector to determine a corresponding speech recognition result 248 of the audio track 210.) (Col. 8, Lines 16-27). (In some examples, the AV-ASR 200 model is further configured provide speaker labels 255 to the transcription 250 to identify a source of the transcribed content. For instance, labeling a speaker of the transcribed content may be referred to as speaker diarization to answer both “who spoke what” and “who spoke when”.) (Col. 8, Lines 43-48). The audio data, video data, and confidence score are used to generate a target speaker features by associating the speech in the audio with the corresponding face and creating labels for the speech/speakers generating an output. (The display 111 associated with the user device 110 may display the transcription 250 generated by the AV-ASR model 200. The AV-ASR model 200 may stream the transcription 250 in real time for output on the display 111 and/or on displays associated with remotely located participants) (Col. 6, Lines 23-27). The output is a displayed transcription to the user device. Regarding Claim 2, Braga teaches the method of claim 1. Braga further teaches receiving a data stream from a third input; and computing a plurality of initial auxiliary features from the data stream, wherein generating a plurality of target speaker features comprises processing the plurality of initial audio features, the plurality of initial video features, the plurality of enhanced audio features and the plurality of initial auxiliary features. (In some implementations, the encoder 260 concatenates the attention-weighted visual feature vector 272 that soft-selects the video face track 230 associated with the active speaking face with the acoustic feature vector to provide a corresponding combined feature vector at each time step. The combined feature vector at each time step indicates an encoding of the audio track 210 and the video face track 230 among the plurality of video face tracks that is associated with the highest confidence score. Accordingly, at each time step, the decoder portion 280 is configured to decode the combined feature vector to determine a corresponding speech recognition result 248 of the audio track 210. The speech recognition result 248 at each time step may include a probability distribution over possible recognition results. In examples when the AV-ASR model 200 is the Audio-Visual RNN-T model, the model 200 may emit the speech recognition result 248 at each time step in a streaming fashion.) (Col. 7, Lines 56-66). A third input is included in the form of timing information for the video and audio. The time is used associate the video and audio vectors at the same moment and area used to create the concatenated representation. Regarding Claim 4, Braga teaches the method of claim 1. Braga further teaches wherein generating the plurality of enhanced audio features comprises performing a computation solely on the plurality of extended video features. (For each video face track 230, the attention mechanism 270 determines a corresponding confidence score indicating a likelihood that the face of the respective person associated with the corresponding video face track 230 includes a speaking face of the audio track 210. In some implementations, the attention mechanism 270 includes a softmax layer having an inverse temperature parameter configured to cause the attention mechanism 270 to converge to a hard-decision rule of selecting the video face track 230 of the plurality of video face tracks 230a-c associated with the highest confidence score as the speaking face of the audio track 110.) (Col. 7, Lines 56-66). The confidence score that a face is speaking is calculated for each video. Regarding Claim 5, Braga teaches the method of claim 1. Braga further teaches wherein generating the plurality of enhanced audio features comprises one or more of: applying at least one cross-attention layer to the plurality of initial audio features and the plurality of extended video features; (In some implementations, the encoder 260 concatenates the attention-weighted visual feature vector 272 that soft-selects the video face track 230 associated with the active speaking face with the acoustic feature vector to provide a corresponding combined feature vector at each time step. The combined feature vector at each time step indicates an encoding of the audio track 210 and the video face track 230 among the plurality of video face tracks that is associated with the highest confidence score.) (Col. 8, Lines 16-24) An attention layer concatenates the visual and audio vectors together. applying a recurrent neural network to the plurality of initial audio features, the plurality of initial video features and a plurality of previously computed target speaker features; and (In examples when the AV-ASR model 200 is the Audio-Visual RNN-T model, the model 200 may emit the speech recognition result 248 at each time step in a streaming fashion. A speech recognition result may include a character, a space, a word-piece, or a word. The multiple speech recognition results 248 may combine to provide the transcription 250 of the audio track 210. Thus, the Audio-Visual RNN-T model is capable of streaming a transcription 250 of the audio track 210 in real time.) (Col. 8, Lines 30-38). The neural network used to create the target speaker features is a recurrent neural network. utilizing a decision process to the plurality of initial audio features, the plurality of initial video features and a plurality of previously computed target speaker features. (For instance, labeling a speaker of the transcribed content may be referred to as speaker diarization to answer both “who spoke what” and “who spoke when”. Accordingly, by leveraging the video portion 220 of the audio-visual data 204, the AV-ASR model 200 may provide diarization results that include a corresponding speaker label 255 assigned to each segment of the transcription 250 to identify “who spoke what” and “who spoke when”.) (Col. 8, Lines 45-53). Labels are created for the speech data from this process. Regarding Claim 6, Braga teaches the method of claim 1. Braga further teaches wherein generating the plurality of target speaker features comprises applying a recurrent neural network to the plurality of initial audio features, the plurality of initial video features and the plurality of enhanced audio features. (The AV-ASR model 200 includes an encoder portion (“encoder”) 260 and a decoder portion (“decoder”) 280. The AV-ASR model 200 may include a sequence-to-sequence model. In some examples, the AV-ASR model 200 includes an Audio-Visual Recurrent Neural Network-Transducer (RNN-T) model.) (Col. 7, Lines 29-34). (In some implementations, the encoder 260 concatenates the attention-weighted visual feature vector 272 that soft-selects the video face track 230 associated with the active speaking face with the acoustic feature vector to provide a corresponding combined feature vector at each time step. The combined feature vector at each time step indicates an encoding of the audio track 210 and the video face track 230 among the plurality of video face tracks that is associated with the highest confidence score. Accordingly, at each time step, the decoder portion 280 is configured to decode the combined feature vector to determine a corresponding speech recognition result 248 of the audio track 210.) (Col. 8, Lines 16-27). The RNN model uses audio, video, and associations between the speaker in them in order to create transcription data corresponding to the speaker. Regarding Claim 7, Braga teaches the method of claim 1. Braga further teaches wherein generating an output comprises: applying a machine learning model to the plurality of target speaker features, thereby generating a target speaker label; and outputting the target speaker label. (In some examples, the AV-ASR 200 model is further configured provide speaker labels 255 to the transcription 250 to identify a source of the transcribed content. For instance, labeling a speaker of the transcribed content may be referred to as speaker diarization to answer both “who spoke what” and “who spoke when”. Accordingly, by leveraging the video portion 220 of the audio-visual data 204, the AV-ASR model 200 may provide diarization results that include a corresponding speaker label 255 assigned to each segment of the transcription 250 to identify “who spoke what” and “who spoke when”. FIG. 3 shows an example training process 300 for training the encoder portion 260 of the AV-ASR model 200 to learn how to gate a correct video face track 230 for each segment of the audio track to aid in speech recognition. The encoder portion 260 is trained on a training data set 302 that includes a training audio track 210T, a first training video face track 230Ta, and one or more second training video face tracks 230Tb. The training audio track 210 includes one or more spoken utterances.) (Col. 8, Lines 43-62). The neural network model is trained, making it a machine learning model. The model generates labels for target speakers which are used in creating the transcription data. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 3 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 11615781 B2 (Braga). in view of China Patent Publication CN 112967713 A (Wang et al.). Regarding Claim 3, Braga teaches the system of claim 1. Braga does not explicitly teach: wherein the neural network comprises at least one convolution layer. However, Wang et al. teaches wherein the neural network comprises at least one convolution layer. (processing the voice enhanced spectrum by the second audio coder to obtain the audio context vector; processing the original video characteristic by the second video coder to obtain the video context vector; the second audio coder and the second video coder are respectively composed of a layer of time convolution block and two layers of Skip LSTM;) (Page 2, Paragraph 12). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the audio/video transcription as taught by Braga to include a convolution layer in the neural network as taught by Wang et al. This would have been an obvious addition/substitution as convolution neural networks are commonly used for extracting features from video information (Wang et al. Page 2, Paragraph 3). Regarding Claim 11, Braga teaches the system of claim 1. Furthermore, Braga teaches wherein computing a plurality of initial audio features comprises: receiving an audio signal from a microphone device; (For instance, an audio capture device 116, 116a (e.g., an array of one or more microphones) is configured to capture utterances 14 spoken by the participants 10a-g and convert the captured utterances 14 into audio data that corresponds to the audio portion 210 of the audio-visual data 204.) (Col. 5, Lines 22-27). Braga does not explicitly teach: computing a spectrogram of the audio signal; and applying at least one convolution layer to the spectrogram. However, Wang et al. teaches computing a spectrogram of the audio signal; (a first decoding module, for decoding the first fusion feature by the first audio decoder, obtaining voice enhanced spectrum;) (Page 3, Paragraph 14). (S203: Referring to FIG. 2, the first fusion feature into the first audio decoder, the fusion feature for decoding, the feature after decoding into the full connection layer to output and initial voice spectrum graph of the same latitude of the speech enhancement spectrum.) (Page 5, Paragraph 8). Wang et al. creates a spectrogram for the input audio. and applying at least one convolution layer to the spectrogram. (a second extracting module, for processing the voice enhanced spectrum by the second audio encoder to obtain audio context vector; processing the original video characteristic by the second video coder to obtain the video context vector; the second audio coder and the second video coder are respectively composed of a layer of time convolution block and two layers of Skip LSTM;) (Page 3, Paragraph 15). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the audio/video transcription as taught by Braga to create a spectrogram and apply a convolution layer as taught by Wang et al. This would have been an obvious addition as it is a method of combining video and audio data while maintaining timestamp information which has improved performance over a common LSTM (Wang et al. Page 5, Paragraph 6). Claims 8 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 11615781 B2 (Braga). in view of US Patent Publication US 20250131915 A1 (Xie et al.). Regarding Claim 8, Braga teaches the system of claim 1. Braga does not explicitly teach: wherein generating an output comprises: applying positional encoding to the plurality of enhanced audio features; applying at least one transformer encoder layer, thereby generating a sequence of encoded vectors; applying a transformer decoder layer to the sequence of encoded vectors, thereby generating a sequence of textual tokens; concatenating the sequence of textual tokens, thereby generating a text transcription; and outputting the text transcription. However, Xie et al. teaches wherein generating an output comprises: applying positional encoding to the plurality of enhanced audio features; (The output data input to the encoder component 310 can facilitate (e.g., enable) sinusoidal positional encoding of the data by the encoder component 310. The sinusoidal positional encoding of the data can be utilized to incorporate the temporal information of the audio features of the audio content (e.g., the spectrogram information representative of the audio content) into the speech recognition model 104, which can assist the speech recognition model 104 to understand the sequence of the input data (e.g., the spectrogram information representative of the audio content).) (Paragraph 82). Sinusoidal positional encoding is used to extract audio features. applying at least one transformer encoder layer, thereby generating a sequence of encoded vectors; (As disclosed, in some embodiments, the data input to the encoder component 310 can be labeled speech-related data (e.g., the textual transcript data, such as labeled textual transcript data, and/or spectrogram information) that can be representative of the mixed language dataset for analysis and processing by the speech recognition model 104 to facilitate training the speech recognition model 104 and/or performing of speech recognition on the input data to generate a transcription comprising the transcribed textual data (e.g., comprising textual words in mixed languages) that can be representative of and/or can correspond to the spoken mixed language words of the audio content that can be representative of the words in the mixed language dataset (e.g., the enhanced mixed language dataset).) (Paragraph 79). An encoder layer is applied to input speech data. applying a transformer decoder layer to the sequence of encoded vectors, thereby generating a sequence of textual tokens; (The speech recognition model 104, employing the encoder component 310 and decoder component 312, can perform next-token prediction to predict a next token in a sequence of tokens based at least in part on the results of analyzing and processing (e.g., encoding and decoding) of the input data, wherein the next token can be representative of a next word or subword in a sequence of words (e.g., a sentence (or other collection of words) of the spoken mixed language words of the audio content) that can be represented by the sequence of tokens, such a described herein.) (Paragraph 79). The decoder tokenizes the input sequence. concatenating the sequence of textual tokens, thereby generating a text transcription; and (The output of the encoder component 310 can be associated with the respective inputs (e.g., respective input ports) of the respective decoder blocks (e.g., 320, 322, and/or 324) of the decoder component 312. The decoder blocks (e.g., 320, 322, and/or 324) can generate output transcriptions (e.g., textual transcriptions) representative of the mixed language words of the mixed language dataset based at least in part on the results of processing or analyzing the encoded audio representations of the audio features.) (Paragraph 81). Textual transcriptions are created from the output. outputting the text transcription. (This disclosure relates generally to systems, mechanisms, methods, and techniques that desirably (e.g., suitably, accurately, quickly, efficiently, reliably, enhancedly, or optimally) can perform speech recognition (e.g., automatic speech recognition (ASR)) on audio of mixed-language speech of a person, convert the respective recognized spoken words (e.g., mixed language words) of the audio to respective textual words or characters that can correspond to, and be representative of, the spoken words, and present the textual words or characters as an output) (Paragraph 20). The generated transcription is provided as output It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the audio/video transcription as taught by Braga to deploy a positional encoder/decoder architecture as taught by Xie et al. This would have been an obvious improvement as it allows the speech recognition model to understand the sequence of the tokens in the audio data (Xie et al. Paragraph 82). Regarding Claim 9, Braga in view of Xie et al. teaches the system of claim 8. Furthermore, Braga teaches applying a machine learning model to the plurality of target speaker features and at least one of a previously obtained text transcription and a previously obtained plurality of textual tokens, thereby generating a target speaker label; (In some examples, the AV-ASR 200 model is further configured provide speaker labels 255 to the transcription 250 to identify a source of the transcribed content. For instance, labeling a speaker of the transcribed content may be referred to as speaker diarization to answer both “who spoke what” and “who spoke when”. Accordingly, by leveraging the video portion 220 of the audio-visual data 204, the AV-ASR model 200 may provide diarization results that include a corresponding speaker label 255 assigned to each segment of the transcription 250 to identify “who spoke what” and “who spoke when”.) (Col. 8, Lines 43-53). (The first training video face track 230Ta is paired with a ground-truth correct face label 232C. Each second training video face track 230Tb includes an incorrect speaking face of the one or more spoken utterances of the audio track 210. Each second training video face track 230Tb is paired with a ground-truth incorrect face label 232I.) (Col. 8, Lines 64-67) The AV-ASR model is a machine learning model as it is a trained neural network. The model performs diarization on inputs in order to create label information for who spoke what and when they spoke it. Thus, speaker features from previous inputs (the last sentence spoken) are constantly being used to generate the transcription information. outputting the target speaker label. (ASR model 200 may provide diarization results that include a corresponding speaker label 255 assigned to each segment of the transcription 250 to identify “who spoke what” and “who spoke when”.) (Col. 8, Lines 43-53). The labels are provided as output from the model. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 11615781 B2 (Braga) in view of US Patent Publication US 20250131915 A1 (Xie et al.) and US Patent Publication US 12033641 B2 (Rikhye et al.). Regarding Claim 10, Braga in view of Xie et al. teaches the system of claim 9. Braga in view of Xie et al. does not explicitly teach: further comprising: determining whether a condition is met, wherein the condition comprises checking whether the target speaker label equals to a predetermined speaker label; and responsive to determining that the condition is met, skipping the step of outputting the text transcription. However, Rikhye et al. teaches further comprising: determining whether a condition is met, wherein the condition comprises checking whether the target speaker label equals to a predetermined speaker label; and (At block 712, the system determines whether the speaker spoke the utterance. For example, the system can determine whether a registered and/or verified speaker spoke the utterance based on the speaker verification output generated at block 704. If so, the system proceeds to block 714. If not, the process ends.) (Col. 23, Lines 1-6). The system determines if the speaker meets the condition of being a registered/verified speaker. responsive to determining that the condition is met, skipping the step of outputting the text transcription. (In some implementations, the system can determine action(s) corresponding to the particular keyphrase by processing the text representation of the utterance by processing the text representation of the utterance using a NLU model to generate an intent of the utterance.) (Col. 23, Lines 15-19). The input audio utterance how processing performed on it if the condition is met. This can be seen in the flowchart shown in Fig. 7 It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the audio/video transcription as taught by Braga in view of Xie et al. to avoid performing processing on utterances from certain speakers as taught by Rikhye et al. This would have been an obvious improvement avoid unnecessary processing (Rikhye et al. Col. 11, Lines 60-67). Claims 12, 14, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 11615781 B2 (Braga) in view of US Patent Publication US 20190130594 A1 (Seyfi et al.). Regarding Claim 12, Braga teaches the system of claim 1. Braga does not explicitly teach: wherein the neural network comprises at least one convolution layer. However, Seyfi et al. teaches wherein computing a plurality of initial video features comprises: receiving a frame sequence from a camera device; responsive to detection of one or more faces within a frame in the frame sequence, determining a bounding box for each detected face; (For each of the detected moving areas 214 generated by motion detection module 204, a CNN-based face detection module 206 can be used to detect some or all faces with the detected moving area. … Face detection module 206 generates a set of detected faces 216 and the corresponding bounding box locations. Note that face tracking module 210 can be used to track previously detected faces of processed video images based on the current output of face detection module 206 associated with a newly processed video image.) (Paragraph 49). Each detected face is identified and a bounding box is created for it. and extracting from the frame a 2-dimensional (2D) array of pixels limited by the respective bounding box for each detected face. (For each detected face in a processed video frame, process 600 locates the corresponding bounding box of the detected face as a reference box and the detected face image within the bounding box as a search block (step 602). Next, in a subsequent unprocessed video frame, process 600 uses the search block to search within a search window of a predetermined size and centered around the same location of the reference box in the unprocessed video frame (step 604). More specifically, within the search window, multiple locations (e.g., 64 different locations) of the size of the reference box can be searched. At each of the search locations within the search window, the search block is compared with the image patch within the reference box (step 606). Hence, process 600 identifies the same detected face in the unprocessed video frame at a search location where the best match between the search block and the corresponding image patch is found (step 608).) (Paragraph 66). The bounding box size is extracted to recreate when searching for the face in the subsequent frames. The pixels within the box are used for identifying if the faces in the bounding boxes on each frame are the same face. It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the audio/video transcription as taught by Braga to identify faces and isolate them with bounding boxes for further processing as taught by Seyfi et al. This would have been an obvious improvement as Braga is already performing its processing videos that have been separated into individual faces. It is beneficial to separate out the face as then processing can only be performed on those pixels, thus cutting down the amount of work that must be done (Seyfi et al. Paragraph 7). Regarding Claim 14, Braga in view of Seyfi et al. teaches the system of claim 12. Furthermore, Seyfi et al. teaches further comprising: detecting a particular face at the center of each frame in the frame sequence of a configured duration; (In face-detection-and-tracking subsystem 200, to find the best pose of each tracked person, it is necessary to track the location of the track person in each frame of the captured video, from the time the person is initially detected in the video until the time the person is determined to have disappeared from the video.) (Paragraph 56). (Next, for each detected face in the processed video frame, face tracking module 212 uses a corresponding search block to search around the estimated new location (i.e., the search location) of the detected face in the unprocessed video frame. Note that due to the improved accuracy of the estimated location, face tracking module 212 does not need to search many positions around the estimated location. At each of the search locations centered around the estimated location, the search block is compared with the image patch within the search box. Hence, the same detected face can be identified in the unprocessed video frame at a search location where the best match between the search block and the corresponding image patch is found.) (Paragraph 69). The bounding boxes are created around each face, in successive frames it searches for the same face within the center of the bounding box. tracking the detected face across subsequent frames; and (In some embodiments, face tracking module 212 can be configured to locate and label the tracked faces within the in-between video frames without the need of applying face-detection module 206. In some embodiments, face tracking module 212 is configured to determine the location of each tracked face within an unprocessed video frame (e.g., Frame 2) immediately following a processed frame 504 (e.g., Frame 1) based on the determined location of the tracked face in the processed frame (e.g., Frame 1).) (Paragraph 65). The face is tracked throughout subsequent frames. prioritizing the target speaker features corresponding to the tracked face across subsequent frames. (Returning to FIG. 3, if it is determined at step 310 that the tracked new person has disappeared from the video, process 300 subsequently transmits a detected face of the track new person corresponding to the determined best pose to a server (step 312). Note that transmitting just the face image corresponding to the best pose without sending all of the detected faces significantly reduces network bandwidth and storage space. Otherwise, if the tracked new person remains in the video, process 300 continues to track this person through the subsequent video images and update the best pose for this person (step 314)) (Paragraph 62). The tracked person is prioritized until the disappear from the video information. Regarding Claim 15, Braga in view of Seyfi et al. teaches the system of claim 12. Furthermore, Seyfi et al. teaches further comprising: receiving a signal to select a particular face from a third input; (Next, process 300 determines that a new person has appeared in the video based on the detected faces (step 306). For example, process 300 can perform a face association operation between a set of labeled detected faces in an immediate preceding video image and the set of unlabeled bounding boxes in the current video image. The process subsequently identifies each of the detected faces not associated with a previously detected face as a new person.) (Paragraph 60). A signal to track a particular person appear in the form of a new person appearing in the video. In this case the system begins tracking that face. tracking the selected face across subsequent frames; and (Next, process 300 tracks the new person through subsequent video images in the capture video (step 308). For example, process 300 can detect a sequence of new locations of the new person in the subsequent video images.) (Paragraph 60), The new face is tracked across subsequent frames. prioritizing the target speaker features corresponding to the tracked face across subsequent frames. (Returning to FIG. 3, if it is determined at step 310 that the tracked new person has disappeared from the video, process 300 subsequently transmits a detected face of the track new person corresponding to the determined best pose to a server (step 312). Note that transmitting just the face image corresponding to the best pose without sending all of the detected faces significantly reduces network bandwidth and storage space. Otherwise, if the tracked new person remains in the video, process 300 continues to track this person through the subsequent video images and update the best pose for this person (step 314)) (Paragraph 62). This person is prioritized until the disappear from the video data. Claims 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 11615781 B2 (Braga) in view of US Patent Publication US 11748579 B2 (Cowburn et al.). Regarding Claim 17, Braga teaches the system of claim 16. Braga does not explicitly teach: further comprising one of an augmented reality display and virtual reality display, wherein generating an output comprises rendering a text transcription on the one of an augmented reality display, an auxiliary screen, an auxiliary projector or a virtual reality display. However, Cowburn et al. teaches further comprising one of an augmented reality display and virtual reality display, (Disclosed is an augmented reality system to generate and cause display of an augmented reality interface at a client device. Various embodiments may detect speech, identify a source of the speech, transcribe the speech to a text string, generate a speech bubble based on properties of the speech and that includes a presentation of the text string, and cause display of the speech bubble at a location in the augmented reality interface based on the source of the speech.) (Col. 2, Lines 33-40). Cowburn et al. teaches a augmented reality display. wherein generating an output comprises rendering a text transcription on the one of an augmented reality display, an auxiliary screen, an auxiliary projector or a virtual reality display. (Operation 712 may be performed by the presentation module 602. At operations 712, the presentation module 602 causes display of the speech bubble at a position in the presentation of the space, based on the location of the source of the speech. In some example embodiments, the presentation module 602 identifies the position to display the speech bubble based on the location of the source of the speech, as well as locations of significant elements in the presentation. For example, the presentation module 602 may identify a position in the presentation of the space that does not include any significant elements (e.g., faces). The presentation module 602 may thereby display the speech bubble at the position without obstructing any significant elements in the presentation.) (Col. 15, Lines 52-65). The system displays a speech bubble (representing transcription data) on the augmented reality display. It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the audio/video transcription as taught by Braga to display the output on an augmented reality display as taught by Cowburn et al. This would have been an obvious substitution as Cowburn et al. is also transcribing speech and displaying it using video information only on an augmented reality display rather than on a user’s device (Cowburn et al. Col. 2, Lines 33-40). Regarding Claim 18, Braga teaches the system of claim 16. Furthermore, Cowburn et al. teaches wherein at least one of the one or more microphone devices and the one or more camera devices is coupled to the processor via at least in part a wireless link. (Communication may be implemented using a wide variety of technologies. The I/O components 1518 may include communication components 1540 operable to couple the machine 1500 to a network 1532 or devices 1520 via coupling 1522 and coupling 1524 respectively. For example, the communication components 1540 may include a network interface component or other suitable device to interface with the network 1532. In further examples, communication components 1540 may include wired communication components, wireless communication components, cellular communication components, Near Field Communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components to provide communication via other modalities. The devices 1520 may be another machine or any of a wide variety of peripheral devices (e.g., a peripheral device coupled via a Universal Serial Bus (USB)).) (Col. 22, Line 60 to Col. 23, Line 10). The I/O components such as a microphone can be connected wirelessly. This can be seen in the architecture in Fig. 15. Regarding Claim 19, Braga in view of Cowburn et al. teaches the system of claim 18. Furthermore, Cowburn et al. teaches wherein the wireless link comprises one of a Bluetooth, Wi-Fi, Wi-Fi Direct, Wi-Fi HaLow, Ultra-Wideband, mmWave, 5G, and LiFi connection. (Communication may be implemented using a wide variety of technologies. The I/O components 1518 may include communication components 1540 operable to couple the machine 1500 to a network 1532 or devices 1520 via coupling 1522 and coupling 1524 respectively. For example, the communication components 1540 may include a network interface component or other suitable device to interface with the network 1532. In further examples, communication components 1540 may include wired communication components, wireless communication components, cellular communication components, Near Field Communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components to provide communication via other modalities. The devices 1520 may be another machine or any of a wide variety of peripheral devices (e.g., a peripheral device coupled via a Universal Serial Bus (USB)).) (Col. 22, Line 60 to Col. 23, Line 10). Wireless communication can be done with Bluetooth, Wi-Fi, or a plurality of alternative wireless communication methods. Allowable Subject Matter Claim 13 would be allowable if not for being dependent on the rejected base claims 1 and 12. Furthermore, the claim would need to be rewritten to overcome the rejection(s) under 35 U.S.C. 101 set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The closest prior art of record for claim 13 is US Patent Publication US 11615781 B2 (Braga) in view of US Patent Application Publication US 20190130594 A1 (Seyfi et al.). Braga teaches generating enhanced audio features, generating speaker features, and producing an output. However, it does not teach skipping these steps and still producing an output in response to the speaker not appearing in the video data. Thus, Braga does not teach responsive to detecting absence of a target speaker in the plurality of video streams: skipping the step of generating a plurality of extended video features; skipping the step of generating a plurality of enhanced audio features; and processing the plurality of initial audio features, thereby generating a plurality of target speaker features; and generating an output. Seyfi et al. teaches responsive to detecting absence of a target speaker in the plurality of video streams: (Paragraph 62). However, none of the prior at, either alone or in combination, overcomes the limitations as presented in claims 13. Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS DANIEL LOWEN whose telephone number is (571)272-5828. The examiner can normally be reached Mon-Fri 8:00am - 4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NICHOLAS D LOWEN/Examiner, Art Unit 2653 /DOUGLAS GODBOLD/Primary Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Feb 04, 2025
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705910
SYSTEMS AND METHODS FOR A VISION-LANGUAGE PRETRAINING FRAMEWORK
3y 6m to grant Granted Aug 11, 2026
Patent 12693779
TRAINING AND USING A SENTIMENT MACHINE LEARNING MODULE TO RECEIVE AS INPUT HAPTIC METRIC VALUES TO DETERMINE A SENTIMENT SCORE FOR TEXT TO PROVIDE TO AN INTERACTIVE PROGRAM
3y 11m to grant Granted Jul 28, 2026
Patent 12657381
MULTI-LAYERED CUSTOMIZATION FRAMEWORK
2y 8m to grant Granted Jun 16, 2026
Patent 12614025
Authorship Source Analysis for Large Language Models (LLM) Using a Distributed Ledger
2y 6m to grant Granted Apr 28, 2026
Patent 12592224
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND COMPUTER PROGRAM PRODUCT
2y 1m to grant Granted Mar 31, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
64%
Grant Probability
99%
With Interview (+68.9%)
2y 8m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month