Prosecution Insights
Last updated: October 01, 2026
Application No. 19/068,764

Accelerated Audio Separation and Classification for On-Device Machine-Learned Systems

Non-Final OA §101§103
Filed
Mar 03, 2025
Priority
Mar 01, 2024 — provisional 63/560,491
Examiner
LOWEN, NICHOLAS DANIEL
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
10 granted / 15 resolved
+6.7% vs TC avg
Strong +56% interview lift
Without
With
+55.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
16 currently pending
Career history
38
Total Applications
across all art units

Statute-Specific Performance

§101
34.7%
-5.3% vs TC avg
§103
46.7%
+6.7% vs TC avg
§102
14.6%
-25.4% vs TC avg
§112
3.0%
-37.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 15 resolved cases

Office Action

§101 §103
DETAILED ACTION This communication is in response to the Application filed on 03/03/2025. Claims 1-20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 13, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant claims the benefit of US Provisional Application No. 63/560,491, filed March 01, 2024. Claims 1-20 have been afforded the benefit of this filing date. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 1 and 11 recite A computer-implemented method implemented by one or more [processors], the method comprising: obtaining media including audio and video; providing decoded audio from the media to a [machine-learned audio separation model]; generating a plurality of separated sound components from the decoded audio using the machine-learned audio separation model; providing decoded video from the media and the plurality of separated sound components to a [machine-learned audio classification model]; and generating a class label for each of the plurality of separated sound components using the machine-learned audio classification model. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The human mind can identify sound sources and label them. The method could be performed mentally by creating a physical representation (drawing, numbers, etc.) of the sound. They could then create physical representations for each sound and compare them to images to associate them with the objects in the image. Finally, they could use prior knowledge to label each sound source in the image. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims recite the additional components of processors, a machine-learned audio separation model, and a machine-learned audio classification model. The processors are merely being used to apply the method via a generic computing device. The processor is detailed in paragraph 49 of the specification with a generic description of the component. The machine-learned audio separation model and a machine-learned audio classification model are merely being used to apply the method via a generic computing device. The machine-learned models are detailed in paragraph 115 and are described as general purpose neural network models. Claim 11 specifically lists the additional component of a computer-readable storage media. The computer-readable storage media are merely being used to apply the method via a generic computing device. The computer-readable storage media is detailed in paragraph 49 of the specification with a generic description of the component. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 2 and 12 recite wherein: the media includes a plurality of frames of audio and video; the method further comprises performing keyframe-only decoding of the media to generate the decoded video, the decoded video corresponding to less than all of the plurality of frames of video from the media. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. A human can view a series of images and identify which ones contain key information which is worth analyzing further. The further analyzing could be decoding in the form of noting objects in the image on a piece of paper. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 3 and 13 recite wherein performing keyframe- only decoding comprises: performing seek operations to locate keyframes nearest to required video frames; and decoding the keyframes nearest to the required video frames. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. A human can locate a section of images based on how close they are to key features and a certain end point. For example, selecting the images where an object first appears and including images from that point onward. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 4 and 14 recite generating a graphical user interface including the class label and a user interface element for each of the plurality of separated sound components, wherein the user interface element enables user modification of a corresponding separated sound component. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The human equivalent would be creating a piece of paper with the different sounds listed where someone can fill it out to select which sounds they’d like to change. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 5 and 15 recite wherein: the machine-learned audio separation model is executed by a graphical processing unit; and the machine-learned audio classification model is executed by a tensor processing unit. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The operations these models perform are mental processes, thus, the GPU and TPU are additional components used to apply the mental process. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims recite the additional components of graphical processing unit and tensor processing unit. The graphical processing unit is merely being used to apply the method via a generic computing device. The graphical processing unit is detailed in paragraph 37 of the specification with a generic description of the component. The tensor processing unit is merely being used to apply the method via a generic computing device. The tensor processing unit is detailed in paragraph 37 of the specification with a generic description of the component. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 6 and 16 recite decoding all audio data from the media prior to passing the decoded audio to the machine- learned audio separation model. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The human mind can break down the information found in an audio (writing features down on paper) prior to performing any sound separation tasks. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 7 and 17 recite decoding video from the media in parallel with generating the plurality of separated sound components from the decoded audio using the machine-learned audio separation model. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. A human can identify features in images at the same time that they separate sounds from an audio. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 8 and 18 recite decoding video from the media in parallel with decoding audio from the media. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. A human can analyze images and audio simultaneously and write down features. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claims 9 and 19 recite wherein the media includes a video file. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. A human can analyze information found in a video. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claim 10 recites wherein: each separated sound component corresponds to a distinct source of audio in the media. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. A human can identify sources of audio in images and audio. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims do not recite any additional components that were not cited in the independent claim. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claim 20 recites A computer-implemented method implemented by one or more processors, the method comprising: obtaining media including a plurality of frames of audio and video; providing decoded audio from the media to a machine-learned audio separation model; generating a plurality of separated sound components from the decoded audio using the machine-learned audio separation model; performing keyframe-only decoding of the media to generate decoded video corresponding to less than all of the plurality of frames of video from the media; providing the decoded video and the plurality of separated sound components to a machine-learned audio classification model; generating an audio class label for each of the plurality of separated sound components using the machine-learned audio separation model; and generating a graphical user interface including the audio class label and a user interface element for each of the plurality of separated sound components, wherein the user interface element enables user modification of a corresponding separated sound component. The limitations in these claims, as drafted, are a process that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The human mind can identify sound sources and label them. The method could be performed mentally by creating a physical representation (drawing, number, etc.) of the sound. They could then create physical representations for each sound and compare them to images to associate them with the objects in the image. They could use prior knowledge to label each sound source in the image. A human can also view a series of images and identify which ones contain key information which is worth analyzing further. The further analyzing could be decoding in the form of noting objects in the image on a piece of paper. A section of images can be selected based on how close they are to key features and a certain end point. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application. The claims recite the additional components of processors, a machine-learned audio separation model, and a machine-learned audio classification model. The processors are merely being used to apply the method via a generic computing device. The processor is detailed in paragraph 49 of the specification with a generic description of the component. The machine-learned audio separation model and a machine-learned audio classification model are merely being used to apply the method via a generic computing device. The machine-learned models are detailed in paragraph 115 and are described as general purpose neural network models. Accordingly, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 4, 7-11, 14, and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 20230402055 A1 (Kulasekaran et al.) in view of “Learning to Separate Object Sounds by Watching Unlabeled Video” (Gao et al.). Regarding Claims 1 and 11, Kulasekaran et al. teaches A computer-implemented method implemented by one or more processors, the method comprising: obtaining media including audio and video; (The method includes receiving, by a processor, a real-world sound input and an image, wherein the real-world sound input comprises a plurality of sound signals corresponding to a plurality of sound sources) (Paragraph 15). Claim 11 presents the alternative A system, comprising: one or more processors; and one or more computer-readable storage media that store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising: (Among other capabilities, the processor 302 is adapted to fetch and execute computer-readable instructions and data stored in the memory 304.) (Paragraph 42). providing decoded audio from the media to a machine-learned audio separation model; (The separating module 312 of the neural network 309 may include the encoder 404 to receive the real-world sound input 402 and generate the spectrogram 406. The spectrogram 406 is then subjected to the multi-class (or multi label) classifier 408.) (Paragraph 60). The audio is converted to a spectrogram and separated by a separating module. generating a plurality of separated sound components from the decoded audio using the machine-learned audio separation model; (Further, upon identification of the classes for the sound signal, the heat map 410 may be generated using the Class Activation Mapping (CAM) technique. The separation module 312 uses the heat map 410 to separate the sound signal 412. A masking loss 414 may also be computed based on the separated sound signal 412.) (Paragraph 61). (The method 600 may include splitting the spectrogram into ‘n’ number of horizontal patches based on a plurality of predefined classes. The neural network is trainable for splitting the spectrogram based on the predefined classes.) (Paragraph 74). The separation module separates the sounds. providing (decoded video from the media) (Taught by Gao et al.) and the plurality of separated sound components to a machine-learned audio classification model; (At block 502, the process flow 500 may include splitting the spectrogram 406 into the horizontal patches. For instance, the spectrogram 406 may be split into ‘n’ number of patches equivalent to the number pre-defined classes in the multi label classifier 408.) (Paragraph 65). (At operation 608, the method 600 may include generating an association between each of the sound generating object 108 and the separated sound signals 412. In an example, the association is generated based on contrastive learning and wherein the contrastive learning is based on the permutation invariant contrastive learning.) (Paragraph 77). Kulasekaren et al. teaches a classifier that labels each sound and then in the following step the labelled sounds are associated with sound emitting objects in the images (Fig. 4). The exact architecture of both separated sounds and video being input to the classifier is taught by Gao et al. below. generating a class label for each of the plurality of separated sound components using the machine-learned audio classification model. (At operation 610, the method 600 may include matching in real-time, each of the detected sound generating object 108 with the respective sound signals from the separated sound signals 412 based on the association generated.) (Paragraph 78). Labelled sounds are matched with respective objects in the image data. Kulasekaren et al. does not explicitly teach: providing decoded video from the media and the plurality of separated sound components to a machine-learned audio classification model; However, Gao et al. teaches providing decoded video from the media and the plurality of separated sound components to a machine-learned audio classification model; (To recover the association, we construct a neural network for multi-instance multi-label learning (MIML) that maps audio bases to the distribution of detected visual objects.) (Section 1, Pargraph 4). (Unsupervised training pipeline. For each video, we perform NMF on its audio magnitude spectrogram to get M basis vectors. An ImageNet-trained ResNet-152 network is used to make visual predictions to find the potential objects present in the video. Finally, we perform multi-instance multi-label learning to disentangle which extracted audio basis vectors go with which detected visible object(s).) (Fig. 2 Description). Fig. 2 shows the architecture of separating the audio and providing both the separated audio and the decoded image data to multi-label learning model (classifier). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the sound source classification method as taught by Kulasekaren et al. to include use a classifier based on both audio and video as taught by Gao et al. This would have been an obvious improvement the model can then associate sounds with image data rather than processing the audio by itself (Gao et al. Section 1, Paragraph 5). Regarding Claims 4 and 14, Kulasekaren et al. in view of Gao et al. teaches the system of claims 1 and 11. Furthermore, Kulasekaren et al. teaches generating a graphical user interface including the class label and a user interface element for each of the plurality of separated sound components, wherein the user interface element enables user modification of a corresponding separated sound component. (The visual source 106 is illustrated. The sound generating object 108 are marked by the bounding box using the object detection bounding box 416. The user 102 may control the sliding knobs 702 to increase or decrease the magnitude of the each of the respective sound signals corresponding to the sound generating object 108 in the visual source 106.) (Paragraph 80). Fig. 7 shows a user interface where individual sound sources are represented with a slider to adjust them. Regarding Claims 7 and 17, Kulasekaren et al. in view of Gao et al. teaches the system of claims 1 and 11. Furthermore, Kulasekaren et al. teaches decoding video from the media in parallel with generating the plurality of separated sound components from the decoded audio using the machine-learned audio separation model. (The separating module 312 of the neural network 309 may include the encoder 404 to receive the real-world sound input 402 and generate the spectrogram 406. The spectrogram 406 is then subjected to the multi-class (or multi label) classifier 408.) (Paragraph 60). (In another embodiment, the detection module 314 may be adapted to receive the visual source 106. The detection module 314 may be adapted to apply the object detection bounding box 416 technique to mark a box area around each of the sound generating object 108 in the visual source 106.) (Paragraph 62). Fig. 4 shows the audio separation and video decoding happening in parallel withing the system. Regarding Claims 8 and 18, Kulasekaren et al. in view of Gao et al. teaches the system of claims 1 and 11. Furthermore, Kulasekaren et al. teaches decoding video from the media in parallel with decoding audio from the media. Fig. 4, shows the audio and video data being decoded in parallel to be separated/classified. Regarding Claims 9 and 19, Kulasekaren et al. in view of Gao et al. teaches the system of claims 1 and 11. Furthermore, Kulasekaren et al. teaches wherein the media includes a video file. (In an example, the visual source 106 may include a camera preview, a video, an image, exhibiting a sound generating object 108.) (Paragraph 38). The system takes in video data as input. Regarding Claim 10, Kulasekaren et al. in view of Gao et al. teaches the system of claims 1. Furthermore, Kulasekaren et al. teaches wherein: each separated sound component corresponds to a distinct source of audio in the media. (At operation 604, the method 600 may include separating the sound signals from the real-world sound input. In an example, for separating the sound signals, the method 600 may include generating the spectrogram 406 of the real-world sound input. The method 600 may include identifying the class for each of the sound signals in the spectrogram 406. In an example, the class is indicative of a type of the sound generating object 108. The method 600 may include splitting the spectrogram into ‘n’ number of horizontal patches based on a plurality of predefined classes) (Paragraph 74). The audio is split up by each sound source. Claims 2, 3, 12, 13, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 20230402055 A1 (Kulasekaran et al.) in view of “Learning to Separate Object Sounds by Watching Unlabeled Video” (Gao et al.) and further in view of US Patent Publication US 20220078514 A1 (Bustamante et al.). Regarding Claims 2 and 12, Kulasekaran et al. in view of Gao et al. teaches the method of claims 1 and 11. Furthermore, Kulasekaren et al. teaches wherein: the media includes a plurality of frames of audio and video; (In an example, the visual source 106 may include a camera preview, a video, an image, exhibiting a sound generating object 108.) (Paragraph 38). The system takes in video data as input which includes a plurality of frames of audio and video. Kulasekaren et al. in view of Gao et al. does not explicitly teach: the method further comprises performing keyframe-only decoding of the media to generate the decoded video, the decoded video corresponding to less than all of the plurality of frames of video from the media. However, Gao et al. teaches the method further comprises performing keyframe-only decoding of the media to generate the decoded video, the decoded video corresponding to less than all of the plurality of frames of video from the media. (A real-time media stream can include at least one key frame (e.g., data encoded for rendering a complete frame by a client device) and a number of predictive or delta frames that represent differences relative to the key frame.) (Paragraph 29). (Key frames can be decoded without reference to any other frame in a sequence; that is, the decoder reconstructs such frames beginning from a default state. Key frames provide random access (or seeking) points in a media stream. Prediction frames are encoded with a reference to prior frames, specifically all prior frames up to and including the most recent key frame.) (Paragraph 30). Bustamante et al. teaches a method of modifying video data which performs key frame decoding to identify which frames are most important to perform processing on. It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the sound source classification method as taught by Kulasekaren et al. in view of Gao et al. to include keyframe decoding as taught by Bustamante et al. This would have been an obvious improvement to avoid processing frames that are not near keyframes such as corrupt frames (Bustamante et al. Paragraph 30). Regarding Claims 3 and 13, Kulasekaren et al. in view of Gao et al. teaches the system of claims 2 and 12. Furthermore, Bustamante et al. teaches wherein performing keyframe- only decoding comprises: performing seek operations to locate keyframes nearest to required video frames; and decoding the keyframes nearest to the required video frames. (A real-time media stream can include at least one key frame (e.g., data encoded for rendering a complete frame by a client device) and a number of predictive or delta frames that represent differences relative to the key frame.) (Paragraph 29) (Key frames can be decoded without reference to any other frame in a sequence; that is, the decoder reconstructs such frames beginning from a default state. Key frames provide random access (or seeking) points in a media stream. Prediction frames are encoded with a reference to prior frames, specifically all prior frames up to and including the most recent key frame.) (Paragraph 30). Keyframes represent seeking points within the video and a range of frames around the keyframe is processed by the system Regarding Claim 20, Kulasekaran et al. teaches A computer-implemented method implemented by one or more processors, the method comprising: (The method includes receiving, by a processor, a real-world sound input and an image, wherein the real-world sound input comprises a plurality of sound signals corresponding to a plurality of sound sources) (Paragraph 15). providing decoded audio from the media to a machine-learned audio separation model; (The separating module 312 of the neural network 309 may include the encoder 404 to receive the real-world sound input 402 and generate the spectrogram 406. The spectrogram 406 is then subjected to the multi-class (or multi label) classifier 408.) (Paragraph 60). The audio is converted to a spectrogram and separated by a separating module. generating a plurality of separated sound components from the decoded audio using the machine-learned audio separation model; (Further, upon identification of the classes for the sound signal, the heat map 410 may be generated using the Class Activation Mapping (CAM) technique. The separation module 312 uses the heat map 410 to separate the sound signal 412. A masking loss 414 may also be computed based on the separated sound signal 412.) (Paragraph 61). (The method 600 may include splitting the spectrogram into ‘n’ number of horizontal patches based on a plurality of predefined classes. The neural network is trainable for splitting the spectrogram based on the predefined classes.) (Paragraph 74). The separation module separates the sounds. providing (decoded video from the media) (Taught by Gao et al.) and the plurality of separated sound components to a machine-learned audio classification model; (At block 502, the process flow 500 may include splitting the spectrogram 406 into the horizontal patches. For instance, the spectrogram 406 may be split into ‘n’ number of patches equivalent to the number pre-defined classes in the multi label classifier 408.) (Paragraph 65). (At operation 608, the method 600 may include generating an association between each of the sound generating object 108 and the separated sound signals 412. In an example, the association is generated based on contrastive learning and wherein the contrastive learning is based on the permutation invariant contrastive learning.) (Paragraph 77). Kulasekaren et al. teaches a classifier that labels each sound and then in the following step the labelled sounds are associated with sound emitting objects in the images (Fig. 4). The exact architecture of both separated sounds and video being input to the classifier is taught by Gao et al. below. generating an audio class label for each of the plurality of separated sound components using the machine-learned audio separation model. (At operation 610, the method 600 may include matching in real-time, each of the detected sound generating object 108 with the respective sound signals from the separated sound signals 412 based on the association generated.) (Paragraph 78). Labelled sounds are matched with respective objects in the image data. generating a graphical user interface including the audio class label and a user interface element for each of the plurality of separated sound components, wherein the user interface element enables user modification of a corresponding separated sound component. (The visual source 106 is illustrated. The sound generating object 108 are marked by the bounding box using the object detection bounding box 416. The user 102 may control the sliding knobs 702 to increase or decrease the magnitude of the each of the respective sound signals corresponding to the sound generating object 108 in the visual source 106.) (Paragraph 80). Fig. 7 shows a user interface where individual sound sources are represented with a slider to adjust them. Kulasekaren et al. does not explicitly teach: performing keyframe-only decoding of the media to generate the decoded video corresponding to less than all of the plurality of frames of video from the media. providing decoded video from the media and the plurality of separated sound components to a machine-learned audio classification model; However, Gao et al. teaches providing decoded video from the media and the plurality of separated sound components to a machine-learned audio classification model; (To recover the association, we construct a neural network for multi-instance multi-label learning (MIML) that maps audio bases to the distribution of detected visual objects.) (Section 1, Pargraph 4). (Unsupervised training pipeline. For each video, we perform NMF on its audio magnitude spectrogram to get M basis vectors. An ImageNet-trained ResNet-152 network is used to make visual predictions to find the potential objects present in the video. Finally, we perform multi-instance multi-label learning to disentangle which extracted audio basis vectors go with which detected visible object(s).) (Fig. 2 Description). Fig. 2 shows the architecture of separating the audio and providing both the separated audio and the decoded image data to multi-label learning model (classifier). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the sound source classification method as taught by Kulasekaren et al. to include use a classifier based on both audio and video as taught by Gao et al. This would have been an obvious improvement the model can then associate sounds with image data rather than processing the audio by itself (Gao et al. Section 1, Paragraph 5). Kulasekaren et al. in view of Gao et al. does not explicitly teach: performing keyframe-only decoding of the media to generate the decoded video corresponding to less than all of the plurality of frames of video from the media. However, Gao et al. teaches performing keyframe-only decoding of the media to generate the decoded video corresponding to less than all of the plurality of frames of video from the media. (A real-time media stream can include at least one key frame (e.g., data encoded for rendering a complete frame by a client device) and a number of predictive or delta frames that represent differences relative to the key frame.) (Paragraph 29). (Key frames can be decoded without reference to any other frame in a sequence; that is, the decoder reconstructs such frames beginning from a default state. Key frames provide random access (or seeking) points in a media stream. Prediction frames are encoded with a reference to prior frames, specifically all prior frames up to and including the most recent key frame.) (Paragraph 30). Bustamante et al. teaches a method of modifying video data which performs key frame decoding to identify which frames are most important to perform processing on. It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the sound source classification method as taught by Kulasekaren et al. in view of Gao et al. to include keyframe decoding as taught by Bustamante et al. This would have been an obvious improvement to avoid processing frames that are not near keyframes such as corrupt frames (Bustamante et al. Paragraph 30). Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 20230402055 A1 (Kulasekaran et al.) in view of “Learning to Separate Object Sounds by Watching Unlabeled Video” (Gao et al.) and further in view of US Patent Publication US 20180122403 A1 (Koretzky et al.). Regarding Claims 6 and 16, Kulasekaran et al. in view of Gao et al. teaches the method of claims 1 and 11. Kulasekaren et al. in view of Gao et al. does not explicitly teach: decoding all audio data from the media prior to passing the decoded audio to the machine- learned audio separation model. However, Koretzky et al. teaches decoding all audio data from the media prior to passing the decoded audio to the machine- learned audio separation model. (Server 102 may optionally include reading/decoding logic 110. Reading/decoding logic 110 is programmed or configured to read an audio source 105 and generate a plurality of raw audio samples based on the audio source 105.) (Paragraph 60). (In an embodiment, the spectrogram generated by transform logic 114 may be sent to audio source separation logic 116 for further processing.) (Paragraph 77) Koretzky et al. teaches an audio source separation method in which the audio is decoded prior to being given to an audio separation model. It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the sound source classification method as taught by Kulasekaren et al. in view of Gao et al. to decode the audio prior to separating it as taught by Koretzky et al. This would have been an obvious improvement reduce the size of the raw audio so the separating model has less to compute (Koretzky et al. Paragraph 67). Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication US 20230402055 A1 (Kulasekaran et al.) in view of “Learning to Separate Object Sounds by Watching Unlabeled Video” (Gao et al.) and further in view of US Patent Publication US 20180122403 A1 (Koretzky et al.) and “Environmental Sound Recognition on Embedded Systems: From FPGAs to TPUs” (Vandendriessche et al.). Regarding Claims 5 and 15, Kulasekaran et al. in view of Gao et al. teaches the method of claims 1 and 11. Kulasekaren et al. in view of Gao et al. does not explicitly teach: wherein: the machine-learned audio separation model is executed by a graphical processing unit; and the machine-learned audio classification model is executed by a tensor processing unit. However, Koretzky et al. teaches wherein: the machine-learned audio separation model is executed by a graphical processing unit; (In an embodiment, the previously-described DNN inference of audio source separation logic 116 may be performed, at least in part, by one or more Graphics Processing Units (GPUs).) (Paragraph 106). Koretzky et al. teaches an audio source separation method that is executed using a GPU. It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the sound source classification method as taught by Kulasekaren et al. in view of Gao et al. to execute the audio separation using a GPU as taught by Koretzky et al. This would have been an obvious addition as this is a common processing unit found in general purpose computing devices (Koretzky et al. Paragraph 164). Kulasekaren et al. in view of Gao et al. and Koretzky et al. does not explicitly teach: the machine-learned audio classification model is executed by a tensor processing unit. However, Vandendriessche et al. teaches the machine-learned audio classification model is executed by a tensor processing unit. (In this section, an overview is given of the methodology used to implement embedded sound classifiers and for their evaluation. First, an overview of the well-known audio datasets is given. Second, an overview of different audio features and their extraction to be used for the classifiers’ training are explained. Finally, several metrics to evaluate the sound classifiers are discussed.) (Section 3, Paragraph 1). (It is important to note that the TPU compiler currently has an important limitation in that the TPU does not support all operations, and thus some operations are executed on the CPU[45].) (Section 4.3.1, Paragraph 2) Vandrendriessche et al. teaches common methods and hardware for sound classification in which a TPU is detailed for performing this task. It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the sound source classification method as taught by Kulasekaren et al. in view of Gao et al. and Koretzky et al. to use a TPU for audio classification as taught by Vandrendriessche et al. This would have been an obvious improvement as TPU’s are know to be used for this task and a limitation of them is that they have to offload tasks to other processing units (GPU in this case) (Vandrendriessche et al. Section 4.3.1, Paragraph 2). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS DANIEL LOWEN whose telephone number is (571)272-5828. The examiner can normally be reached Mon-Fri 8:00am - 4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NICHOLAS D LOWEN/Examiner, Art Unit 2653 /Paras D Shah/Supervisory Patent Examiner, Art Unit 2653 09/04/2026
Read full office action

Prosecution Timeline

Mar 03, 2025
Application Filed
Sep 09, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748926
MACHINE LEARNING BASED SYSTEMS AND METHODS FOR ANALYZING INTENT OF EMAILS
2y 9m to grant Granted Sep 29, 2026
Patent 12705910
SYSTEMS AND METHODS FOR A VISION-LANGUAGE PRETRAINING FRAMEWORK
3y 6m to grant Granted Aug 11, 2026
Patent 12693779
TRAINING AND USING A SENTIMENT MACHINE LEARNING MODULE TO RECEIVE AS INPUT HAPTIC METRIC VALUES TO DETERMINE A SENTIMENT SCORE FOR TEXT TO PROVIDE TO AN INTERACTIVE PROGRAM
3y 11m to grant Granted Jul 28, 2026
Patent 12657381
MULTI-LAYERED CUSTOMIZATION FRAMEWORK
2y 8m to grant Granted Jun 16, 2026
Patent 12614025
Authorship Source Analysis for Large Language Models (LLM) Using a Distributed Ledger
2y 6m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+55.6%)
2y 7m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 15 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month