DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Introduction
This office action is in response to communications filed 11/05/2024. Claims 1-20 are pending and likewise have been examined.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 11/05/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 10 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 10 recites the limitation "the at least one transformation operation" in line 4. A “at least one transformation operation” has been established twice prior to this point in the claim. One on line 2-3 of claim 10 and another on line 10 of claim 1. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5, 11-12, 15 and 17-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omran et al. “Disentangling speech from surroundings with neural embeddings”, and further in view of Lim et al. (US 20210366497 A1).
Regarding Claim 1:
Omran teaches a method comprising: receiving an audio signal at a transmitting endpoint(Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame);
using a trained neural network, source separating the audio signal to generate a first source stream comprising first embedding vectors representative of a first source of the audio signal and a second source stream comprising second embedding vectors representative of a second source of the audio signal(Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame. Pg 2, 3.1. Model, Para 1, Ln 1-5, Our model architecture is shown in Fig. 1A. Single-channel audio waveforms x ∈ RT are processed by the encoder, which is a convolutional network with blocks consisting of residual units and a strided convolutional layer. All convolutions have causal padding to allow for real-time streaming of the input waveforms. Pg 2, 3.2. Disentanglement scheme, Para 1, Ln 1-4, Disentangling speech from its environment requires careful de sign of the training objectives, in order to both achieve high quality audio transmission and create structured embeddings with no content leakage between different partitions);
applying at least one transformation operation to at least one of the first source stream and the second source stream(Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12, We also find that multiplying a partition with a weight factor between 0 and 1 leads to a reduction in volume of the corresponding audio content, a feature that can be used for adjusting the target level of denoising);
vector quantizing the first source stream to generate first codewords(Pg 2, 3.1. Model, Para 2, Ln 1-5, We apply residual vector quantization to each embedding partition separately. A residual vector quantizer [4] is a stack of vector quantization layers, each of which replaces its input by a vector from a learned discrete codebook and passes the residual error to the next layer.);
vector quantizing the second source stream to generate second codewords(Pg 2, 3.1. Model, Para 2, Ln 1-5, We apply residual vector quantization to each embedding partition separately. A residual vector quantizer [4] is a stack of vector quantization layers, each of which replaces its input by a vector from a learned discrete codebook and passes the residual error to the next layer.);
and transmitting, in one or more packets, the first codewords and the second codewords, or respective indexes thereof, to a remote endpoint(Pg 2, 3.1. Model, Para 3, Ln 1-3, The quantized embeddings are concatenated together along the feature axis and fed into the decoder, which converts them into a reconstructed audio waveform).
Omran does not teach adjusting an allocation of bandwidth for at least one of the first source stream and the second source stream based on at least one of entropy and dynamism of the at least one of the first source stream and the second source stream; applying at least one transformation operation to at least one of the first source stream and the second source stream.
In the same field of source separation and audio coding, Lim teaches adjusting an allocation of bandwidth for at least one of the first source stream and the second source stream based on at least one of entropy and dynamism of the at least one of the first source stream and the second source stream(Para [0084], Ln 1-5, A theoretical lower bound of a bitrate caused by Huffman coding may be defined by entropies of the quantized sound source signals 205 and 206. A frequency of a cluster mean may define the entropies of the sound source signals 205 and 206 as shown in Equation 14 below. Para [0090], Ln 1-7, may denote the total entropy of the input signal 201 of the k-th sound source, and H(μ.sup.(k)) may denote entropies of the sound source signals 203 and 204 of the k-th sound source. As an experimental result, a loss value for entropy may not guarantee an exact bitrate, but may not have a great influence on performance despite the difference between the actual entropy and the target entropy. Para [0087], Ln 1-6, In an example, by performing quantization with different bitrates for each sound source, a sound source signal of an important sound source recognized by humans may be quantized at a higher bitrate than a sound source signal of other sound sources such as noise, so that a reconstruction quality may increase).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify Omran with the entropy based bitrate of Lim, as it can improve audio quality(Para [0087], Ln 1-6).
Regarding Claim 2:
The combination of Omran and Lim teaches the method of claim 1, but does not teach further comprising: using the trained neural network, source separating the audio signal to generate a third source stream comprising third embedding vectors representative of a third source of the audio signal; and at least one of selecting the third source stream to not send to the remote endpoint and adjusting an allocation of bandwidth for the third source stream.
In the same field of source separation and audio coding, Lim teaches further comprising: using the trained neural network, source separating the audio signal to generate a third source stream comprising third embedding vectors representative of a third source of the audio signal(Para [0064], Ln 1-9, each of the sound source signals 203 and 204 for the plurality of sound sources may satisfy Equations 5 and 6 shown below. Eq 5 & 6. Para [0065], Ln 1-5, In Equation 6, z.sup.(k) may denote the sound source signals 203 and 204 for the k-th sound source. Also, z may denote the latent signal 202, and m.sup.(k) may denote a masking vector for the k-th sound source.);
and at least one of selecting the third source stream to not send to the remote endpoint and adjusting an allocation of bandwidth for the third source stream(Para [0084], Ln 1-5, A theoretical lower bound of a bitrate caused by Huffman coding may be defined by entropies of the quantized sound source signals 205 and 206. A frequency of a cluster mean may define the entropies of the sound source signals 205 and 206 as shown in Equation 14 below. Para [0090], Ln 1-7, may denote the total entropy of the input signal 201 of the k-th sound source, and H(μ.sup.(k)) may denote entropies of the sound source signals 203 and 204 of the k-th sound source. As an experimental result, a loss value for entropy may not guarantee an exact bitrate, but may not have a great influence on performance despite the difference between the actual entropy and the target entropy.).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Omran and Lim with the entropy based coding of Lim, as it can improve audio quality(Para [0087], Ln 1-6).
Regarding Claim 5:
The combination of Omran and Lim teaches the method of claim 1, and Omran teaches wherein applying the at least one transformation operation to the at least one of the first source stream and the second source stream comprises controlling at least one of volume, tone, and temporal and spectral characteristics for at least one of the first source stream and the second source stream(Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12, We also find that multiplying a partition with a weight factor between 0 and 1 leads to a reduction in volume of the corresponding audio content, a feature that can be used for adjusting the target level of denoising).
Regarding Claim 11:
Omran teaches receive an audio signal(Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame);
using a trained neural network, source separate the audio signal to generate a first source stream comprising first embedding vectors representative of a first source of the audio signal and a second source stream comprising second embedding vectors representative of a second source of the audio signal(Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame. Pg 2, 3.1. Model, Para 1, Ln 1-5, Our model architecture is shown in Fig. 1A. Single-channel audio waveforms x ∈ RT are processed by the encoder, which is a convolutional network with blocks consisting of residual units and a strided convolutional layer. All convolutions have causal padding to allow for real-time streaming of the input waveforms. Pg 2, 3.2. Disentanglement scheme, Para 1, Ln 1-4, Disentangling speech from its environment requires careful de sign of the training objectives, in order to both achieve high quality audio transmission and create structured embeddings with no content leakage between different partitions);
apply at least one transformation operation to at least one of the first source stream and the second source stream(Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12, We also find that multiplying a partition with a weight factor between 0 and 1 leads to a reduction in volume of the corresponding audio content, a feature that can be used for adjusting the target level of denoising);
vector quantize the first source stream to generate first codewords(Pg 2, 3.1. Model, Para 2, Ln 1-5, We apply residual vector quantization to each embedding partition separately. A residual vector quantizer [4] is a stack of vector quantization layers, each of which replaces its input by a vector from a learned discrete codebook and passes the residual error to the next layer.);
vector quantize the second source stream to generate second codewords(Pg 2, 3.1. Model, Para 2, Ln 1-5, We apply residual vector quantization to each embedding partition separately. A residual vector quantizer [4] is a stack of vector quantization layers, each of which replaces its input by a vector from a learned discrete codebook and passes the residual error to the next layer.);
and transmit, in one or more packets, the first codewords and the second codewords, or respective indexes thereof, to a remote endpoint(Pg 2, 3.1. Model, Para 3, Ln 1-3, The quantized embeddings are concatenated together along the feature axis and fed into the decoder, which converts them into a reconstructed audio waveform).
Omran does not explicitly teach a device comprising: an interface configured to enable network communications; a memory; and one or more processors coupled to the interface and the memory, and configured to: ….adjust an allocation of bandwidth for at least one of the first source stream and the second source stream based on at least one of entropy and dynamism of the at least one of the first source stream and the second source stream.
In the same field of source separation and audio coding, Lim teaches a device comprising: an interface configured to enable network communications; a memory; and one or more processors coupled to the interface and the memory, and configured to(Para [0144], Ln 1-19, information carrier, e.g., in a machine-readable storage device…..programmable processor….interconnected by a communication network):
adjust an allocation of bandwidth for at least one of the first source stream and the second source stream based on at least one of entropy and dynamism of the at least one of the first source stream and the second source stream(Para [0084], Ln 1-5, A theoretical lower bound of a bitrate caused by Huffman coding may be defined by entropies of the quantized sound source signals 205 and 206. A frequency of a cluster mean may define the entropies of the sound source signals 205 and 206 as shown in Equation 14 below. Para [0090], Ln 1-7, may denote the total entropy of the input signal 201 of the k-th sound source, and H(μ.sup.(k)) may denote entropies of the sound source signals 203 and 204 of the k-th sound source. As an experimental result, a loss value for entropy may not guarantee an exact bitrate, but may not have a great influence on performance despite the difference between the actual entropy and the target entropy. Para [0087], Ln 1-6, In an example, by performing quantization with different bitrates for each sound source, a sound source signal of an important sound source recognized by humans may be quantized at a higher bitrate than a sound source signal of other sound sources such as noise, so that a reconstruction quality may increase).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify Omran with the computer components and entropy based bitrate of Lim, as it provides a environment for the system to be realized(Para [0144], Ln 1-19) and can improve audio quality(Para [0087], Ln 1-6).
Regarding Claim 12:
Claim 12 contains similar limitations as Claim 2, and is therefore rejected for the same reasons.
Regarding Claim 15:
Claim 15 contains similar limitations as Claim 5, and is therefore rejected for the same reasons.
Regarding Claim 17:
Omran teaches receive an audio signal at a transmitting endpoint(Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame);
using a trained neural network, source separate the audio signal to generate a first source stream comprising first embedding vectors representative of a first source of the audio signal and a second source stream comprising second embedding vectors representative of a second source of the audio signal(Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame. Pg 2, 3.1. Model, Para 1, Ln 1-5, Our model architecture is shown in Fig. 1A. Single-channel audio waveforms x ∈ RT are processed by the encoder, which is a convolutional network with blocks consisting of residual units and a strided convolutional layer. All convolutions have causal padding to allow for real-time streaming of the input waveforms. Pg 2, 3.2. Disentanglement scheme, Para 1, Ln 1-4, Disentangling speech from its environment requires careful de sign of the training objectives, in order to both achieve high quality audio transmission and create structured embeddings with no content leakage between different partitions);
apply at least one transformation operation to at least one of the first source stream and the second source stream(Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12, We also find that multiplying a partition with a weight factor between 0 and 1 leads to a reduction in volume of the corresponding audio content, a feature that can be used for adjusting the target level of denoising);
vector quantize the first source stream to generate first codewords(Pg 2, 3.1. Model, Para 2, Ln 1-5, We apply residual vector quantization to each embedding partition separately. A residual vector quantizer [4] is a stack of vector quantization layers, each of which replaces its input by a vector from a learned discrete codebook and passes the residual error to the next layer.);
vector quantize the second source stream to generate second codewords(Pg 2, 3.1. Model, Para 2, Ln 1-5, We apply residual vector quantization to each embedding partition separately. A residual vector quantizer [4] is a stack of vector quantization layers, each of which replaces its input by a vector from a learned discrete codebook and passes the residual error to the next layer.);
and transmit, in one or more packets, the first codewords and the second codewords, or respective indexes thereof, to a remote endpoint(Pg 2, 3.1. Model, Para 3, Ln 1-3, The quantized embeddings are concatenated together along the feature axis and fed into the decoder, which converts them into a reconstructed audio waveform).
Omran does not explicitly teach one or more non-transitory computer readable storage media encoded with instructions that, when executed by a processor, cause the processor to: ….adjust an allocation of bandwidth for at least one of the first source stream and the second source stream based on at least one of entropy and dynamism of the at least one of the first source stream and the second source stream.
In the same field of source separation and audio coding, Lim teaches one or more non-transitory computer readable storage media encoded with instructions that, when executed by a processor, cause the processor to: (Para [0144], Ln 1-19, information carrier, e.g., in a machine-readable storage device…..programmable processor….interconnected by a communication network):
adjust an allocation of bandwidth for at least one of the first source stream and the second source stream based on at least one of entropy and dynamism of the at least one of the first source stream and the second source stream(Para [0084], Ln 1-5, A theoretical lower bound of a bitrate caused by Huffman coding may be defined by entropies of the quantized sound source signals 205 and 206. A frequency of a cluster mean may define the entropies of the sound source signals 205 and 206 as shown in Equation 14 below. Para [0090], Ln 1-7, may denote the total entropy of the input signal 201 of the k-th sound source, and H(μ.sup.(k)) may denote entropies of the sound source signals 203 and 204 of the k-th sound source. As an experimental result, a loss value for entropy may not guarantee an exact bitrate, but may not have a great influence on performance despite the difference between the actual entropy and the target entropy. Para [0087], Ln 1-6, In an example, by performing quantization with different bitrates for each sound source, a sound source signal of an important sound source recognized by humans may be quantized at a higher bitrate than a sound source signal of other sound sources such as noise, so that a reconstruction quality may increase).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify Omran with the computer components and entropy based bitrate of Lim, as it provides an environment for the system to be realized(Para [0144], Ln 1-19) and can improve audio quality(Para [0087], Ln 1-6).
Regarding Claim 18:
Claim 18 contains similar limitations as Claim 2 and is therefore rejected for the same reasons.
Claim(s) 3, 13 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Omran and Lim as applied to claim 2 above, and further in view of Atti et al. (US 20190013028 A1).
Regarding Claim 3:
The combination of Omran and Lim teaches the method of claim 2, but does not teach further comprising allocating an increased amount of bandwidth to at least one of the first source stream and the second source stream in response to at least one of the third source stream not being sent to the remote endpoint and the allocation of bandwidth for the third source stream being reduced.
In the same filed of audio coding, Atti teaches further comprising allocating an increased amount of bandwidth to at least one of the first source stream and the second source stream in response to at least one of the third source stream not being sent to the remote endpoint and the allocation of bandwidth for the third source stream being reduced(Para [0065], Ln 1-20, streams having priority 5 may be assigned a highest estimated bit rate, streams having priority 4 may be assigned a next-highest estimated bit rate, and streams having priority 1 may be assigned a lowest estimated bit rate. Estimated bit rate may be determined at least partially based on a total bitrate available for the output bitstream 126, such as by partitioning the total bitrate into larger-sized bit allocations for higher-priority streams and smaller-sized bit allocations for lower-priority streams. Para [0030], Ln 1-13, determine a priority of each of the streams based on one or more characteristics of the signal corresponding to the stream, such as signal energy, foreground vs. background, content type, or entropy).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Omran and Lim with the bitrate control of Atti, as it can help improve the quality of desired streams(Para [0060], Ln 1-23).
Regarding Claim 13:
Claim 13 contains similar limitations as Claim 3 and is therefore rejected for the same reasons.
Regarding Claim 19:
Claim 19 contains similar limitations as Claim 3 and is therefore rejected for the same reasons.
Claim(s) 4, 14 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Omran and Lim as applied to claim 1 above, and further in view of Kim et al. (US 20240127835 A1).
Regarding Claim 4:
The combination of Omran and Lim teaches the method of claim 1, but does not teach further comprising adjusting an amount of bandwidth allocated to at least one of the first source stream and the second source stream in response to receiving information regarding a channel condition between the transmitting endpoint and the remote endpoint.
In the same field of audio coding, Kim teaches further comprising adjusting an amount of bandwidth allocated to at least one of the first source stream and the second source stream in response to receiving information regarding a channel condition between the transmitting endpoint and the remote endpoint(Para [0020], Ln 1-13, For example, the external electronic device 102 may obtain the audio bitstream based on executing the coding based on a bitrate corresponding to a quality (or state) of the channel 110. For example, on a condition that a value indicating a quality of the channel 110 is a first value, the external electronic device 102 may obtain the audio bitstream by executing the coding based on a first bitrate corresponding to the first value).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Omran and Lim, with the bitrate adjustments of Kim, as it can improve audio quality and coding efficiency(Para [0020], Ln 1-22).
Regarding Claim 14:
Claim 14 contains similar limitations as Claim 4 and is therefore rejected for the same reasons.
Regarding Claim 20:
Claim 20 contains similar limitations as Claim 4 and is therefore rejected for the same reasons.
Claim(s) 6, 9-10 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Omran and Lim as applied to claim 5 above, and further in view of Kukosa (US 20190268552 A1).
Regarding Claim 6:
The combination of Omran and Lim teaches the method of claim 5, and Omran teaches wherein applying the at least one transformation operation to the at least one of the first source stream and the second source stream is in response to input(Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12, We also find that multiplying a partition with a weight factor between 0 and 1 leads to a reduction in volume of the corresponding audio content, a feature that can be used for adjusting the target level of denoising).
The combination of Omran and Lim does not explicitly teach in response to input received via a user interface.
In the same field of audio processing, Kukosa teaches applying the at least one transformation operation to the at least one of the first source stream and the second source stream is in response to input received via a user interface(Para [0036], Ln 1-8, control information is defined by the central node. Alternatively, the control information may be fixedly predefined. Definition may be done by algorithm and/or user input. Para [0038], Ln 1-8, The device of this aspect is adapted to execute a method according to any of the preceding aspects, and may in particular be a server, preferably conference server, client, desktop computer, portable computer, tablet, telephone, mobile phone, smart phone, PDA, or the like. Para [0037], Ln 1-5, The characteristic parameter may be a volume parameter, and/or may be adapted to said control information).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Omran and Lim, with the stream control method of Kukosa, as it can improve scalability(Para [0010], Ln 1-7).
Regarding Claim 9:
The combination of Omran and Lim teaches the method of claim 1, but does not teach further comprising: in response to receiving user input at the remote endpoint, selecting at least one of a reproduced first source stream and a reproduced second source stream to not play back as audio.
In the same field of audio processing, Kukosa teaches further comprising: in response to receiving user input at the remote endpoint, selecting at least one of a reproduced first source stream and a reproduced second source stream to not play back as audio(Para [0036], Ln 1-8, control information is defined by the central node. Alternatively, the control information may be fixedly predefined. Definition may be done by algorithm and/or user input. Para [0038], Ln 1-8, The device of this aspect is adapted to execute a method according to any of the preceding aspects, and may in particular be a server, preferably conference server, client, desktop computer, portable computer, tablet, telephone, mobile phone, smart phone, PDA, or the like. Para [0037], Ln 1-5, The characteristic parameter may be a volume parameter, and/or may be adapted to said control information. Abstract, Ln 1-22, receiving audio streams from audio clients connected to the distributed multipoint audio processing node and generating evaluated audio streams by analyzing packets of the received audio streams in terms of at least one audio communication characteristic, and attaching an analysis result information of said analysis to said packets, in each audio stream. Audio streams are selected by deciding on whether or not any evaluated audio stream is to be transmitted further, based on the received control information and/or the analysis result information contained in said evaluated audio streams. Then selected audio streams are transmitted further while discarding evaluated audio streams decided not to be to be transmitted further, without mixing any transmitted audio streams. Para [0017], Ln 1-5, preselecting audio streams by deciding on whether or not any evaluated audio stream is to be transmitted upstream for mixing, based on said received control information. Also, Omran teaches not playing back a source Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Omran and Lim, with the stream control method of Kukosa, as it can improve scalability(Para [0010], Ln 1-7).
Regarding Claim 10:
The combination of Omran and Lim teaches the method of claim 1, and Omran teaches further comprising:….wherein the at least one transformation operation is executed in an embedding domain prior to decoding the reproduced first source stream and the reproduced second source stream(Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12, We also find that multiplying a partition with a weight factor between 0 and 1 leads to a reduction in volume of the corresponding audio content, a feature that can be used for adjusting the target level of denoising. Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame. See Fig 2 partitions).
The combination of Omran and Lim does not explicitly teach in response to receiving user input at the remote endpoint, applying at least one transformation operation to at least one of a reproduced first source stream and a reproduced second source stream.
In the same field of audio processing, Kukosa teaches in response to receiving user input at the remote endpoint, applying at least one transformation operation to at least one of a reproduced first source stream and a reproduced second source stream(Para [0036], Ln 1-8, control information is defined by the central node. Alternatively, the control information may be fixedly predefined. Definition may be done by algorithm and/or user input. Para [0038], Ln 1-8, The device of this aspect is adapted to execute a method according to any of the preceding aspects, and may in particular be a server, preferably conference server, client, desktop computer, portable computer, tablet, telephone, mobile phone, smart phone, PDA, or the like. Para [0037], Ln 1-5, The characteristic parameter may be a volume parameter, and/or may be adapted to said control information. Abstract, Ln 1-22, receiving audio streams from audio clients connected to the distributed multipoint audio processing node and generating evaluated audio streams by analyzing packets of the received audio streams in terms of at least one audio communication characteristic, and attaching an analysis result information of said analysis to said packets, in each audio stream. Audio streams are selected by deciding on whether or not any evaluated audio stream is to be transmitted further, based on the received control information and/or the analysis result information contained in said evaluated audio streams. Then selected audio streams are transmitted further while discarding evaluated audio streams decided not to be to be transmitted further, without mixing any transmitted audio streams. Para [0017], Ln 1-5, preselecting audio streams by deciding on whether or not any evaluated audio stream is to be transmitted upstream for mixing, based on said received control information).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Omran and Lim, with the stream control method of Kukosa, as it can improve scalability(Para [0010], Ln 1-7).
Regarding Claim 16:
Claim 16 contains similar limitations as Claim 6 and is therefore rejected for the same reasons.
Claim(s) 7-8 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Omran and Lim as applied to claim 1 above, and further in view of Liu et al. (US 20250042431 A1).
Regarding Claim 7:
The combination of Omran and Lim teaches the method of claim 1, but does not teach further comprising passing the at least one of the first source stream and the second source stream through a classifier.
In the same field of source separation, Liu teaches further comprising passing the at least one of the first source stream and the second source stream through a classifier(Para [0031], Ln 1-16, number of sources K and types of those sources currently present in the driving environment are not known apriori. The perception system 130 can also deploy a sound classification model (SCM) 134 that performs classification of sources j, e.g., among various predefined (during training of SCM 134) classes, such as sirens, noise, private speech, valid public speech, and/or the like).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Omran and Lim, with the source classification of Liu, as it can enhance user privacy(Para [0031], Ln 1-22).
Regarding Claim 8:
The combination of Omran, Lim and Liu teaches the method of claim 7, and Omran teaches wherein the first source of the audio signal comprises a main speaker and wherein the second source of the audio signal comprises at least one of a far-field talker, human noises, background noises, and music(Pg 2, Fig 1, Ln 1-2, An input waveform is processed by the encoder to produce one embedding vector per frame. Pg 2, 3.1. Model, Para 1, Ln 1-5, Our model architecture is shown in Fig. 1A. Single-channel audio waveforms x ∈ RT are processed by the encoder, which is a convolutional network with blocks consisting of residual units and a strided convolutional layer. All convolutions have causal padding to allow for real-time streaming of the input waveforms. Pg 2, 3.2. Disentanglement scheme, Para 1, Ln 1-4, Disentangling speech from its environment requires careful de sign of the training objectives, in order to both achieve high quality audio transmission and create structured embeddings with no content leakage between different partitions. Pg 3, 4.1. Separating speech from noise, Para 4, Ln 9-12, We also find that multiplying a partition with a weight factor between 0 and 1 leads to a reduction in volume of the corresponding audio content, a feature that can be used for adjusting the target level of denoising).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Mao et al. (US 20250131919 A1).
Audio codec involving the encoding of embeddings, and entropy based encoding.
Mesgarani et al.(US 20190066713 A1).
Neural network based audio source separation.
Yang et al. “SOURCE-AWARE NEURAL SPEECH CODING FOR NOISY SPEECH COMPRESSION”.
Neural network based source separation and audio coding, involving entropy based bitrate allocation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER G MARLOW whose telephone number is (571)272-4536. The examiner can normally be reached Monday - Thursday 10:00 am - 8:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richmond Dorvil can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEXANDER G MARLOW/ Assistant Examiner, Art Unit 2658
/RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658