DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on 05/27/2026. Claims 1-3, 5-17 and 19-22 are pending and have been examined.
Any previous objection/rejection not mentioned in this Office Action has been withdrawn by the Examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Applicant’s Amendments and Arguments
The Applicant’s amendments to overcome the 35 USC 101 rejections of claims 16 and 22 have been considered and withdrawn as a result of the amendments.
The Applicant has amended the independent claims to recite further limitations “wherein the encoder transforms the feature by reducing a size of the feature, the size of the feature is divided, and features with a size after the size of the feature being divided are each input to a corresponding one of the plurality of sub-neural networks.” Hence, a new grounds for rejection has been made in view of the newly added limitation and reference as mapped below. Therefore, these arguments are moot.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 5, 8-11, 14-17, and 20-22 are rejected under 35 U.S.C. 103 as being unpatentable over Sawata (“All for One and One for All: Improving Music Separation by Bridging Networks”, cited in IDS) in view of Brocal et al. (“Conditioned-U-NET: Introducing a Control Mechanism in the U-NET for Multiple Source Separations”).
As to claim 1 and 15-16, Sawata teaches a program stored on a non-transitory computer-readable medium for causing a computer to execute an information processing method (see Figure 3b, sect 3.1, where implementation described which requires an inherent processor and associated program to operate and realize the model and where description to sampling of audio based on and training is described), the information processing method comprising:
generating, by a neural network, sound source separation information for separating a predetermined sound source signal from a mixed sound signal containing a plurality of sound source signals (see Figure 3b, where input mixture is inputted to the architecture and then output of the separated sources provided);
transforming, by an encoder included in the neural network, a feature extracted from the mixed sound signa (see Figure 3b, where left side of the figure the input mixture is passed through affine + BN and then Nonlinearity block);
inputting a process result from the encoder to each of a plurality of sub-neural networks included in the neural network (see Figure 3b, where output from averaging block into the middle BLSTM blocks where various blocks can be seen parallel to the top to bottom and then in between has ellipsis); and
inputting the process result from the encoder and a process result from each of the plurality of sub-neural networks to a decoder included in the neural network (see Figure 3b, output of the 2nd averaging block on the right side, where input into Affine+BN followed by Nonlinearity blocks are provided in sequence)
wherein the encoder transforms the feature by reducing a size of the feature (see sect. 1, right column, 2ⁿᵈ paragraph, last four lines, where input is in frequency domain as a spectrogram and see sect 3.1, 1st paragraph where STFT magnitude domain and see Figure 4b, where affine+BN as well as the averaging inputs performed on the input data which reduces the size from the frequency domain to vectorized form which is then inputted into the BLSTMs)
However, Sawata does not specifically teach, the size of the feature is divided, and features with a size after the size of the feature being divided are each input to a corresponding one of the plurality of sub-neural networks. The Examiner notes that in Sawata the affine transformation +BN are interpreted as the encoder and in Brocal below such affine layers are part of the encoder. This is mentioned in Figure 3 and in sect 3.1, right column, 2nd full paragraph of Brocal) thus signifying an overlap of subject matter.
Brocal does teach wherein the encoder transforms the feature by reducing a size of the feature (see page 3, right column, “encoder” section, where encoder creates a compressed and deep representation of the mixture thus noting a size reduction of the input feature).
the size of the feature is divided (see page 3, right column, “encoder” section, where each layer halves the size of the input and doubles the number of channels, see output of encoder in Figure 3), and
features with a size after the size of the feature being divided are each input to a corresponding one of the plurality of sub-neural networks (see Figure 3, the feature maps of the encoder are connected to the layers of the decoder and see page 3, right column, “skip-connections” sections).
Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the encoder as taught by Sawata with the size as taught by Brocal in order to reduce dimensionality while preserving the relevant information for separation (see Brocal page 3, right column, “encoder” section).
As to claim 14, apparatus claim 1 and 15-16 and method claim 14 are related as apparatus and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 14 is similarly rejected under the same rationale as applied above with respect to each function of the apparatus claim. As to claim 15, Sawata further teaches outputting the predetermined sound source signal (see Figure 3b, Source 1-j is output from the right side).
As to claim 2, Sawata in view of Brocal teach all of the limitations as in claim 1.
Furthermore, Sawata does teach wherein each of the plurality of sub-neural networks includes a recurrent neural network that uses at least one of a temporally past process result or a temporally future process result for current input (see Figure 3b, BLSTM blocks in the middle) (e.g. The examiner notes that RNNs operate on past data which is an intrinsic feature of RNN architecture, where BLSTM is a form of RNN).
As to claim 3, Sawata in view of Brocal teach all of the limitations as in claim 1.
Furthermore, Sawata does teach Wherein the recurrent neural network includes a neural network using a gated recurrent (GRU) algorithm or a long short term memory (LSTM) algorithm (see Figure 3b, where BLSTM is shown in the middle blocks).
As to claim 5, Sawata in view of Brocal teaches all of the limitations as in claim 1, above.
Brocal does teach wherein the feature and the size of the feature are defined by a multidimensional vector and a number of dimensions of the multidimensional vector, respectively, and the encoder reduces the number of dimensions of the multidimensional vector (see page 3, right column, “encoder” section, multidimensional signal is input as seen in Figure 3, and encoder reduces number of dimensions)
As to claim 8, Sawata in view of Brocal teach all of the limitations as in claim 1.
Furthermore, Sawata does teach wherein the encoder includes affine transformation circuitry (see Figure 3b. where affine transformation is included in the left most blocks).
As to claim 9, Sawata in view of Brocal teach all of the limitations as in claim 1.
Furthermore, Sawata does wherein the decoder generates the sound source separation information based on the process result from the encoder and the process result from each of the plurality of sub-neural networks (see Figure 3b, where right most side of the figure shows the separated sources J through 1. based on propagation of the input through the architecture).
As to claim 10, Sawata in view of Brocal teach all of the limitations as in claim 1.
Furthermore, Sawata does teach wherein the decoder includes affine transformation circuitry (see Figure 3b, where affine transformation is included in the right most blocks).
As to claim 11, Sawata in view of Brocal teach all of the limitations as in claim 1.
Furthermore, Sawata does teach wherein a feature extraction circuitry extracts the feature from the mixed sound signal (see sect. 1, right column, 2nd paragraph, last four lines, where input is in frequency domain as a spectrogram and see sect 3.1, 1st paragraph where STFT magnitude domain).
As to claim 17 and 21-22, Sawata in view of Brocal teach all of the limitations as in claim 1.
Furthermore, Sawata teaches a program stored on a non-transitory computer-readable medium for causing a computer to execute an information processing method (see Figure 3b, sect 3.1, where implementation described which requires an inherent processor and associated program to operate and realize the model and where description to sampling of audio based on and training is described), the information processing method comprising:
generating, by each of a plurality of neural networks, sound source separation information for separating a different sound source signal from a mixed sound signal containing a plurality of sound source signals (see Figure 3b, where input mixture is inputted to the architecture and then output of the separated sources provided based on the different networks that lead to the respective determined sources);
transforming, by an encoder included in one of the plurality of neural networks, a feature extracted from the mixed sound signal (see Figure 3b, where left side of the figure the input mixture is passed through affine + BN and then Nonlinearity block); and
inputting a process result from the encoder to each of a plurality of sub-neural networks included in each of the plurality of neural networks (see Figure 3b, where output from averaging block into the middle BLSTM blocks where various blocks can be seen parallel to the top to bottom and then in between has ellipsis).
As to claim 20, apparatus claim 17 and 21-22 and method claim 20 are related as apparatus and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 20 is similarly rejected under the same rationale as applied above with respect to each function of the apparatus claim. As to claim 21, Sawata further teaches outputting the different sound source signals from each of the plurality of neural networks (see Figure 3b, Source 1-j is output from the right side).
Claim(s) 12, 13, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Sawata in view of Brocal, as applied in claim 1 and 17, above, and further in view of Soler (“Music Source Separation Using Deep Neural Networks”, 2020)
As to claim 12, Sawata in view of Brocal teaches all of the limitations as in claim 1 above.
However, Sawata in view of Brocal do not specifically teach wherein operation circuitry multiplies the feature of the mixed sound signal by the sound source separation information output from the decoder.
Soler does teach wherein operation circuitry multiplies the feature of the mixed sound signal by the sound source separation information output from the decoder (see Fig. 3.2, where input from mix spectrograms is multiplied with output from decoding, and see page 25, sections skipped connection and output stage describe where output spectrogram from decoder is multiplied with input magnitude spectrogram).
Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the NN architecture as taught by Sawata in view of Brocal with the multiplication as taught by Soler in order to be able to learn how much each TF bin belongs to the target source (see Soler, page 25, Skipper Connection, last bullet).
As to claim 13, Sawata in view of Brocal in view of Soler teaches all of the limitations as in claim 12 above.
Furthermore, Soler teaches wherein separated sound source signal generation circuitry generates the predetermined sound source signal based an operation result from the operation circuitry (see Figure 3.2, where upward arrow shows target spectrograms generated after the multiplication).
As to claim 19, Sawata in view of Brocal teaches all of the limitations as in claim 17, above.
However, Sawata in view of Brocal do not teach wherein operation circuitry included in each of the plurality of neural networks multiplies the feature of the mixed sound signal by the sound source separation information output from a decoder, and filter circuitry separates the predetermined sound source signal based on process results from the operation circuitry.
Soler does teach wherein operation circuitry included in each of the plurality of neural networks multiplies the feature of the mixed sound signal by the sound source separation information output from a decoder (see Fig. 3.2, where input from mix spectrograms is multiplied with output from decoding, and see page 25, sections skipped connection and output stage describe where output spectrogram from decoder is multiplied with input magnitude spectrogram), and filter circuitry separates the predetermined sound source signal based on process results from the operation circuitry (see Figure 3.2, where upward arrow shows target spectrograms generated after the multiplication)..
Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the NN architecture as taught by Sawata in view of Brocal with the multiplication as taught by Soler in order to be able to learn how much each TF bin belongs to the target source (see Soler, page 25, Skipper Connection, last bullet).
Allowable Subject Matter
Claims 6-7 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
None of the prior art cited above teaches the combination of limitations as recited from the independent claims and the limitations set forth in claims 6 and 7.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PARAS D SHAH whose telephone number is (571)270-1650. The examiner can normally be reached Monday-Thursday 7:30AM-2:30PM, 5PM-7PM (EST), Friday 8AM-noon (EST).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PARAS D SHAH can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Paras D Shah/ Supervisory Patent Examiner, Art Unit 2653
08/20/2026