DETAILED ACTION
This action is in response to the amendment filed on 06/25/2026.
Response to Amendment
Applicant’s amendment filed on 06/25/2026 has been entered. Claims 1, 4, 6, 8, 11, 13, 14, 17 and 20 have been amended. Claims 3, 10 and 16 have been canceled. Claims 21 – 23 have been added. Claims 1, 2, 4 – 9, 11 – 15 and 17 – 23 are still pending in this application, with claims 1, 8 and 14 being independent.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 4, 5, 7- 9, 11, 12, 14, 15, 17 – 19 and 21 - 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ayrapetian et al. (US 10,522,167) (“Ayrapetian”) in view of YU (US 2017/0178666) and further in view of Matsoukas et al. (US 10,522,134) (“Matsoukas”).
For claims 1 and 8, Ayrapetian discloses a system (Abstract), comprising: one or more processors (Fig.12, 1204; column 21 lines 45 – 48); one or more non-transitory computer-readable (Fig.12, 1206)media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations (column 21 lines 45 – 56; column 22 lines 29 – 54) comprising:
receiving second audio data generated by the first device (Fig.1, 110; column 2 lines 57 – column 3 lines 12) (Fig.1, 130, Fig.7B, 750; column 8 lines 30 – 36, column 16 lines 40 - 42) and associated with a communication session between the first device and a second device (column 3 lines 38- 42); analyzing the second audio data using a deep learning model (single DNN and/or multiple DNNs interpreted as a single deep learning model) that is configured to recognize a first user’s speech (target speech) (Fig.1, 132, Fig.7B, 752; column 5 lines 19 – 63, column 6 lines 18 – 30, column 7 lines 12 – 47, 65 – column 8 line 30, column 8 lines 37 – 43, column 16 lines 41 - 50); based on the analyzing, identifying the first voice of the first user represented in the second audio data (First mask data corresponding to a first user is determined based on the microphone audio data. The mask data corresponds to binary flags which classify each time-frequency unit of the microphone audio data as speech associated with a first user., Fig.1, 132, Fig.5A and 5B, Fig.7B, 752; column 5 lines 19 – 63; column 6 lines 18 – 30; column 7 lines 12 – 47, 65 – column 8 line 30, column 8 lines 37 – 43, column 14 lines 5 – 31, column 16 lines 41 - 50); analyzing the second audio data using the deep learning model to identify a second voice associated with a second user represented in the second audio data (Second mask data corresponding to a second user is determined based on the microphone audio data., Fig.1, 136, Fig.5A and 5B, Fig.7B, 758; column 5 lines 19 – 63; column 6 lines 18 – 30; column 7 lines 12 – 47, 65 – column 8 line 30, column 8 lines 55 – 64, column 14 lines 5 – 31, column 16 lines 59 - 67); using the deep learning model, enhancing the first voice of the first user represented in the second audio data (The binary mask data generated by the deep learning model is used to determine beamformer parameters and a first vector corresponding to a first direction associated with the first user. A beamformed audio corresponding to this first direction/first user is generated. This beamformed audio is enhanced by removing noise/unwanted speech audio., Fig.1, 134, Fig.7B, 754 and 756; column 8 lines 36 – 54, column 16 lines 50 - 58, column 17 lines 29 - 56); using the deep learning model, suppressing the second voice of the second user represented in the second audio data (The binary mask data generated by the deep learning model is used to determine beamformer parameters and a second vector corresponding to a second direction associated with the second user. A beamformed audio corresponding to the second user/second direction is generated. This beamformed audio is suppressed by removing it from the beamformed audio corresponding to the first user/first direction , Fig.1, 138 and 144, Fig.7B, 760, 762, 770, 772; column 8 lines 36 – 55, column 16 lines 60 – column 17 line 9, 29 - 56); and sending the second audio data to the second device via the communication session (column 3 lines 38- 42; column 4 lines 1 - 13; column 9 lines 25 - 35).
Yet, Ayrapetian fails to teach the following: receiving, using the microphone, first audio data generated by a first device, the first audio data representing a first voice of a first user; generating a voice profile that represents voice characteristics of the first voice of the first user represented in the first audio data; generating an embedding from the voice profile; and configuring the deep learning model by augmenting the deep learning model with the embedding.
However, YU discloses a system and method for performing multi-speaker speech separation (Abstract), comprising the following: a machine learning model (RNN model, Fig.1, 120) is configured to perform speech separation to recognize a user’s (target speaker) speech by training the model using features (voice characteristics) of the user’s speech ([0045 – 0054]). Furthermore, the user’s speech features used to perform the training are provided by a device capturing the user’s speech (deployment feedback loop, [0021]).
Moreover, Matsoukas discloses a system and method for recognizing a user (Abstract), comprising the following: a voice profile/training data comprising an embedding (feature vector) is generated based on a user’s speech captured by a first device (column 3 lines 30 – column 4 line 15; column 19 lines 50 – column 20 line 3; (column 23 lines 15 – column 24 line 5).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve Ayrapetian’s invention in the same way that YU’s invention has been improved to achieve the following, predictable results for the purpose of improving beamforming and/or noise cancellation in electronic device by using trained deep learning models (Ayrapetian, column 1 lines 55 – column 2 line 40) (Yu, [0001]): the deep learning model is further configured by augmenting (training) the model using voice characteristics/features of a first user (target speaker). Furthermore, the voice characteristics/features are captured and obtained from the first device (feedback deployment loop).
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Ayrapetian and YU in the same way that Matsouka’s invention has been improved to achieve the following, predictable results for the purpose improving beamforming and/or noise cancellation in electronic device by using trained deep learning models which have been configured to accept embeddings as input (Ayrapetian, Feature vectors are generated by neural network techniques, column 1 lines 55 – column 2 line 40 and column 7 lines 12 – 40) (Yu, [0001]): the voice characteristics/features further comprise a voice profile of the first user (target speaker), wherein the voice profile is generated by embedding the first user’s speech which has been captured by the first device.
For claim 14, Ayrapetian discloses a first device (Abstract; Fig.1, 110 and Fig.12; column 2 lines 57 – column 3 lines 12) comprising: one or more processors (Fig.12, 1204; column 21 lines 45 - 48); a microphone (Fig.12, 112; column 21 lines 64 – column 22 line 5); and one or more computer-readable media (Fig.12, 1206; column 21 lines 48 - 52) storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations (column 21 lines 52 – 56) comprising: generating, using the microphone, second audio data (Fig.1, 130, Fig.7B, 750; column 8 lines 30 – 36, column 16 lines 40 - 42) associated with a communication session between the first device and a second device (column 3 lines 38- 42); analyzing the second audio data using a deep learning model (single DNN and/or multiple DNNs interpreted as a single deep learning model) that is configured to recognize a first user’s speech (target speech) (Fig.1, 132, Fig.7B, 752; column 5 lines 19 – 63, column 6 lines 18 – 30, column 7 lines 12 – 47, 65 – column 8 line 30, column 8 lines 37 – 43, column 16 lines 41 - 50); based on the analyzing, identifying the first voice of the first user represented in the second audio data (First mask data corresponding to a first user is determined based on the microphone audio data. The mask data corresponds to binary flags which classify each time-frequency unit of the microphone audio data as speech associated with a first user., Fig.1, 132, Fig.5A and 5B, Fig.7B, 752; column 5 lines 19 – 63; column 6 lines 18 – 30; column 7 lines 12 – 47, 65 – column 8 line 30, column 8 lines 37 – 43, column 14 lines 5 – 31, column 16 lines 41 - 50); analyzing the second audio data using the deep learning model to identify a second voice associated with a second user represented in the second audio data (Second mask data corresponding to a second user is determined based on the microphone audio data., Fig.1, 136, Fig.5A and 5B, Fig.7B, 758; column 5 lines 19 – 63; column 6 lines 18 – 30; column 7 lines 12 – 47, 65 – column 8 line 30, column 8 lines 55 – 64, column 14 lines 5 – 31, column 16 lines 59 - 67); using the deep learning model, enhancing the first voice of the first user represented in the second audio data (The binary mask data generated by the deep learning model is used to determine beamformer parameters and a first vector corresponding to a first direction associated with the first user. A beamformed audio corresponding to this first direction/first user is generated. This beamformed audio is enhanced by removing noise/unwanted speech audio., Fig.1, 134, Fig.7B, 754 and 756; column 8 lines 36 – 54, column 16 lines 50 - 58, column 17 lines 29 - 56); using the deep learning model, suppressing the second voice of the second user represented in the second audio data (The binary mask data generated by the deep learning model is used to determine beamformer parameters and a second vector corresponding to a second direction associated with the second user. A beamformed audio corresponding to the second user/second direction is generated. This beamformed audio is suppressed by removing it from the beamformed audio corresponding to the first user/first direction , Fig.1, 138 and 144, Fig.7B, 760, 762, 770, 772; column 8 lines 36 – 55, column 16 lines 60 – column 17 line 9, 29 - 56); and sending the second audio data to the second device via the communication session (column 3 lines 38- 42; column 4 lines 1 - 13; column 9 lines 25 - 35).
Yet, Ayrapetian fails to teach the following: receiving, using the microphone, first audio data generated by a first device, the first audio data representing a first voice of a first user; generating a voice profile that represents voice characteristics of the first voice of the first user represented in the first audio data; and configuring the deep learning model by augmenting the deep learning model with the voice profile.
However, YU discloses a system and method for performing multi-speaker speech separation (Abstract), comprising the following: a machine learning model (RNN model, Fig.1, 120) is configured to perform speech separation to recognize a user’s (target speaker) speech by training the model using features (voice characteristics) of the user’s speech ([0045 – 0054]). Furthermore, the user’s speech features used to perform the training are provided by a device capturing the user’s speech (deployment feedback loop, [0021]).
Moreover, Matsoukas discloses a system and method for recognizing a user (Abstract), comprising the following: a voice profile/training data is generated based on a user’s speech captured by a first device (column 3 lines 30 – column 4 line 15; column 19 lines 50 – column 20 line 3; (column 23 lines 15 – column 24 line 5).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve Ayrapetian’s invention in the same way that YU’s invention has been improved to achieve the following, predictable results for the purpose of improving beamforming and/or noise cancellation in electronic device by using trained deep learning models (Ayrapetian, column 1 lines 55 – column 2 line 40) (Yu, [0001]): the deep learning model is further configured by augmenting (training) the model using voice characteristics/features of a first user (target speaker). Furthermore, the voice characteristics/features are captured and obtained from the first device (feedback deployment loop).
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Ayrapetian and YU in the same way that Matsouka’s invention has been improved to achieve the following, predictable results for the purpose improving beamforming and/or noise cancellation in electronic device by using trained deep learning models (Ayrapetian, Feature vectors are generated by neural network techniques, column 1 lines 55 – column 2 line 40 and column 7 lines 12 – 40) (Yu, [0001]): the voice characteristics/features further comprise a voice profile of the first user (target speaker), wherein the voice profile is generated from first user’s speech which has been captured by the first device.
For claims 2, 9 and 15, Ayrapetian further discloses: analyzing the second audio data using the model to identify an unwanted noise signal represented in the second audio data (Ayrapetian, Fig.1, 140, Fig.5A and 5B, Fig.7B, 764; column 5 lines 19 – 63; column 6 lines 18 – 30; column 7 lines 12 – 47, 65 – column 8 line 30, column 9 lines 7 - 16, column 14 lines 5 – 31, column 17 lines 10 - 19); and prior to sending the second audio data, using the model to suppress the unwanted noise signal represented in the second audio data (Ayrapetian, Fig.1, 142 and 144, Fig.7B, 766, 768, 770, 772; column 9 lines 16 - 3 , column 17 lines 10 – 56).
For claims 4, 11 and 17, Ayrapetian, YU and Matsouka further disclose, wherein the voice profile is a first voice profile (Ayrapetian, speaker data used to train deep learning model, column 2 lines 25 - 40) (YU, [0019 – 0021] [0045 – 0054]) (Matsouka, Fig.5, user 1; column 3 lines 30 – column 4 line 15; column 19 lines 50 – column 20 line 3), the operations further comprising: receiving third audio data representing a third voice of a third user (Ayrapetian, column 2 lines 40 – 56) (Matsouka, Fig.5, user 3; column 3 lines 30 – column 4 line 15; column 19 lines 50 – column 20 line 3); generating a second voice profile that represents second voice characteristics of the third voice of the third user represented in the third audio data (Ayrapetian, speaker data used to train deep learning model, column 2 lines 25 - 40) (Yu, [0019 – 0021] [0045 – 0054]) (Matsouka, Fig.5, user 1; column 3 lines 30 – column 4 line 15; column 19 lines 50 – column 20 line 3); and determining, using the first voice profile and the second voice profile, that particular voice characteristics represented in the second audio data are more closely correlated to the voice characteristics than the second voice characteristics (Matsouka, The voice profiles/training data are compared against the incoming data to determine a match., column 23 lines 15 – column 24 line 33).
For claims 5, 12 and 18, Ayrapetian further discloses: receiving third audio data (loudspeaker noise) generated by the first device (Ayrapetian, Fig.1, 130; column 2 lines 41 – 56, column 8 lines 30 - 36) and associated with the communication session (Ayrapetian, column 3 lines 38- 42), the third audio data representing noise in an environment of the first device (Ayrapetian, column 2 lines 41 – 56); determining, using the voice profile (Ayrapetian, Speaker data for a first speaker has been used to train the deep learning model. By using the deep learning model, the speaker data is also used.,, column 2 lines 25 - 40) (YU, [0019 – 0021] [0045 – 0054]) (Matsouka, Fig.5, user 1; column 3 lines 30 – column 4 line 15; column 19 lines 50 – column 20 line 3), that the third audio data does not represent the first voice of the first user (Ayrapetian, Fig.5B, column 13 lines 42 – 63; column 14 lines 20 - 31; and based at least in part on the third audio data not representing the first voice of the first user, refraining from sending the third audio data to the second device (Ayrapetian, The binary mask data generated by the deep learning model is used to determine beamformer parameters and a third vector corresponding to a third direction associated with noise. A beamformed audio corresponding to the noise/third direction is generated. This beamformed audio is suppressed by removing it from the beamformed audio corresponding to the first user/first direction., Fig.1, 142 and 144, Fig.7B, 766, 768, 770, 772; column 9 lines 16 - 3 , column 17 lines 10 – 56).
For claim 7, Ayrapetian further discloses, wherein one or more steps are performed by the first device (Ayrapetian, column 5 lines 19 – 41 and column 6 lines 57- 63; claim 12 – claim 18).
For claim 19, Ayrapetian further discloses, wherein the model is a deep learning model (Ayrapetian, column 5 lines 19 – 67).
For claim 21 - 23, Ayrapetian, YU and Matsouka further disclose, wherein the one or more models comprise a speaker-specific neural network that has been trained using speech signals of the first user (Ayrapetian, speaker data used to train deep learning model, column 2 lines 25 – 40; column 5 lines 42 - 67) (Yu, [0019 – 0021] [0045 – 0054]) (Matsouka, Fig.5, user 1; column 3 lines 30 – column 4 line 15; column 19 lines 50 – column 20 line 3)
Claim(s) 6, 13 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ayrapetian et al. (US 10,522,167) (“Ayrapetian”) in view of YU (US 2017/0178666), and further in view of Matsoukas et al. (US 10,522,134) (“Matsoukas”) and further in view of Tachibana (US 2008/0086301).
For claims 6, 13 and 20, the combination of Ayrapetian, YU and Matsoukas fails to teach the following: receiving video data generated by a first device that is associated with the second audio data of the communication session; adding a tag to the video data that indicates that noise suppression associated with speaker identification is being performed for the communication session; and sending the video data that includes the tag to the second device via the communication session such that a visual representation of the tag is presented on the second device.
However, Tachibana discloses a discloses an audio communication apparatus and method (Abstract), comprising the following: receiving video data generated by a first device that is associated with audio data of a communication session ([0059] [0078]); adding a tag to the video data that indicates that noise suppression associated with speaker identification is being performed for the communication session (In a TV phone, voice data is transmitted with the image data. A voice correction function is applied to the voice data. A type and parameter of the voice correction is added to the voice data. Since the image and voice data are transmitted together, both are tagged with the type and parameter of the voice correction ([0030 - 0032] [0078]); and sending the video data that includes the tag to the second device via the communication session such that a visual representation of the tag is presented on the second device ([0032 – 0037] [0078]).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Ayrapetian, YU and Matsoukas in the same way that Tachibana’s invention has been improved to achieve the following, predictable results for the purpose of increasing user satisfaction by notifying a receiving device that speech enhancement is being performed so that the receiving device can request a modification of the speech enhancement process based on preference (Tachibana, [0007 – 0010]): further receiving video data generated by a first device that is associated with the second audio data of the communication session; adding a tag to the video data that indicates that noise suppression associated with speaker identification is being performed for the communication session; and sending the video data that includes the tag to the second device via the communication session such that a visual representation of the tag is presented on the second device.
Response to Arguments
Applicant’s arguments with respect to claim(s) 14, 15, 17 – 20 and 23 have been considered but are moot in view of the new ground(s) of rejection
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Tu et al. (“Speech Separation Based on Improved Deep Neural Networks with Dual Outputs of Speech Features for Both Target and Interfering Speakers”)- the inventive concept of separating audio signals using a DNN.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SONIA L GAY whose telephone number is (571)270-1951. The examiner can normally be reached Monday-Friday 9-5 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SONIA L GAY/Primary Examiner, Art Unit 2657