Prosecution Insights
Last updated: October 01, 2026
Application No. 18/887,462

VOICE CUSTOMIZATION FOR SYNTHETIC SPEECH GENERATION

Non-Final OA §102§103§DOUBLEPATENT
Filed
Sep 17, 2024
Priority
Feb 14, 2022 — provisional 63/309,762 +1 more
Examiner
SERRAGUARD, SEAN ERIN
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Amazon Technologies Inc.
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
112 granted / 162 resolved
+7.1% vs TC avg
Strong +34% interview lift
Without
With
+34.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
20 currently pending
Career history
188
Total Applications
across all art units

Statute-Specific Performance

§101
8.1%
-31.9% vs TC avg
§103
50.1%
+10.1% vs TC avg
§102
19.7%
-20.3% vs TC avg
§112
20.0%
-20.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 162 resolved cases

Office Action

§102 §103 §DOUBLEPATENT
CTNF 18/887,462 CTNF 95611 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Double Patenting 08-33 AIA The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg , 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman , 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi , 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum , 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel , 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington , 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA. A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA/25, or PTO/AIA/26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 5 and 13 of U.S. Pat. No. 12,100,383 in view of combinations of Zhang (CN 113436608 A, hereinafter Zhang ), Aggarwal (U.S. Pat. No. 11017763, hereinafter Aggarwal), Minnen (U.S. Pat. App. Pub. No. 2020/0027247, hereinafter Minnen ), and Golman (U.S. Pat. App. Pub. No. 2023/0223011, hereinafter Golman ). Claims 5 and 13 of the issued patent match to those of the instant application, as shown in claim chart provided below. Limitations not expressly recited in the claims of the issued patent are mapped to Zhang , Minnen , Aggarwal , and Golman , respectively. All mappings are listed in the claim chart. Explanations of how said limitations are taught and motivations for the combination are substantially similar to the explanations and motivations provided in the mapping of the rejections below. To avoid unnecessary duplication, said explanations, mappings, and motivations are not duplicated here. U.S. App. No.: 18/887,462 U.S. Pat. No. 12,100,383 Claim 1 A computer-implemented method comprising Claim 5 A method comprising: (the method described in the 12/100,383 patent is inherently computer implemented, as necessitated by, at least, the density based model) : receiving first audio data representing first speech receiving first audio data representing first speech; ; receiving first data representing first voice characteristics of the first speech receiving first data representing first voice characteristics of the first speech; ; based at least in part on the first audio data and the first data , generating first compressed data representing the first speech processing the first audio data using a first model and the first data to generate second data representing a latent representation of the first speech, the first model representing a density-based model; performing data compression on the second data to generate first compressed data; ; decompressing the first compressed data to determine second data representing the first speech decompressing the first compressed data to determine second data representing the latent representation of the first speech; ; receiving third data representing second voice characteristics for synthesized speech, receiving third data representing second voice characteristics for synthesized speech; wherein the second voice characteristics are different from the first voice characteristics Zhang at [0001]-[0002], [0077], [0088], [0091] ; processing the second data and the third data to generate second audio data processing the second data using a second model and the third data to generate second audio data , the second model representing an inverse of the first model; and ; and generating, using the second audio data, audio representing synthesized speech corresponding to the second voice characteristics generating, using the second audio data, an audio signal representing synthesized speech, the synthesized speech corresponding to the second voice characteristics Claim 2 The computer-implemented method of claim 1, Claim 5 wherein the first speech is captured by a first device and the audio is output by a second device different from the first device. Aggarwal at Col. 1, lines 53-65; Col. 2, lines 45-56; Col. 3, lines 1-15, and 35-38; Col. 16, lines 28-36 Claim 3 The computer-implemented method of claim 1 Claim 5 , further comprising: causing the first compressed data to be sent from a first device to a second device. Minnen at [0060] Claim 4 The computer-implemented method of claim 3, Claim 5 wherein the first data is not sent to the second device. Zhang at [0089], [0091] Claim 5 The computer-implemented method of claim 1 Claim 5 , further comprising: processing, by a first device, the first data representing first voice characteristics of the first speech to determine modified voice characteristics Zhang at [0089], [0091] ; and sending, from the first device to a second device, data representing the modified voice characteristics. Minnen at [0060] Claim 6 The computer-implemented method of claim 5, Claim 5 wherein the data representing the modified voice characteristics comprises the third data. Zhang at [0089], [0091] Claim 7 The computer-implemented method of claim 1, Claim 5 wherein: generation of the first compressed data uses a first machine learning model processing the first audio data using a first model and the first data to generate second data representing a latent representation of the first speech, the first model representing a density-based model [ machine learning] ; performing data compression on the second data to generate first compressed data; ; and generation of the second audio data uses a second machine learning model corresponding to an inverse of the first machine learning model. processing the second data using a second model and the third data to generate second audio data, the second model representing an inverse of the first model ; Claim 8 The computer-implemented method of claim 1 Claim 5 , further comprising: processing the first audio data to determine fourth data representing a latent representation of the first speech, Minnen at [0059], [0066] wherein the first compressed data is based at least in part on the latent representation. Minnen at [0061], [0067] Claim 9 The computer-implemented method of claim 8, Claim 5 wherein decompressing the first compressed data to determine the second data comprises decompressing the first compressed data to determine the latent representation. Minnen at [0080]-[0081] Claim 10 The computer-implemented method of claim 1, Claim 5 wherein receiving the third data comprises: receiving an input to a user interface component Golman at [0110] , the input corresponding to at least one voice characteristic Golman at [0070]-[0071] ; and based at least in part on the input, determining the third data. Golman at [0070] Claim 11 A system comprising: at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system Claim 13 A system, comprising: at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to: receive first audio data representing first speech to: receive first audio data representing first speech ; receive first data representing first voice characteristics of the first speech ; receive first data representing first voice characteristics of the first speech ; based at least in part on the first audio data and the first data, generate first compressed data representing the first speech ; process the first audio data using a first model and the first data to generate second data representing a latent representation of the first speech, the first model representing a density-based model; perform data compression on the second data to generate first compressed data ; decompress the first compressed data to determine second data representing the first speech ; decompress the first compressed data to determine second data representing the latent representation of the first speech ; receive third data representing second voice characteristics for synthesized speech, ; receive third data representing second voice characteristics for synthesized speech wherein the second voice characteristics are different from the first voice characteristics Zhang at [0001]-[0002], [0077], [0088], [0091] ; process the second data and the third data to generate second audio data ; process the second data using a second model and the third data to generate second audio data , the second model representing an inverse of the first model ; and generate, using the second audio data, audio representing synthesized speech corresponding to the second voice characteristics ; and generate, using the second audio data, an audio signal representing synthesized speech, the synthesized speech corresponding to the second voice characteristics Claim 12 The system of claim 11, Claim 13 wherein the first speech is captured by a first device and the audio is output by a second device different from the first device. Aggarwal at Col. 1, lines 53-65; Col. 2, lines 45-56; Col. 3, lines 1-15, and 35-38; Col. 16, lines 28-36 Claim 13 The system of claim 11, Claim 13 wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to: causing the first compressed data to be sent from a first device to a second device. Minnen at [0060] Claim 14 The system of claim 13, Claim 13 wherein the first data is not sent to the second device. Zhang at [0089], [0091] Claim 15 The system of claim 11, Claim 13 wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to: processing, by a first device, the first data representing first voice characteristics of the first speech to determine modified voice characteristics Zhang at [0089], [0091] ; and sending, from the first device to a second device, data representing the modified voice characteristics. Minnen at [0060] Claim 16 The system of claim 15, Claim 13 wherein the data representing the modified voice characteristics comprises the third data. Zhang at [0089], [0091] Claim 17 The system of claim 11, Claim 13 wherein: generation of the first compressed data uses a first machine learning model processing the first audio data using a first model and the first data to generate second data representing a latent representation of the first speech, the first model representing a density-based model [machine learning] ; performing data compression on the second data to generate first compressed data; ; and generation of the second audio data uses a second machine learning model corresponding to an inverse of the first machine learning model. processing the second data using a second model and the third data to generate second audio data, the second model representing an inverse of the first model ; Claim 18 The system of claim 11, Claim 13 wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to: processing the first audio data to determine fourth data representing a latent representation of the first speech, Minnen at [0059], [0066] wherein the first compressed data is based at least in part on the latent representation. Minnen at [0061], [0067] Claim 19 The system of claim 18, Claim 13 wherein the instructions that cause the system to decompress the first compressed data to determine the second data comprise instructions that, when executed by the at least one processor, cause the system to decompress the first compressed data to determine the latent representation. Minnen at [0080]-[0081] Claim 20 The system of claim 11, Claim 13 wherein the instructions that cause the system to receive the third data comprise instructions that, when executed by the at least one processor, cause the system to: receive an input to a user interface component Golman at [0110] , the input corresponding to at least one voice characteristic Golman at [0070]-[0071] ; and based at least in part on the input, determine the third data. Golman at [0070] Claim Objections 07-29-01 AIA Claim s 13, 15, and 18 are objected to because of the following informalities: Regarding claim 13, the phrase “causing the first compressed data…” should read as “ causing cause the first compressed data…”. Regarding claim 15, the phrase “processing, by a first device…” should read as “ processing process , by a first device…”. Regarding claim 15, the phrase “sending, from the first device…” should read as “ sending send , from the first device…”. Regarding claim 18, the phrase “processing the first audio data…” should read as “ processing process the first audio data…” . Appropriate correction is required. Claim Rejections - 35 USC § 102 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-15-aia AIA Claim(s) 1 and 11 is/are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Zhang (CN 113436608 A, hereinafter Zhang ) . Regarding claim 1, Zhang discloses A computer-implemented method comprising (Discloses systems and methods described with reference to a “dual-stream speech conversion device, comprising: a memory and at least a processor, in which instructions are stored; At least one processor calls the instruction in the memory to enable the dual-stream speech conversion device to execute the dual-stream speech conversion method described above”; Zhang, ¶ [0059]): receiving first audio data representing first speech (The system receives the “voice signal includes the time domain signal of the voice data and the frequency domain signal of the voice data” and where “The discrete Fourier transform and logarithmic spectral conversion are performed on the time-domain signals in the preprocessed speech signals to obtain the first spectral information {first audio data}.”; Zhang, ¶ [0072]); receiving first data representing first voice characteristics of the first speech (Discloses “extracting the features of the speech signal to obtain the... target voiceprint features, target emotional features {...representing the first voice characteristics}” where the above described target “features of the speech signal” are the first data.; Zhang, ¶ [0091]); based at least in part on the first audio data and the first data, generating first compressed data representing the first speech (The sampled data including “voiceprint feature vector” and “emotional feature vector” is “compressed through the vector compression layer in the vector processing layer to obtain the audio compression vector”; Zhang, ¶ [0089]); decompressing the first compressed data to determine second data representing the first speech (“The server calls the left side of the affine coupling channel in the affine coupling layer… to perform upsampling” and “upsampling, vector division, mapping relationship learning and affine transformation based on speech features,” which expands the latent space data for transformation {decompressing the first compressed data}.; Zhang, ¶ [0079], [0097]); receiving third data representing second voice characteristics for synthesized speech, (The system relies on voice parameters (also referred to as “preset speech features” or “the value of the target learning variable”) which is used to “indicate the weight ratio for adaptive addition,” where “two-way mapping learning based on the preset speech features is carried out to obtain the candidate dual-flow speech synthesis model” and “the preset speech features are used to indicate the converted speech,” thereby defining the new voice the audio will be converted into.; Zhang, ¶ [0081], [0087]-[0088]) wherein the second voice characteristics are different from the first voice characteristics (Discloses using “the preset speech features” to “obtain the converted voice feature vector,” as part of a voice conversion process. As such, the “preset speech features{second voice characteristics}” are understood to be different from the “target…features {first voice characteristics}”; Zhang, ¶ [0001]-[0002], [0077], [0088], [0091]); processing the second data and the third data to generate second audio data (The system uses the target features and variables within “the affine coupling layer” and “the dual-flow affine transform” based on the “preset speech features” is performed on “the treated vector to obtain the converted voice feature vector,”; Zhang, ¶ [0077]); and generating, using the second audio data, audio representing synthesized speech corresponding to the second voice characteristics (“Obtain the value of the target learning variable, and use the normalization layer to weighted normalize the converted voice feature vector based on the value of the target learning variable to obtain the target speech.”; Zhang, ¶ [0080]). Regarding claim 11, Zhang discloses A system comprising: at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system (Discloses systems and methods described with reference to a “dual-stream speech conversion device, comprising: a memory and at least a processor, in which instructions are stored; At least one processor calls the instruction in the memory to enable the dual-stream speech conversion device to execute the dual-stream speech conversion method described above”; Zhang, ¶ [0059]) to: receive first audio data representing first speech (The system receives the “voice signal includes the time domain signal of the voice data and the frequency domain signal of the voice data” and where “The discrete Fourier transform and logarithmic spectral conversion are performed on the time-domain signals in the preprocessed speech signals to obtain the first spectral information {first audio data}.”; Zhang, ¶ [0072]); receive first data representing first voice characteristics of the first speech (Discloses “extracting the features of the speech signal to obtain the... target voiceprint features, target emotional features {...representing the first voice characteristics}” where the above described target “features of the speech signal” are the first data.; Zhang, ¶ [0091]); based at least in part on the first audio data and the first data, generate first compressed data representing the first speech (The sampled data including “voiceprint feature vector” and “emotional feature vector” is “compressed through the vector compression layer in the vector processing layer to obtain the audio compression vector”; Zhang, ¶ [0089]); decompress the first compressed data to determine second data representing the first speech (“The server calls the left side of the affine coupling channel in the affine coupling layer… to perform upsampling” and “upsampling, vector division, mapping relationship learning and affine transformation based on speech features,” which expands the latent space data for transformation {decompressing the first compressed data}.; Zhang, ¶ [0079], [0097]); receive third data representing second voice characteristics for synthesized speech, (The system relies on voice parameters (also referred to as “preset speech features” or “the value of the target learning variable”) which is used to “indicate the weight ratio for adaptive addition,” where “two-way mapping learning based on the preset speech features is carried out to obtain the candidate dual-flow speech synthesis model” and “the preset speech features are used to indicate the converted speech,” thereby defining the new voice the audio will be converted into.; Zhang, ¶ [0081], [0087]-[0088]) wherein the second voice characteristics are different from the first voice characteristics (Discloses using “the preset speech features” to “obtain the converted voice feature vector,” as part of a voice conversion process. As such, the “preset speech features{second voice characteristics}” are understood to be different from the “target…features {first voice characteristics}”; Zhang, ¶ [0001]-[0002], [0077], [0088], [0091]); process the second data and the third data to generate second audio data (The system uses the target features and variables within “the affine coupling layer” and “the dual-flow affine transform” based on the “preset speech features” is performed on “the treated vector to obtain the converted voice feature vector,”; Zhang, ¶ [0077]); and generate, using the second audio data, audio representing synthesized speech corresponding to the second voice characteristics (“Obtain the value of the target learning variable, and use the normalization layer to weighted normalize the converted voice feature vector based on the value of the target learning variable to obtain the target speech.”; Zhang, ¶ [0080]) . Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-22-aia AIA Claim s 2 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang as applied to claim (s) 1 and 11 above, and further in view of Aggarwal (U.S. Pat. No. 11017763, hereinafter Aggarwal) . Regarding claim 2, the rejection of claim 1 is incorporated. Zhang discloses all of the elements of the current invention as stated above. However, Zhang fails to expressly recite wherein the first speech is captured by a first device and the audio is output by a second device different from the first device. Aggarwal teaches systems and methods for speech synthesis using normalizing flow. ( Aggarwal , ¶ Col. 1, lines 5-10). Regarding claim 2, Aggarwal teaches Regarding claim 2, Aggarwal discloses wherein the first speech is captured by a first device and the audio is output by a second device different from the first device (Discloses a system for “transform text and/or other audio into synthesized speech” comprising “multiple devices... employed in a single system” where each of the “multiple devices may include overlapping components,” and each of the disclosed components may be “disposed on either the user device 110 and/or the remote system 120.” Thus, the system can include two user devices 110, where the first “user device 110... receives” the “text data 14 (and/or audio data) for transformation into audio data that includes a representation of synthesized speech” and the second “user device 110” then “processes (132) the text data using a trained sequence-to-sequence model (and/or other trained model)” to “output a series of power spectrograms, such as Mel-spectrograms, that each correspond to a certain duration of output audio” which are processed “to determine the audio data.”; Aggarwal, ¶ Col. 1, lines 53-65; Col. 2, lines 45-56; Col. 3, lines 1-15, and 35-38; Col. 16, lines 28-36). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the flow-based voice conversion systems of Zhang to incorporate the teachings of Aggarwal to include wherein the first speech is captured by a first device and the audio is output by a second device different from the first device. Zhang teaches systems and methods for flow-based voice conversion, which mathematically fuse target voice print features and phoneme features into a single latent vector. However, though Zhang considers remote implementation at a server for interaction with a terminal, Zhang fails to expressly contemplate the details of specific hardware topology and the distribution of this process to devices and/or between devices. It is well known that normalizing flow networks are computationally expensive, and a person having ordinary skill in the art (PHOSITA) would recognize that forcing a local device, such as a smart phone, to perform the entire conversion pipeline would result in unacceptable battery drain, overheating, and latency. Aggarwal explicitly addresses the physical hardware topology of distributed speech processing. Specifically, Aggarwal teaches a “synthetic speech processing” system that divides computational components between any combination of user devices and remote systems communicating over a network. As such, a PHOSITA would be motivated to modify the flow-based voice conversion process of Zhang with the distributed multi-device processing scheme of Aggarwal to optimize computational resource allocation and minimize latency in a voice conversion system, as recognized by Aggarwal . ( Aggarwal , ¶ Col. 15, lines 10-36). Regarding claim 12, the rejection of claim 11 is incorporated. Claim 12 is substantially the same as claim 2 and is therefore rejected under the same rationale as above . 07-22-aia AIA Claim s 3-9 and 13-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang as applied to claim (s) 1 and 11 above, and further in view of Minnen (U.S. Pat. App. Pub. No. 2020/0027247, hereinafter Minnen) . Regarding claim 3, the rejection of claim 1 is incorporated. Zhang discloses all of the elements of the current invention as stated above. However, Zhang fails to expressly recite further comprising: causing the first compressed data to be sent from a first device to a second device. Minnen teaches systems and methods for compression and decompression of data and data representations. ( Minnen , ¶ [0059]-[0061]). Regarding claim 3, Minnen teaches further comprising: causing the first compressed data to be sent from a first device to a second device (“the compression and decompression systems may be co-located or remotely located, and compressed data generated by the compression system can be provided to the decompression system in any of a variety of ways. For example, the compressed data may be stored (e.g., in a physical data storage device or logical data storage area), and then subsequently retrieved from storage and provided to the decompression system. As another example, the compressed data may be transmitted over a communications network (e.g., the Internet) to a destination, where it is subsequently retrieved and provided to the decompression system.”; Minnen, ¶ [0060]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the flow-based voice conversion systems of Zhang to incorporate the teachings of Minnen to include further comprising: causing the first compressed data to be sent from a first device to a second device. Zhang teaches systems and methods for flow-based voice conversion, which mathematically fuse target voice print features and phoneme features into a single latent vector. However, though Zhang considers remote implementation at a server, Zhang fails to expressly contemplate the details of the distribution of this process to devices and/or between devices across a network. In distributed speech processing, it is well known in the art that transmitting raw audio requires significant bandwidth, and transmitting uncompressed, high dimensional neural network vectors introduces unacceptable latency. Minnen specifically addresses the technical problem of transmitting high dimensional neural network data across networks, including a specialized architecture designed to generate a compressed representation from a latent representation, and transmit that compressed representation to a second device for further use and processing of the underlying data. A PHOSITA would be motivated to modify the flow-based voice conversion process of Zhang with the transmission architecture of Minnen to further enable the server embodiments described in Zhang , while avoiding known latency difficulties and increasing transmission efficiency across a distributed network, as recognized by Minnen . ( Minnen , ¶ [0046]-[0048]). Regarding claim 4, the rejection of claim 3 is incorporated. Zhang and Minnen disclose all of the elements of the current invention as stated above. Zhang further discloses wherein the first data is not sent to the second device ( Zhang in light of Minnen teaches the transmission of the compressed data of Zhang to a second device. Zhang teaches “extracting the features of the speech signal to obtain the… target voiceprint features, [and] target emotional features” which are the sampled data which is mapped to the first data. The sampled data including “semantic feature vector, voiceprint feature vector, emotional feature vector and phonemic feature vector” is then “compressed through the vector compression layer in the vector processing layer to obtain the audio compression vector,” where the audio compression vector is the first compressed data which is transmitted to the second device in light of the combination with Minnen . As the first compressed data is based on, but does not include, that sampled data, the first data is not part of the data which is sent to the second device.; Zhang, ¶ [0089], [0091]). Regarding claim 5, the rejection of claim 1 is incorporated. Zhang discloses all of the elements of the current invention as stated above. Zhang further discloses further comprising: processing, by a first device, the first data representing first voice characteristics of the first speech to determine modified voice characteristics (Further discloses “extracting the features of the speech signal to obtain the target semantic features… and target phoneme features {...representing the first voice characteristics}” where the “features of the speech signal” are further voice characteristic which will be modified {the modified voice characteristics}, and the sampled data comprising the “semantic feature vector... and phonemic feature vector” are “compressed through the vector compression layer in the vector processing layer” with the “voiceprint feature vector” and “emotional feature vector” to “obtain the audio compression vector”; Zhang, ¶ [0089], [0091]). However, Zhang fails to expressly recite sending, from the first device to a second device, data representing the modified voice characteristics. The relevance of Minnen is described above with relation to claim 3. Regarding claim 5, Minnen teaches sending, from the first device to a second device, data representing the modified voice characteristics (“the compression and decompression systems may be co-located or remotely located, and compressed data generated by the compression system can be provided to the decompression system in any of a variety of ways. For example, the compressed data may be stored (e.g., in a physical data storage device or logical data storage area), and then subsequently retrieved from storage and provided to the decompression system. As another example, the compressed data may be transmitted over a communications network (e.g., the Internet) to a destination, where it is subsequently retrieved and provided to the decompression system.”; Minnen, ¶ [0060]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the flow-based voice conversion systems of Zhang to incorporate the teachings of Minnen to include sending, from the first device to a second device, data representing the modified voice characteristics. Zhang teaches systems and methods for flow-based voice conversion, which mathematically fuse target voice print features and phoneme features into a single latent vector. However, though Zhang considers remote implementation at a server, Zhang fails to expressly contemplate the details of the distribution of this process to devices and/or between devices across a network. In distributed speech processing, it is well known in the art that transmitting raw audio requires significant bandwidth, and transmitting uncompressed, high dimensional neural network vectors introduces unacceptable latency. Minnen specifically addresses the technical problem of transmitting high dimensional neural network data across networks, including a specialized architecture designed to generate a compressed representation from a latent representation, and transmit that compressed representation to a second device for further use and processing of the underlying data. A PHOSITA would be motivated to modify the flow-based voice conversion process of Zhang with the transmission architecture of Minnen to further enable the server embodiments described in Zhang , while avoiding known latency difficulties and increasing transmission efficiency across a distributed network, as recognized by Minnen . ( Minnen , ¶ [0046]-[0048]). Regarding claim 6, the rejection of claim 5 is incorporated. Zhang and Minnen disclose all of the elements of the current invention as stated above. Zhang further discloses wherein the data representing the modified voice characteristics comprises the third data (As previously indicated, the sampled data includes both “voiceprint feature vector, [and] emotional feature vector {third data}” and “semantic feature vector... and phonemic feature vector {data representing the modified third voice}” which are “compressed through the vector compression layer in the vector processing layer to obtain the audio compression vector.” As such, the data representing the modified voice characteristics comprises the third data.; Zhang, ¶ [0089], [0091]). Regarding claim 7, the rejection of claim 1 is incorporated. Zhang discloses all of the elements of the current invention as stated above. Zhang further discloses wherein: generation of the first compressed data uses a first machine learning model (The sampled data including “voiceprint feature vector” and “emotional feature vector” is “compressed through the vector compression layer in the vector processing layer to obtain the audio compression vector” where the vector compression layer is a machine learning model.; Zhang, ¶ [0089]). However, Zhang fails to expressly recite generation of the second audio data uses a second machine learning model corresponding to an inverse of the first machine learning model. The relevance of Minnen is described above with relation to claim 3. Regarding claim 7, Minnen teaches wherein: generation of the first compressed data uses a first machine learning model (“The system receives the [audio data] to be compressed” and “processes the data using an encoder neural network to generate a latent representation of the data (704). “; Minnen, ¶ [0098]-[0099]); and generation of the second audio data uses a second machine learning model (“The system determines the reconstruction of the data by processing the quantized latent representation of the data using the decoder neural network… to generate an (approximate or exact) reconstruction of the [audio data]”; Minnen, ¶ [0098], [0114]) corresponding to an inverse of the first machine learning model (“operations performed by the decoder neural network 204” can be an inverse of the operations performed by “the encoder neural network 106”, where network 106 performs an analysis transform to map the original data domain to a latent space and network 204 performs a synthesis transform to map the latent space back to the data domain. Further, with reference to specific examples, in FIG. 3, the encoder 302 is the inverse of decoder 304.; Minnen, ¶ [0066], [0086], [0091], FIG. 3). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the flow-based voice conversion systems of Zhang to incorporate the teachings of Minnen to include generation of the second audio data uses a second machine learning model corresponding to an inverse of the first machine learning model. Zhang teaches systems and methods for flow-based voice conversion, which mathematically fuse target voice print features and phoneme features into a single latent vector. However, though Zhang considers remote implementation at a server, Zhang fails to expressly contemplate the details of the distribution of this process to devices and/or between devices across a network. In distributed speech processing, it is well known in the art that transmitting raw audio requires significant bandwidth, and transmitting uncompressed, high dimensional neural network vectors introduces unacceptable latency. Minnen specifically addresses the technical problem of transmitting high dimensional neural network data across networks, including a specialized architecture designed to generate a compressed representation from a latent representation, and transmit that compressed representation to a second device for further use and processing of the underlying data. A PHOSITA would be motivated to modify the flow-based voice conversion process of Zhang with the transmission architecture of Minnen to further enable the server embodiments described in Zhang , while avoiding known latency difficulties and increasing transmission efficiency across a distributed network, as recognized by Minnen . ( Minnen , ¶ [0046]-[0048]). Regarding claim 8, the rejection of claim 1 is incorporated. Zhang discloses all of the elements of the current invention as stated above. However, Zhang fails to expressly recite further comprising: processing the first audio data to determine fourth data representing a latent representation of the first speech, wherein the first compressed data is based at least in part on the latent representation. The relevance of Minnen is described above with relation to claim 3. Regarding claim 8, Minnen teaches further comprising: processing the first audio data to determine fourth data representing a latent representation of the first speech, (“The encoder neural network 106 is configured to process the input data 102 (x)” which is audio data “to generate a latent representation 116 (y) of the input data 102.”; Minnen, ¶ [0059], [0066]) wherein the first compressed data is based at least in part on the latent representation (“To facilitate compression of the latent representation 116 of the input data using entropy encoding techniques, the compression system 100 quantizes the latent representation 116 of the input data using a quantizer Q 118 to generate an ordered collection of code symbols 120” and then “maps the input data to a quantized latent representation as an ordered collection of ‘code symbols’...[and] then generates the compressed representation of the input data based on: (i) the compressed code symbols, and (ii) ‘side-information’ characterizing the conditional entropy model used to compress the code symbols.”; Minnen, ¶ [0061], [0067]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the flow-based voice conversion systems of Zhang to incorporate the teachings of Minnen to include further comprising: processing the first audio data to determine fourth data representing a latent representation of the first speech, wherein the first compressed data is based at least in part on the latent representation. Zhang teaches systems and methods for flow-based voice conversion, which mathematically fuse target voice print features and phoneme features into a single latent vector. However, though Zhang considers remote implementation at a server, Zhang fails to expressly contemplate the details of the distribution of this process to devices and/or between devices across a network. In distributed speech processing, it is well known in the art that transmitting raw audio requires significant bandwidth, and transmitting uncompressed, high dimensional neural network vectors introduces unacceptable latency. Minnen specifically addresses the technical problem of transmitting high dimensional neural network data across networks, including a specialized architecture designed to generate a compressed representation from a latent representation, and transmit that compressed representation to a second device for further use and processing of the underlying data. A PHOSITA would be motivated to modify the flow-based voice conversion process of Zhang with the transmission architecture of Minnen to further enable the server embodiments described in Zhang , while avoiding known latency difficulties and increasing transmission efficiency across a distributed network, as recognized by Minnen . ( Minnen , ¶ [0046]-[0048]). Regarding claim 9, the rejection of claim 8 is incorporated. Zhang and Minnen disclose all of the elements of the current invention as stated above. However, Zhang fails to expressly recite wherein decompressing the first compressed data to determine the second data comprises decompressing the first compressed data to determine the latent representation. The relevance of Minnen is described above with relation to claim 3. Regarding claim 9, Minnen teaches wherein decompressing the first compressed data to determine the second data comprises decompressing the first compressed data to determine the latent representation (“The decompression system 200 processes the compressed data 104 generated by the compression system to generate a reconstruction 202 that approximates the original input data” including “using the entropy decoding engine 206 to entropy decode the compressed representation 126 of the quantized hyper-prior 124 that is included in the compressed data 104. In this example, the entropy decoding engine 206 may entropy decode the compressed representation 126 of the quantized hyper-prior 124 using the same (e.g., predetermined) entropy model that was used to entropy encode it,” which is performed to determine the latent representation, such that the original data can be reconstructed.; Minnen, ¶ [0080]-[0081]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the flow-based voice conversion systems of Zhang to incorporate the teachings of Minnen to include wherein decompressing the first compressed data to determine the second data comprises decompressing the first compressed data to determine the latent representation. Zhang teaches systems and methods for flow-based voice conversion, which mathematically fuse target voice print features and phoneme features into a single latent vector. However, though Zhang considers remote implementation at a server, Zhang fails to expressly contemplate the details of the distribution of this process to devices and/or between devices across a network. In distributed speech processing, it is well known in the art that transmitting raw audio requires significant bandwidth, and transmitting uncompressed, high dimensional neural network vectors introduces unacceptable latency. Minnen specifically addresses the technical problem of transmitting high dimensional neural network data across networks, including a specialized architecture designed to generate a compressed representation from a latent representation, and transmit that compressed representation to a second device for further use and processing of the underlying data. A PHOSITA would be motivated to modify the flow-based voice conversion process of Zhang with the transmission architecture of Minnen to further enable the server embodiments described in Zhang , while avoiding known latency difficulties and increasing transmission efficiency across a distributed network, as recognized by Minnen . ( Minnen , ¶ [0046]-[0048]). Regarding claim 13, the rejection of claim 11 is incorporated. Claim 13 is substantially the same as claim 3 and is therefore rejected under the same rationale as above. Regarding claim 14, the rejection of claim 13 is incorporated. Claim 14 is substantially the same as claim 4 and is therefore rejected under the same rationale as above. Regarding claim 15, the rejection of claim 11 is incorporated. Claim 15 is substantially the same as claim 5 and is therefore rejected under the same rationale as above. Regarding claim 16, the rejection of claim 15 is incorporated. Claim 16 is substantially the same as claim 6 and is therefore rejected under the same rationale as above. Regarding claim 17, the rejection of claim 11 is incorporated. Claim 17 is substantially the same as claim 7 and is therefore rejected under the same rationale as above. Regarding claim 18, the rejection of claim 11 is incorporated. Claim 18 is substantially the same as claim 8 and is therefore rejected under the same rationale as above. Regarding claim 19, the rejection of claim 18 is incorporated. Claim 19 is substantially the same as claim 9 and is therefore rejected under the same rationale as above . 07-22-aia AIA Claim s 10 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang as applied to claim 1 and 11 above, and further in view of Golman (U.S. Pat. App. Pub. No. 2023/0223011, hereinafter Golman ) . Regarding claim 10, the rejection of claim 1 is incorporated. Zhang discloses all of the elements of the current invention as stated above. However, Zhang fails to expressly recite wherein receiving the third data comprises: receiving an input to a user interface component, the input corresponding to at least one voice characteristic; and based at least in part on the input, determining the third data. Golman teaches “systems and methods for real-time correction of accent in speech audio signals.” ( Golman , ¶ [0002]). Regarding claim 10, Golman teaches wherein receiving the third data comprises: receiving an input to a user interface component (“Input devices 1208, in some examples, may be configured to receive input from a user through tactile, audio, video, or biometric channels” including any “device capable of detecting an input from a user... and relaying the input to the computer system 1200, or components thereof.”; Golman, ¶ [0110]), the input corresponding to at least one voice characteristic (“The user 102 may be provided with an option to select speaker embedding 808 form a list of pretrained speaker embeddings corresponding to different speakers” and the selected speaker embedding corresponds to “voice style {at least one voice characteristic}”; Golman, ¶ [0070]-[0071]); and based at least in part on the input, determining the third data (“The speaker embedding 808 of a target speaker and embeddings of the discretized values of energy 208 and pitch 206 (f0) are further added to the output of the encoder 802 to form input for the decoder 804.”; Golman, ¶ [0070]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the flow-based voice conversion systems of Zhang to incorporate the teachings of Golman to include wherein receiving the third data comprises: receiving an input to a user interface component, the input corresponding to at least one voice characteristic; and based at least in part on the input, determining the third data. Zhang teaches systems and methods for flow-based voice conversion, which mathematically fuse target voiceprint features and phoneme features into a single latent vector based on a preset. However, Zhang relies on preset voice characteristics to dictate the target voice, Zhang fails to expressly contemplate user interaction in selecting the one or more voice characteristics which are integrated into the target voice. Golman teaches an accent correction system including a user interface component which presents the user with the option to select a desired speaker embedding from a list of available speaker embeddings. A PHOSITA would be motivated to modify the static flow-based voice conversion process of Zhang with the speaker characteristic selection described in Minnen to transform Zhang’s preset voice customization into an interactive and customizable voice conversion application based on recognizing the limited functionality for preset voices, thus allowing a broader number of use cases for the system (e.g., obfuscating voice for safety, allowing third party markets for voice profiles, etc.) while maintaining flexibility to meet a wide variety of user needs, as recognized in light of the disclosure of Golman . ( Golman , ¶ [0063], [0071]). Regarding claim 20, the rejection of claim 11 is incorporated. Claim 20 is substantially the same as claim 10 and is therefore rejected under the same rationale as above . Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Zhang, Y. (U.S. Pat. App. Pub. No. 20200380952) discloses an end-to-end (E2E) text-to-speech (TS) model for multispeaker, multilingual speech synthesis. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sean E. Serraguard whose telephone number is (313)446-6627. The examiner can normally be reached 07:00-17:00 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C. Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Sean E Serraguard/Primary Examiner, Art Unit 2657 Application/Control Number: 18/887,462 Page 2 Art Unit: 2657 Application/Control Number: 18/887,462 Page 3 Art Unit: 2657 Application/Control Number: 18/887,462 Page 4 Art Unit: 2657 Application/Control Number: 18/887,462 Page 5 Art Unit: 2657 Application/Control Number: 18/887,462 Page 6 Art Unit: 2657 Application/Control Number: 18/887,462 Page 7 Art Unit: 2657 Application/Control Number: 18/887,462 Page 8 Art Unit: 2657 Application/Control Number: 18/887,462 Page 9 Art Unit: 2657 Application/Control Number: 18/887,462 Page 10 Art Unit: 2657 Application/Control Number: 18/887,462 Page 11 Art Unit: 2657 Application/Control Number: 18/887,462 Page 12 Art Unit: 2657 Application/Control Number: 18/887,462 Page 14 Art Unit: 2657 Application/Control Number: 18/887,462 Page 16 Art Unit: 2657 Application/Control Number: 18/887,462 Page 17 Art Unit: 2657 Application/Control Number: 18/887,462 Page 18 Art Unit: 2657 Application/Control Number: 18/887,462 Page 19 Art Unit: 2657 Application/Control Number: 18/887,462 Page 20 Art Unit: 2657 Application/Control Number: 18/887,462 Page 21 Art Unit: 2657 Application/Control Number: 18/887,462 Page 22 Art Unit: 2657 Application/Control Number: 18/887,462 Page 23 Art Unit: 2657 Application/Control Number: 18/887,462 Page 25 Art Unit: 2657
Read full office action

Prosecution Timeline

Sep 17, 2024
Application Filed
May 14, 2026
Non-Final Rejection mailed — §102, §103, §DOUBLEPATENT
Jul 14, 2026
Applicant Interview (Telephonic)
Jul 14, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750409
SIMULATED CHORAL AUDIO CHATTER
3y 11m to grant Granted Sep 29, 2026
Patent 12743582
DYNAMIC VOCABULARIES FOR CONDITIONING A LANGUAGE MODEL FOR TRANSFORMING NATURAL LANGUAGE TO A LOGICAL FORM
2y 8m to grant Granted Sep 22, 2026
Patent 12731158
COLLABORATIVE USER SUPPORT PORTAL
4y 6m to grant Granted Sep 08, 2026
Patent 12706081
SIMULATING CROWD NOISE FOR LIVE EVENTS THROUGH EMOTIONAL ANALYSIS OF DISTRIBUTED INPUTS
5y 2m to grant Granted Aug 11, 2026
Patent 12699834
SYSTEM AND METHOD FOR INTERACTIVE DIALOGUE
4y 5m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
99%
With Interview (+34.1%)
3y 0m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 162 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month