Prosecution Insights
Last updated: October 04, 2026
Application No. 18/460,642

COMMUNICATION SYSTEM FOR IMPROVING SPEECH

Non-Final OA §102§103
Filed
Sep 04, 2023
Priority
Sep 04, 2022 — provisional 63/374,553 +3 more
Examiner
SHAH, PARAS D
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Tomato.ai, Inc.
OA Round
3 (Non-Final)
73%
Grant Probability
Favorable
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
481 granted / 655 resolved
+11.4% vs TC avg
Strong +31% interview lift
Without
With
+31.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
33 currently pending
Career history
687
Total Applications
across all art units

Statute-Specific Performance

§101
18.5%
-21.5% vs TC avg
§103
46.8%
+6.8% vs TC avg
§102
13.6%
-26.4% vs TC avg
§112
10.5%
-29.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 655 resolved cases

Office Action

§102 §103
DETAILED ACTION This communication is in response to the RCE with Amendments and Arguments filed on 05/12/2026. Claims 1-2, 7, 10-13, 15-19, 24, 27-30, 32-34, and 123-132 are pending and have been examined. Any previous objection/rejection not mentioned in this Office Action has been withdrawn by the Examiner. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/12/2026 has been entered. Priority The Applicant claims priority to 4 different provisional applications. However, the examiner notes that the currently claimed subject matter is not taught in the specific details as recited by the current claim scope. As a result, these priority dates have not be considered and the earliest effective filing date considered for examination is 09/04/2023. If the Applicant is to dispute this finding, then a claim mapping providing specific support from the provisional applications would be required. Change of Examiner The Examiner of record has changed from Cameron Young to Paras Shah. Response to Amendments and Arguments With respect to the amendments filed on 05/12/2026, the Applicant restructured and added more features. Hence, the Applicant’s arguments are moot in view of new grounds for rejection. A new primary reference has been applied. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-2, 7, 18-19, 25, 123-125 are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Serebyrakov et al. (US 2022/0358903) (Note: The assignee of this reference is Sanas.AI and the current assignee of the instant application is Sanas.AI. However, the reference qualifies as prior art since Sanas.AI was not the original assignee at the time of filing (EFD) and was Tomato.AI). As to claim 1, Serebyrakov teaches a communication system comprising memory comprising instructions stored thereon and one or more processors (see [0017], processor) coupled to the memory and configured to execute the instructions to cause the communication system to (see [0019], where storage medium described which stores program instructions): train an offline training model of an accent conversion model on a first training structure and a second training structure (see [0033], where ASR engine includes machine learning models and trained using previously captured speech content from different speakers having same content and see [0035], where multiple speech accents can be learned), based on input from an input structure (see [0033], captured speech content from speakers), to develop inferences to generate a conversion relationship (see [0033]-[0035], learned linguistic representation developed), wherein the first training structure comprises first pairs of utterances and transcripts in a first accent (see [0033], where training data includes for the first accent transcribed content and the captured speech content), the second training structure comprises second pairs of utterances and transcripts in a second accent (see [0035], where captured speech using other accents such as a second accent is noted and is based on the same concept as the first accent), and the input structure comprises third pairs of utterances and transcripts in the first accent (see [0033], where speech captured from different speakers having the same first accent along with transcript); derive, via a speech reception unit, an input signal from a sound wave generated by a microphone to obtain input speech, wherein the input signal includes input language in the first accent (see [0032], where computing device may receive speech content having a first accent such as Indian English accent captured by microphone); modify the first accent of the input language in the input signal by applying a streaming speech-to-speech model of the accent conversion model to generate an output signal comprising output speech, wherein the streaming speech-to-speech model is configured to convert the first accent to the second accent based on the conversion relationship generated by the offline training model and comprises at least one neural network model configured to convert a first spectrogram of the first accent to a second spectrogram of the second accent (see [0042], where neural network uses the learned linguistic representations developed by ASR to learn how speech content maps from one accent to another and see [0045], where output speech generation engine converts into synthesized speech using a neural network and see [0043], mel spectrogram); and providing the output signal for audio output via a speech output unit, wherein the output signal comprises the output speech in the second accent and preserves speech content and one or more characteristics of the input speech (see [0044]-[0045], where output speech is provided which is a synthesized version of the received speech content having the second accent and where pitch, prosody, and emotions are retained). As to claims 18 and 123, apparatus claim 1 and 123 and method claim 18 are related as apparatus and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claims 18 and 123 are similarly rejected under the same rationale as applied above with respect to method claim. As to claim 123, Furthermore, Serebyrakov teaches a machine readable media comprising instructions that when executed cause the processor to (see [0019]) As to claims 2, 19, and 124, Serebyrakov teaches wherein the one or more characteristics comprise an emotion (see [0044], emotion), a prosody (see [0044]. Prosody), or a voice of the input speech (see [0044]-[0045], where voice type is not changed). As to claims 7, 25, and 125, Serebyrakov teaches wherein the streaming speech-to-speech model has a plurality of neural network models to convert the first spectrogram of the first accent to the second spectrogram of the second accent (see [0030], where VS engine 304 and vocoder 306 noted and see [0042], where VS 304 comprises a NN and see [0045], where vocoder 306 comprises a NN), wherein the neural network models have different parameters for execution and function in series (see [0042], [0045], each of the engine 304 and vocoder has a different purpose and therefore inherently makes use of different parameters, and where Figure 3, VC 304 and vocoder as the output speech generation engine 306) shows these two models connected in serial). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 10-11, 27-28, and 126-127 are rejected under 35 U.S.C. 103 as being unpatentable over Serebyrakov in view of Biadsy (US 2022/0122579). As to claims 10, 27, and 126, Serebyrakov teaches all of the limitations as in claims 1, 18, and 123. Serebyrakov does teach the use of a transformer based neural network (see [0033], where transformer neural networks comprise self-attention) However, Serebyrakov does not specifically teach wherein the at least one neural network model uses self-attention to model a sequence of input elements and a sequence of output elements by tracking relationships between pairs of the input elements. Biadsy does teach wherein the at least one neural network model uses self-attention to model a sequence of input elements and a sequence of output elements by tracking relationships between pairs of the input elements (see [0083], where decoder uses attention to learn from the prediction of the previous decoder time step). Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the attention as taught by Serebyrakov with the tracking of relationships as taught by Biadsy in order to aide in learning attention which will help create a prediction of the output signal as described (see Biadsy [0083]-[0084]). As to claims 11, 28, and 127, Serebyrakov teaches all of the limitations as in claims 10, 27, and 126. Furthermore, Biadsy teaches wherein, for each output element, the self-attention uses a subset of past input elements and a subset of future input elements (see [0083], where previous decoder timestep information is used to help learn the attention and is used to predict the output spectrogram where the input is described in [0079] as input speech signal). Claim(s) 12-13, 29-30, and 128-129 are rejected under 35 U.S.C. 103 as being unpatentable over Serebyrakov in view of Bianco (US 2015/0100315) As to claims 12, 29, and 128, Serebyrakov teaches all of the limitations as in claims 1, 18, and 123. However, Serebyrakov does not specifically teach a relay server device positioned between first and second stacks of relays in a telephone system and hosting the accent conversion model. Bianco does teach a relay server device positioned between first and second stacks of relays in a telephone system and hosting the accent conversion model (see [0083]-[0085], where a telephony system comprising relays and codecs wherein there are multiple codecs and relays connected together to transmit IP telephony information. As such, a person of ordinary skill in the art would have understood that server-based accent modification systems such as that of Serebyrakov would be placed between the relays and codecs in order to operate on the audio data received from the codecs and relays. As Serebyrakov is related to accent modification as well as Bianco which in para [0097] notes the usage of accent information for generating audio (audio generation unit 406) and making use of codecs and relays. Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the transmission as taught by Serebyrakov with the relays as taught by Bianco in order help correct losses of data packets being transmitted (see Bianco [0083] – [0085]). As to claims 13, 30, and 129, Serebyrakov in view of Bianco teach all of the limitations as in claims 12, 29, and 128. Furthermore, Bianco teaches wherein the relay server device includes first and second codecs that connect the accent conversion device model to the first and second stacks of relays respectively (see [0083]-[0085], where the relay server device includes first and second codecs that connect the accent conversion device to the first and second stacks of relays respectively. (Bianco teaches a telephony system comprising relays and codecs wherein there are multiple codecs and relays connected together to transmit IP telephony information.) As such, a person of ordinary skill in the art would have understood that server-based accent modification systems such as that of Serebyrakov would be placed between the relays and codecs in order to operate on the audio data received from the codecs and relays. As Serebyrakov is related to accent modification as well as Bianco which in para [0097[ notes the usage of accent information for generating audio (audio generation unit 406) and making use of codecs and relays. Claim(s) 15, 32, and 130 are rejected under 35 U.S.C. 103 as being unpatentable over Serebyrakov in view of Richter (US 11756527). As to claims 15, 32, and 130, Serebyrakov teaches wherein the one or more processors are further configured to execute the instructions to cause the communication system to modify the input language based on at least a first knowledge base (see [0042], accent conversion process is described using learned representations) (e.g. It is known that in order for learned representations to be used they must be stored). However, per the specification and in the interest of compact prosecution, the “modifying” is a removal/correction of speech prior to outputting. Richter does teach modify the input language based on at least a first knowledge base (see col. 14, lines 24-31, where certain phrases are detected such as those related to profanity or offensive phrases, where it is known that these types of phrases or words need to be stored in order to detect, and then are further removed from the speech). Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the input audio as taught by Serebyrakov with the modifying as taught by Richter in order to remove certain phrases that should not be presented to the other individual due to their relationship (see Richter col. 14, lines 24-31). Claim(s)16, 33, and 131 are rejected under 35 U.S.C. 103 as being unpatentable over Serebyrakov in view of Hijazi (US 2022/0392478). As to claims 16, 33, and 131, Serebyrakov teaches all of the limitations as in claims 1, 18, and 123. However, Serebyrakov does not specifically teach wherein the one or more processors are further configured to execute the instructions to cause the communication system to detect an overlap of the input speech from first and second input signals and to suppress the suppressing second speech in the input speech in favor of not suppressing first speech in the input speech. Hijazi does teach wherein the one or more processors are further configured to execute the instructions to cause the communication system to detect an overlap of the input speech from first and second input signals (see [0020]-[0025], where secondary speakers are detected apparat from the single talker) and to suppress the suppressing second speech in the input speech in favor of not suppressing first speech in the input speech (see [0020] and see [0030], where primary talker is determined and other speech is suppressed). Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the audio capturing as taught by Serebyrakov with the suppression as taught by Hijazi in order to provide a useful mechanism of when to suppress background talkers vs an intended participant (see Hijzai [0020]). Claim(s) 17, 34, and 132 are rejected under 35 U.S.C. 103 as being unpatentable over Serebyrakov in view of De Carney (US 2016/0065711). As to claims 17, 34, and 132, Serebyrakov teaches all of the limitations as in claims 1, 18, and 123. However, Serebyrakov does not specifically teach wherein the one or more processors are further configured to execute the instructions to cause the communication system to determine that whether a gap between time segments in the input speech requires an injected utterance and merging an utterance with the time segments so that the utterance is between corresponding time segments in the output speech, wherein the utterance is generated or selected based on a determined context. De Carney does teach wherein the one or more processors are further configured to execute the instructions to cause the communication system to determine that whether a gap between time segments in the input speech requires an injected utterance and merging an utterance with the time segments so that the utterance is between corresponding time segments in the output speech, wherein the utterance is generated or selected based on a determined context (see [0035], [0038]-[0040], where when the user is not talking such as periods of silence, a speech message is injected into the conversation into the voice conversation which is in context of the conversation since it’s a response to the question posed by the caller (see [0036])). Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the audio capturing as taught by Serebyrakov with the injection of speech as taught by De Carney in order to allow the transmittal of important information to the second party of the conversation (e.g., the information that one of the speakers has chosen to use a text-to-speech interface for the call) (see De Carney [0008] – [0013]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Nguyen (US 2024/0404505) is cited to disclose synthesizing multi accent speech using weights. Erdmann (US 2024/0304175) is cited to disclose speech modification using accent embeddings. Ingel (US 2021/0224319) is cited to disclose revoicing/dubbing audio into another accent/dialect. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PARAS D SHAH whose telephone number is (571)270-1650. The examiner can normally be reached Monday-Thursday 7:30AM-2:30PM, 5PM-7PM (EST), Friday 8AM-noon (EST). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Paras D Shah/ Supervisory Patent Examiner, Art Unit 2653 09/11/2026
Read full office action

Prosecution Timeline

Sep 04, 2023
Application Filed
Jul 01, 2025
Non-Final Rejection mailed — §102, §103
Oct 01, 2025
Response Filed
Jan 13, 2026
Final Rejection mailed — §102, §103
May 12, 2026
Request for Continued Examination
May 13, 2026
Response after Non-Final Action
Sep 15, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748922
DERIVING HETEROGENOUS CONTEXT FOR EVALUATING GENERATED CODE
2y 6m to grant Granted Sep 29, 2026
Patent 12749497
INFORMATION PROCESSING SYSTEM THAT ALLOWS USER TO ESTABLISH CONVERSATION WITH DESIRED PERSON THROUGH AVATAR EVEN IN NOISY ENVIRONMENT IN VIRTUAL SPACE, EDGE DEVICE, SERVER, CONTROL METHOD, AND STORAGE MEDIUM
2y 0m to grant Granted Sep 29, 2026
Patent 12737546
SYSTEM AND METHOD FOR K-NUGGET DISCOVERY AND RETROFITTING FRAMEWORK
2y 6m to grant Granted Sep 15, 2026
Patent 12738272
SYSTEM FOR DISTRIBUTION OF CONTENT AND ANALYSIS OF CONTENT ENGAGEMENT
2y 2m to grant Granted Sep 15, 2026
Patent 12730983
RELATION EXTRACTION SYSTEM AND METHOD ADAPTED TO FINANCIAL ENTITIES AND FUSED WITH PRIOR KNOWLEDGE
3y 2m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
73%
Grant Probability
99%
With Interview (+31.0%)
3y 9m (~8m remaining)
Median Time to Grant
High
PTA Risk
Based on 655 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month