Prosecution Insights
Last updated: October 04, 2026
Application No. 18/492,635

PROBABILISTIC MULTI-PARTY AUDIO TRANSLATION

Non-Final OA §103
Filed
Oct 23, 2023
Priority
Dec 28, 2022 — provisional 63/435,701
Examiner
MEIS, JON CHRISTOPHER
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Mass Luminosity Inc.
OA Round
3 (Non-Final)
33%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants only 33% of cases
33%
Career Allowance Rate
11 granted / 33 resolved
-28.7% vs TC avg
Strong +52% interview lift
Without
With
+52.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
17 currently pending
Career history
60
Total Applications
across all art units

Statute-Specific Performance

§101
21.7%
-18.3% vs TC avg
§103
55.9%
+15.9% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
9.2%
-30.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 33 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1-20 are pending. Claims 1, 11, and 20 are independent. This Application was published as US 20240220737. Apparent priority is 28 December 2022. The instant Application is directed to a method of streaming translation. Response to Arguments 35 USC 103 Applicant’s arguments with respect to 35 USC 103 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-3, 5-6, 8-13, 15-16, and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. ("Opportunistic Decoding with Timely Correction for Simultaneous Translation") in view of Kim et al. (US 20110153309 A1) and Kim et al. (US 20210082407 A1), hereinafter as “Kim2”. PNG media_image1.png 222 827 media_image1.png Greyscale Zheng, Fig. 2 Regarding claim 1, Zheng discloses: 1. A method, comprising: receiving input text of a communication session; (Fig. 2 shows input text at the top of the figure. Section 1 mentions "international conferences" which reads on a communication session. ) generating translation data by processing the input text with a prediction model (Fig. 2 shows that two words are predicted at each step. See also "As shown in Fig. 1, our proposed method always decodes more words than the original policy at each step to catch up with the speaker and reduce the latency." bottom of pg. 1 to top of pg. 2. See also pg. 2 section "Correction with Beam Search." Beam search reads on a prediction model. ) and a translation model; (Fig. 2 shows the translated words in English. See pg. 4, section 5 - "Datasets and Implementations" for the translation model details. ) processing the translation data and enunciation data with a sentence similarity model to generate a similarity score; ("At step t + 1, when encoder obtains more information from x6g(t) to x6g(t+1), the decoder is capable to generate more appropriate candidates and may revise and replace the previous outputs from opportunistic decoding." pg. 2, Section 3 - "Timely Correction" – Zheng discloses processing the translation and enunciation data so see if a more appropriate candidate is available. Zheng does not explicitly disclose sentence similarity or a similarity score.) and presenting the enunciation data based on the similarity score (Fig. 2 shows presentation of enunciation data. See also "When there is a disagreement, our model always uses the hypothesis from later step to replace the previous commits." ) and based on rejecting a first translation variant from a plurality of translation variants of the translation data for exceeding a time threshold. (not explicitly disclosed by Zheng) Zheng does not explicitly disclose sentence similarity or a similarity score, or presenting the enunciation data based on the similarity score or a based on rejecting a translation variant of a plurality of variants for exceeding a time threshold. Kim discloses: processing the translation data and enunciation data with a sentence similarity model to generate a similarity score; ("[0018] The similarity calculating unit 120 considers the confidence score for each word processed by the voice recognizing unit 100 and compares the various elements extracted by the language processing unit 110 with various elements stored in the translated sentence DB 150 to calculate the similarity therebetween. ..." ) and presenting the enunciation data based on the similarity score ("[0019] The similarity calculation result by Equation (1) is expressed in the form of probability. A threshold value is set and it is determined whether the calculated similarity is higher than the threshold value. If the calculated similarity is higher than the threshold value, class information of the second-language sentence corresponding to the first-language sentence selected from the translated sentence DB 150 is translated and the translated result is transferred to the voice synthesizing unit 140 without passing through the sentence translating unit 130. On the other hand, if the calculated similarity is lower than the threshold value, user selection is requested or the first-language sentence (i.e., the voice recognition result) is transferred to the sentence translating unit 130. ..." ) Zheng and Kim are considered analogous art to the claimed invention because they disclose methods for translation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Zheng with the sentence similarity calculating unit of Kim in order to use the originally generated sentence if it meets a threshold. Doing so would have been beneficial so that the audience is not overwhelmed by modifications. (Zheng pg. 2, para 1.) Kim does not explicitly disclose presenting data based on rejecting a translation variant of a plurality of variants for exceeding a time threshold. Kim2 discloses: presenting enunciation data based on rejecting a first translation variant from a plurality of translation variants of the translation data for exceeding a time threshold. (“[0076] In this example, when the utterance length of the generated candidate translation result exceeds the delay time, the processor 200 may replace the generated candidate translation result with another candidate translation result. Through the replacing, the processor 200 may reduce the utterance time of the translation result.”.) Zheng, Kim, and Kim2 are considered analogous art to the claimed invention because they disclose methods for translation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Zheng in view of Kim with the time delay threshold and candidate replacement disclosed by Kim2. Doing so would have been beneficial to keep audio and video in sync and reduce a time delay (Kim2 [0053] and [0057]). Regarding claim 2, Zheng and Kim do not disclose the additional limitations. Kim2 discloses: 2. The method of claim 1, further comprising: presenting the enunciation data using the time threshold identifying when the enunciation data for the input text is to be presented. ("[0107] The delay time controller 231 may expand the silence interval received from the storage device 210 to be a spare time, and apply the spare time to the delay time. Additionally, the delay time controller 231 may control a length of a translation result generated by the machine translator 235, and control the speech synthesizer 237 to manage an utterance time for each token.”) See claim 1 for motivation statement Regarding claim 3, Zheng and Kim do not disclose the additional limitations. Kim2 discloses: 3. The method of claim 1, further comprising: determining a playback rate using the time threshold and a time value of the input text; and presenting the enunciation data using the playback rate. ("[0132] In operation 555, the real-time translation engine 230 compares an utterance length of the rephrased candidate translation result to the delay time. In operation 556, when the utterance length of the rephrased candidate translation result is greater than the delay time, the real-time translation engine 230 adjusts a ratio of an utterance speed of a TTS synthesis.”) Zheng, Kim, and Kim2 are considered analogous art to the claimed invention because they disclose methods for translation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Zheng in view of Kim to adjust the utterance speed of TTS synthesis as disclosed by Kim2. Doing so would have been beneficial to keep audio and video in sync and reduce a time delay (Kim2 [0053] and [0057]). Regarding claim 5, Zheng discloses: 5. The method of claim 1, further comprising: processing the input text with the translation model; and processing output from the translation model with the prediction model to generate the translation data. ("When the opportunistic decoding window is w at decoding step t, we define the beam search over w + 1 (include the original output) as follows: ... where nextb n+w(·) performs a beam search with n + w steps, and generate y 0 t as the outputs which include both original and opportunistic decoded words. n represents the length of yt" pg. 3, Section 3 "Correction with Beam Search" – the beam search uses the output translation to predict next words.) Regarding claim 6, Zheng does not explicitly disclose the additional limitations. Kim discloses: 6. The method of claim 1, further comprising: presenting the enunciation data as synthesized audio in an audio stream of a live media stream of the communication session. ("[0021] The voice synthesizing unit 140 receives the second-language sentence from the similarity calculating unit 120 or the second-language sentence from the sentence translating unit 130, synthesizes the prestored voice data mapping to the received second-language sentence, and outputs the synthesized voice data in the form of analog signals." ) Zheng and Kim are considered analogous art to the claimed invention because they disclose methods for translation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified the method of Zheng in view of Kim with the voice synthesizing unit taught by Kim. Doing so would have been beneficial so that the user could hear the translated audio. Regarding claim 8, Zheng discloses: 8. The method of claim 1, further comprising: adjusting the enunciation data with the translation data when the similarity score satisfies a similarity threshold. ("At the same time, it also employs a timely correction mechanism to review the extra outputs from previous steps with more source context, and revises these outputs with current preference when there is a disagreement." pg. 2, para 1) Zheng discloses that the enunciation data is always replaced with the translation if there is a disagreement. Zheng does not disclose adjusting it based on the similarity score satisfying a threshold. Kim discloses: 8. The method of claim 1, further comprising: adjusting the enunciation data with the translation data when the similarity score satisfies a similarity threshold. ("[0019] The similarity calculation result by Equation (1) is expressed in the form of probability. A threshold value is set and it is determined whether the calculated similarity is higher than the threshold value. If the calculated similarity is higher than the threshold value, class information of the second-language sentence corresponding to the first-language sentence selected from the translated sentence DB 150 is translated and the translated result is transferred to the voice synthesizing unit 140 without passing through the sentence translating unit 130. On the other hand, if the calculated similarity is lower than the threshold value, user selection is requested or the first-language sentence (i.e., the voice recognition result) is transferred to the sentence translating unit 130. ..." ) See claim 1 for motivation statement. Regarding claim 9, Zheng discloses: 9. The method of claim 1, further comprising: presenting a correction from the enunciation data after adjustment of the enunciation data when the similarity score satisfies a similarity threshold. ("At the same time, it also employs a timely correction mechanism to review the extra outputs from previous steps with more source context, and revises these outputs with current preference when there is a disagreement." pg. 2, para 1) Zheng does not disclose: the similarity score satisfies a threshold. Kim discloses: 9. The method of claim 1, further comprising: presenting a correction from the enunciation data after adjustment of the enunciation data when the similarity score satisfies a similarity threshold. ("[0019] The similarity calculation result by Equation (1) is expressed in the form of probability. A threshold value is set and it is determined whether the calculated similarity is higher than the threshold value. If the calculated similarity is higher than the threshold value, class information of the second-language sentence corresponding to the first-language sentence selected from the translated sentence DB 150 is translated and the translated result is transferred to the voice synthesizing unit 140 without passing through the sentence translating unit 130. On the other hand, if the calculated similarity is lower than the threshold value, user selection is requested or the first-language sentence (i.e., the voice recognition result) is transferred to the sentence translating unit 130. ..." ) See claim 1 for motivation statement. Regarding claim 10, Zheng discloses: 10. The method of claim 1, further comprising: presenting a correction from the enunciation data after adjustment of the enunciation data and adjustment of a playback rate. (" At the same time, it also employs a timely correction mechanism to review the extra outputs from previous steps with more source context, and revises these outputs with current preference when there is a disagreement." pg. 2, para 1 – Zheng discloses a simultaneous translation system; therefore any correction would continue to be presented after the playback rate has been adjusted as detailed in claim 3.) Claim 11 is a system claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally, “at least one processor; and an application” of the Claim are taught by Zheng (“Transformer based wait-k model” pg. 4, section 5 implies use of a processor). Claim 12 is a system claim with limitations corresponding to the limitations of Claim 2 and is rejected under similar rationale. Claim 13 is a system claim with limitations corresponding to the limitations of Claim 3 and is rejected under similar rationale. Claim 15 is a system claim with limitations corresponding to the limitations of Claim 5 and is rejected under similar rationale. Claim 16 is a system claim with limitations corresponding to the limitations of Claim 6 and is rejected under similar rationale. Claim 18 is a system claim with limitations corresponding to the limitations of Claim 8 and is rejected under similar rationale. Claim 19 is a system claim with limitations corresponding to the limitations of Claim 9 and is rejected under similar rationale. Claim 20 is a computer readable medium claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally, “A non-transitory computer readable medium comprising instructions executable by a computer processor” of the Claim are taught by Zheng (“Transformer based wait-k model” pg. 4, section 5 implies use of a computer which has memory). Claim(s) 4 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zheng in view of Kim and Kim2 as applied in claim 1 above, in further view of Hamid et al. (US 20120253785 A1). Regarding claim 4, Zheng discloses: 4. The method of claim 1, further comprising: processing the input text with the prediction model; (see claim 1) and processing output from the prediction model with the translation model to generate the translation data. (not explicitly disclosed) Zheng and Kim do not disclose a prediction model which outputs to the translation model. Kim2 broadly discloses this approach as a known method (see Kim2 [0054]) but does not provide explicit detail of how it is performed. Hamid discloses: processing the input text with the prediction model; and processing output from the prediction model with the translation model to generate the translation data. ("[0033] … An example language model may be generated based on determining the probability that a sequence would occur in a natural language conversation or natural language document based on observing a large amount of text in the language. Thus, the language model may statistically predict a "next" word or phrase in a sequence of words associated with a natural language that forms the basis of the language model. The language model probability value may be based on information included in a language model repository 144, which may be configured to store information obtained based on a large corpus of documents by analyzing contexts of words in documents that have been translated from a source language to a target language. Based on such analyses, the language model probability value may predict a particular word or phrase that would be expected next in a sequence of object words in a source language. According to an example embodiment, a ranked listing of suggested translations may be obtained." ) Zheng, Kim, Kim2, and Hamid are considered analogous art to the claimed invention because they disclose methods for translation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Zheng in view of Kim and Kim2 with the language model disclosed by Hamid. Doing so would have been beneficial in order to use the information included in a large parallel corpus of documents to provide the highest ranking candidate translation. (Hamid [0034]) Claim 14 is a system claim with limitations corresponding to the limitations of Claim 4 and is rejected under similar rationale. Claim(s) 7 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zheng in view of Kim and Kim2 as applied in claim 1 above, in further view of Potkonjak (US 20100324894 A1). Regarding claim 7, Zheng, Kim, and Kim2 do not disclose the additional limitations. Potkonjak discloses: 7. The method of claim 1, further comprising: presenting the enunciation data as subtitle text in a video stream of a live media stream of the communication session. ("[0029] The constraints 192 and objective functions 194 can be specified for use in capturing lectures or other presentations using single or multiple distributed microphones. The constraints 192 and objective functions 194 can be specified for use in the operation of call centers where one or more of the processing stages 120-170 may be applied to voice signals generated by call center personnel. The constraints 192 and objective functions 194 can be specified for use in the operation of call centers where one or more of the processing stages 120-170 may be applied to voice signals generated by call center customers. Text generated by the V2T processing stage 130 or the T2T processing stage 140 may be displayed to a speaker or a listener. Such a display of text may support closed caption applications or other services for captions, subtitles, or the hearing impaired." ) Zheng, Kim, Kim2 and Potkonjak are considered analogous art to the claimed invention because they disclose methods for translation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Zheng in view of Kim and Kim2 to output subtitles as disclosed by Potkonjak. Doing so would have been beneficial so that hearing impaired users could read the output. (Potkonjak [0029]) Claim 17 is a system claim with limitations corresponding to the limitations of Claim 7 and is rejected under similar rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JON C MEIS whose telephone number is (703)756-1566. The examiner can normally be reached Monday - Thursday, 8:30 am - 5:30 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JON CHRISTOPHER MEIS/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Show 1 earlier event
Sep 22, 2025
Non-Final Rejection mailed — §103
Dec 19, 2025
Examiner Interview Summary
Dec 19, 2025
Applicant Interview (Telephonic)
Dec 22, 2025
Response Filed
Mar 24, 2026
Final Rejection mailed — §103
Jul 02, 2026
Request for Continued Examination
Jul 06, 2026
Response after Non-Final Action
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711961
SYSTEM AND METHOD FOR DIGITAL VOICE DATA PROCESSING AND AUTHENTICATION
3y 4m to grant Granted Aug 18, 2026
Patent 12603087
VOICE RECOGNITION USING ACCELEROMETERS FOR SENSING BONE CONDUCTION
3y 8m to grant Granted Apr 14, 2026
Patent 12579975
Detecting Unintended Memorization in Language-Model-Fused ASR Systems
2y 11m to grant Granted Mar 17, 2026
Patent 12482487
MULTI-SCALE SPEAKER DIARIZATION FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
3y 0m to grant Granted Nov 25, 2025
Patent 12475312
FOREIGN LANGUAGE PHRASES LEARNING SYSTEM BASED ON BASIC SENTENCE PATTERN UNIT DECOMPOSITION
2y 9m to grant Granted Nov 18, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
33%
Grant Probability
86%
With Interview (+52.4%)
2y 10m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 33 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month