Prosecution Insights
Last updated: October 02, 2026
Application No. 18/337,984

TEXT NORMALIZATION AND INVERSE TEXT NORMALIZATION FOR MULTI-LINGUAL LANGUAGE MODELS

Non-Final OA §103
Filed
Jun 20, 2023
Examiner
CAUDLE, PENNY LOUISE
Art Unit
2657
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
70%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
85%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
59 granted / 84 resolved
+8.2% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
17 currently pending
Career history
97
Total Applications
across all art units

Statute-Specific Performance

§101
21.3%
-18.7% vs TC avg
§103
47.1%
+7.1% vs TC avg
§102
15.0%
-25.0% vs TC avg
§112
15.9%
-24.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 84 resolved cases

Office Action

§103
DETAILED ACTION This examination is in response to the communication filed on 05/04/2026. Claims 1-20 are currently pending, wherein claims 1, 8, 16 and 17 have been amended. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment/Argument Applicant’s amendments and arguments, filed 05/04/2026, with respect to the rejections of claims 1-20 under §103 have been fully considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-5 and 7-15 are rejected under 35 U.S.C. 103 as being unpatentable over by Beliga et al. "Text normalization for Croatian speech synthesis," 2011 Proceedings of the 34th International Convention MIPRO, Opatija, Croatia, 2011, pp. 1664-1669 (herein “Beliga”) in view of Zhang et al. “NeMo Inverse Text Normalization: From Development to Production” arXiv:2104.05055v2 [cs.CL] 17 May 2021 (herein “Zhang”), further in view of Huang (US 2003/0202641 A1; herein “Huang”). Regarding claim 1, Beliga teaches a method, comprising: obtaining a textual input corresponding to one or more semiotic classes (p. 1665, Fig. 1 input text; p. 1667 TABLE 1, step 1 “Read the input text…” and p. 1668, Section V teaches “the performance…was tested on the input text which contains typical Croatian NSW classes…”); determining, based at least in part on the one or more semiotic classes, a set of tokens for the textual input (p. 1665, Figure 1, “Tokenizer” and section B. teaches “The text pre-processing module initially has to identify the NSW, and separate it from standard words, as shown in Figure 1” the tokenizer of Figure 1 inherently determines a set of tokens for the input text); determining, for individual tokens of the set of tokens, a classification (P. 1665, Figure 1 “NSW classification” and Section B. teaches “The text pre-processing module initially has to identify the NSW, and separate it from standard words as shown in Figure 1” As shown in Fig. 1, the output of the tokenizer is input to the NSW classification, accordingly, the classification is determined for individual tokens); determining, using one or more first rule-based algorithms, respective plain text representations for the individual tokens (p. 1665, Figure 1, “NORMALIZED FORM”; p. 1666, first column, teaches “The module initially classifies NSW as letters, numerals or combination as shown in the main classification tree in Figure 2…With suggested classification it is possible to retrieve algorithms that make the normalization more achievable”; p. 1667, Table 1, steps 2b-2c. teaches “If it is a NEW, find its class in the classification tree…The algorithms for normalization of subclass NUMBER are applied…The algorithms for normalization of subclass LETTER are applied…The algorithms for normalization of subclass COMBINED are applied…Write the expanded form of the NSW obtained in the process of normalization to the output” ); determining, using one or more second rule-based algorithms , a combined plain text representation based at least on the respective plain text representations for each token generating, based at least on the combined plain text representation, an auditory representation corresponding to the textual input (p. 1667 TABLE 1, step 4 teaches “the normalized text is sent to the module for grapheme-to-phoneme conversion” the phoneme conversation is interpreted as an auditory representation of the textual input). Beliga fails to disclose determining a language for the textual input or that the one or more second rule-based algorithms are selected based, at least in part, on the language of the textual input and the classification of the individual tokens into the one or more semiotic classes. Zhang teaches a rule-based Inverse Text Normalization framework which uses a two stage normalization pipeline that first detects semiotic tokens (classification) and then converts these to written form (verbalization). More specifically, Zhang teaches “This framework is general enough to be applied to text normalization and other languages. In fact, Kestrel [4] and Sparrowhawk [8] were originally intended for text normalization across a variety of languages. To do this, we replace the grammars for all classes. The classification grammars will be mostly language-specific” (Zhang, p.3, section 3.6). Thus, Zhang teaches that the grammars, i.e., rules, utilized are based on both the semiotic classification and the specific-language. Beliga differs from the claimed invention, as defined by claim 1, in that Beliga fails to explicitly disclose the rule-based algorithms are selected based on semiotic class and language of the input as claimed. Text normalization and inverse text normalization system which the rule-based algorithms are selected based on semiotic class and are language specific are known in the art as evidenced by Zhang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system of Beliga to include the text normalization framework taught by Zhang as it merely constitutes the substitution/combination of known processes to achieve the predication result of converting non-standard words. The combination of Beliga and Zhang fails to explicitly disclose determining a language for the textual input. Although the combination teaches utilizing language-specific algorithms for text normalization and inverse text normalization, it fails to explicitly disclose determining the input language. Huang teaches TTS system architectures which function as a synthesizer for multiple languages where the language-specific information, e.g., special rules for linguistic analysis are loaded by the TTS engine at run-time so that it is possible to switch voices and languages as desired at run-time (Huang, ¶[0066]). The combination of Beliga and Zhang differs from the claimed invention, as defined by claim 1, in that combination fails to explicitly disclose determining a corresponding language and selecting the rule-based algorithms based on the determined language as claimed. TTS systems which determine/select the language specific information/rules needed to generate synthesized speech in the desired language are known in the art as evidenced by Huang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system taught by the combination of Beliga and Zhang to detect/determine a corresponding input language and selected the rule-based algorithms based on the determined language in order to make it possible to switch voices and languages of the TTS system at run-time (Huang, ¶[0066].). Regarding claim 3, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 1. In addition, Beliga further teaches the one or more semiotic classes include at least a first level class and one or more sub-classes (p. 1666 Figure 4 and p. 1667 Table 1, Step 2, subclass determination). Regarding claim 4, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 1. In addition, Beliga further teaches providing, to a trained vocalizer, the combined plain text representation (p. 1667, Table 1, step 4 “The normalized text is sent to the module for grapheme-to-phoneme conversation” and p. 1669, first column teaches “The proposed normalization can easily be integrated with existing grapheme-to-phoneme conversation system [14] and speech generation modules of TTS synthesis system [15]…” the speech generation modules of a TTS synthesis inherently includes a trained vocalizer or vocoder). Regarding claim 5, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 1. In addition, Beliga further teaches the one or more first rule-based algorithms are the same as the one or more second rule-based algorithms (p. 1667, TABLE 1, step 3 teaches “Repeat the procedure iteratively until the input text completely normalized” the iteration of the classification steps inherently includes utilization of the same rule-based algorithms as utilized in the previous classification step). Regarding claim 7, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 1(See detailed element mapping above). In addition, Huang further teaches determining, for the auditory output, a desired language (¶[0066] teaches “Some language-specific information is necessary; there are acoustic inventories unique to each language and there are also special rules for linguistic analysis. These data, however, are stored externally in tables and parameter files, and are loaded by the TIS engine at run-time. Thus, in applications such as dialog or e-mail reading, it is possible to switch voices and languages as desired at run-time”); and selecting, based at least on the desired language, the one or more rule-based algorithms ([0066] teaches “Some language-specific information is necessary; there are acoustic inventories unique to each language and there are also special rules for linguistic analysis. These data, however, are stored externally in tables and parameter files, and are loaded by the TIS engine at run-time. Thus, in applications such as dialog or e-mail reading, it is possible to switch voices and languages as desired at run-time”) The combination of Beliga and Zhang differs from the claimed invention, as defined by claim 7, in that the combination fails to explicitly disclose determining a desired language for the auditory output and selecting, based at least one the desired language, one or more rule-based algorithms as claimed. TTS systems which determine/select the language specific information/rules needed to generate synthesized speech in the desired language are known in the art as evidenced by Huang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system taught by the combination of Beliga and Zhang to detect/determine a desired output language and selected the rule-based algorithms based on the desired language in order to make it possible to switch voices and languages of the TTS system at run-time (Huang, ¶[0066].). Regarding claim 8, Beliga teaches a system comprising: at least one processor (Beliga teaches the text normalization process as an integral part of a text-to-speech (TTS) synthesis system such a system inherently requires at least one processor) to: determine one or more tokens for segments of an input (p. 1665, Figure 1, “Tokenizer” ); classify, using a rule-based grammar model for the language, the one or more tokens into a semiotic class (p. 1665, Figure 1, NSW classification and Section B teaches “The text pre-processing module initially has to identify the NSW, and separate it from standard words…Normalization module then classifies NSW according to the given taxonomy of Croatian language…” ); select a class rule-based grammar model, from one or more rule-based grammar models corresponding to the language (p. 1665, Figure 1, NSW classification and Section B teaches “The text pre-processing module initially has to identify the NSW, and separate it from standard words…Normalization module then classifies NSW according to the given taxonomy of Croatian language…Within each NSW class, certain rules are typical so the normalization can be carried out in a more standard way” and p. 1667, Table 1, step 2); generate, using the class rule-based grammar model, a textual output for the token (p. 1667, Table 1, step 2c. “Write the expanded form of the NSW obtained in the process of normalization to the output” the normalized tokens are interpreted as a textual output); and generate, a combined textual output including respective textual outputs for each token of the one or more tokens (Fig. 1, “List of words” and p. 1665, Section B teaches “The module creates a list of words from the normalized text and it passes them forward to the module for grapheme-to-phoneme conversion”). Beliga fails to disclose determining a language for the input or that the rule-based grammar model is selected based, at least in part, on both the language and a respective semiotic class for a token of the one or more token. Zhang teaches a rule-based Inverse Text Normalization framework which uses a two stage normalization pipeline that first detects semiotic tokens (classification) and then converts these to written form (verbalization). More specifically, Zhang teaches “This framework is general enough to be applied to text normalization and other languages. In fact, Kestrel [4] and Sparrowhawk [8] were originally intended for text normalization across a variety of languages. To do this, we replace the grammars for all classes. The classification grammars will be mostly language-specific” (Zhang, p.3, section 3.6). Thus, Zhang teaches that the grammars, i.e., rules, utilized are based on both the semiotic classification and the specific-language. Beliga differs from the claimed invention, as defined by claim 8, in that Beliga fails to explicitly disclose the rule-based algorithms are selected based on semiotic class and language of the input as claimed. Text normalization and inverse text normalization system which the rule-based algorithms are selected based on semiotic class and are language specific are known in the art as evidenced by Zhang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system of Beliga to include the text normalization framework taught by Zhang as it merely constitutes the substitution/combination of known processes to achieve the predication result of converting non-standard words. The combination of Beliga and Zhang fails to explicitly disclose determining a language for the input. Although the combination teaches utilizing language-specific algorithms for text normalization and inverse text normalization, it fails to explicitly disclose determining the input language. Huang teaches TTS system architectures which function as a synthesizer for multiple languages where the language-specific information, e.g., special rules for linguistic analysis are loaded by the TTS engine at run-time so that it is possible to switch voices and languages as desired at run-time (Huang, ¶[0066]). The combination of Beliga and Zhang differs from the claimed invention, as defined by claim 8, in that combination fails to explicitly disclose determining a corresponding language and selecting the rule-based algorithms based on the determined language as claimed. ASR systems which determine/select the language specific information/rules needed to generate output text in the desired language are known in the art as evidenced by Huang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the system taught by the combination of Beliga and Zhang to detect/determine a corresponding input language and selected the rule-based algorithms based on the determined language in order to make it possible to switch voices and languages of the system at run-time (Huang, ¶[0066].). Regarding claim 9, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 8 (see detailed element mappings above). In addition, Beliga further teaches the system comprises at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs);a system for performing operations for a conversational Al application; a system for performing operations for a generative Al application; a system for performing operations using a language model; a system for performing one or more generative content operations using a large language model (LLM); a system implemented at least partially in a data center; a system for performing hardware testing using simulation; a system for performing one or more generative content operations using a language model; a system for synthetic data generation; a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources (Beliga teaches the text normalization method/module is for integration into existing text-to-speech synthesis systems. Existing text-to-speech system perform operations using a language model, performing one or more generative operations using a LLM and generate synthetic data, thus Beliga teaches the text normalization can be integrated into one or more of these systems). Regarding claim 10, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 8. In addition, Beliga further teaches the one or more processing units are further to provide the textual output to a trained vocalizer (p. 1667, Table 1, step 4 “The normalized text is sent to the module for grapheme-to-phoneme conversation” and p. 1669, first column teaches “The proposed normalization can easily be integrated with existing grapheme-to-phoneme conversation system [14] and speech generation modules of TTS synthesis system [15]…” the speech generation modules of a TTS synthesis inherently includes a trained vocalizer or vocoder). Regarding claims 12, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 8. In addition, Zhang further teaches the input is an auditory input (p. 1, Abstract teaches “Inverse text normalization (ITN) converts spoken-domain automatic speech recognition (ASR) output into written-domain text”). Beliga differs from the claimed invention, as defined by claim 12, in that Beliga fails to explicitly disclose the rule-based algorithms are selected based on semiotic class and language of the input as claimed. Text normalization and inverse text normalization system which the rule-based algorithms are selected based on semiotic class and are language specific are known in the art as evidenced by Zhang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the system of Beliga to include the text normalization framework taught by Zhang as it merely constitutes the substitution/combination of known processes to achieve the predication result of converting non-standard words. Regarding claim 13, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 8. In addition, Beliga further teaches the input is a first textual input that is in a different form from the textual output (p. 1667, Table 1, step 2c teaches “Write the expanded form of the NSW obtained in the process of normalization to the output” Therefore, the expanded textual output in step 2c is in a different form from the text input received/read in Step 1). Regarding claim 14, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 8. In addition, Beliga further teaches the one or more processing units are further to select a classification rule-based grammar model to classify the one or more tokens into the semiotic class (p. 1667, Table 1, steps 2a-2c teaches the normalization algorithm is selected based on the class/subclass of the NSW, i.e., the semiotic class). Regarding claims 2, 11 and 15, the combination of Beliga, Zhang and Huang teaches all of the elements of claims 1, 8 and 14. In addition, Zhang further teaches the one or more first rule-based algorithms are incorporated into a library of weighted finite state transducers (p. 3 Figure 2 classification WFST and pp. 2-3, section 3). Beliga differs from the claimed invention, as defined by claims 2 and 11, in that Beliga fails to explicitly disclose the rule-based algorithms are selected based on semiotic class and language of the input as claimed. Text normalization and inverse text normalization system which the rule-based algorithms are selected based on semiotic class and are language specific are known in the art as evidenced by Zhang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system of Beliga to include the text normalization framework taught by Zhang as it merely constitutes the substitution/combination of known processes to achieve the predication result of converting non-standard words. Claims 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over by Beliga in view of Zhang. Regarding claim 16, Beliga teaches a processor comprising: processing circuitry (Beliga teaches the text normalization process as an integral part of a text-to-speech (TTS) synthesis system such a system inherently requires at least one processor) to classify a set of tokens into respective semiotic classes (p. 1665, Figure 1, NSW classification and Section B teaches “The text pre-processing module initially has to identify the NSW, and separate it from standard words…Normalization module then classifies NSW according to the given taxonomy of Croatian language…”) using a classification rule-based grammar model for an identified language associated with the set of tokens (p. 1667, Table 1, steps 2a-2c teaches the normalization algorithm is selected based on the class/subclass of the NSW, i.e., the semiotic class) and to process each token, based at least on the respective semiotic class, with a trained class rule-based grammar model (p. 1665, Figure 1, NSW classification and Section B teaches “The text pre-processing module initially has to identify the NSW, and separate it from standard words…Normalization module then classifies NSW according to the given taxonomy of Croatian language…Within each NSW class, certain rules are typical so the normalization can be carried out in a more standard way” and p. 1667, Table 1, step 2), and to combine each processed token into a combined output text sequence (p. 1667, TABLE 1, step 3 teaches “Repeat the procedure iteratively until the input text completely normalized” the iteration of the classification steps results in a combined plain text representation). Beliga fails to disclose using rules for the identified language and the respective semiotic class. Zhang teaches a rule-based Inverse Text Normalization framework which uses a two stage normalization pipeline that first detects semiotic tokens (classification) and then converts these to written form (verbalization). More specifically, Zhang teaches “This framework is general enough to be applied to text normalization and other languages. In fact, Kestrel [4] and Sparrowhawk [8] were originally intended for text normalization across a variety of languages. To do this, we replace the grammars for all classes. The classification grammars will be mostly language-specific” (Zhang, p.3, section 3.6). Thus, Zhang teaches that the grammars, i.e., rules, utilized are based on both the semiotic classification and the specific-language. Beliga differs from the claimed invention, as defined by claim 16, in that Beliga fails to explicitly disclose the rule-based algorithms are selected based on semiotic class and language of the input as claimed. Text normalization and inverse text normalization system which the rule-based algorithms are selected based on semiotic class and are language specific are known in the art as evidenced by Zhang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system of Beliga to include the text normalization framework taught by Zhang as it merely constitutes the substitution/combination of known processes to achieve the predication result of converting non-standard words. Regarding claim 17, the combination of Beliga and Zhang teaches all of the elements of claim 16 (see detailed element mappings above). In addition, Beliga further teaches the system comprises at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs);a system for performing operations for a conversational Al application; a system for performing operations for a generative Al application; a system for performing operations using a language model; a system for performing one or more generative content operations using a large language model (LLM); a system implemented at least partially in a data center; a system for performing hardware testing using simulation; a system for performing one or more generative content operations using a language model; a system for synthetic data generation; a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources (Beliga teaches the text normalization method/module is for integration into existing text-to-speech synthesis systems. Existing text-to-speech system perform operations using a language model, performing one or more generative operations using a LLM and generate synthetic data, thus Beliga teaches the text normalization can be integrated into one or more of these systems). Regarding claim 18, the combination of Beliga and Zhang teaches all of the elements of claim 16. In addition, Zhang further teaches wherein each of the classification rule-based grammar model to classify the set of tokens and the trained class rule-based grammar model include weighted finite state transducers (WFSTs)(p. 3 Figure 2 classification WFST and pp. 2-3, section 3). Beliga differs from the claimed invention, as defined by claim 18, in that Beliga fails to explicitly disclose the rule-based algorithms are selected based on semiotic class and language of the input as claimed. Text normalization and inverse text normalization system which the rule-based algorithms are selected based on semiotic class and are language specific are known in the art as evidenced by Zhang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system of Beliga to include the text normalization framework taught by Zhang as it merely constitutes the substitution/combination of known processes to achieve the predication result of converting non-standard words. Regarding claim 19, the combination of Beliga and Zhang teaches all of the elements of claim 16. In addition, Zhang further teaches wherein the set of tokens is extracted from an auditory input (p. 1, Abstract teaches “Inverse text normalization (ITN) converts spoken-domain automatic speech recognition (ASR) output into written-domain text”). Beliga differs from the claimed invention, as defined by claim 18, in that Beliga fails to explicitly disclose the rule-based algorithms are selected based on semiotic class and language of the input as claimed. Text normalization and inverse text normalization system which the rule-based algorithms are selected based on semiotic class and are language specific are known in the art as evidenced by Zhang. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the TTS system of Beliga to include the text normalization framework taught by Zhang as it merely constitutes the substitution/combination of known processes to achieve the predication result of converting non-standard words. Regarding claim 20, the combination of Beliga and Zhang teaches all of the elements of claim 19. In addition, Beliga further teaches the set of tokens is extracted from a textual input (p. 1665, Figure 1, “Text input” and “Tokenizer” ). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Beliga, Zhang and Huang as applied to claim 1 above, and further in view of Vu et al. (US 2023/0141853 A1; herein “Vu”). Regarding claim 6, the combination of Beliga, Zhang and Huang teaches all of the elements of claim 1 (See detailed element mapping above). Although the combination of Beliga, Zhang and Huang teaches selecting, based at least on the corresponding language (See Huang ¶[0066] and discussion above with respect to claim 1), the combination fails to explicitly teach determining, for the textual input, the language based on characters represented in the one or more segments as claimed. Vu teaches a language detection system and method that includes, inter alia, determining, for the textual input, the language based on characters represented in the one or more segments (¶[0004] teaches “Language detection is the task of identifying the language of a textual input” and ¶[005] teaches “…presenting the input text to a wide network as a sequence of characters, or a sequence of n-grams or subwords. Techniques disclose herein can provide language detection for textual inputs” ). The combination of Beliga, Zhang and Huang differs from the claimed invention, as defined by claim 6, in that the combination fails to explicitly disclose determining a corresponding language based on the characters of the text input as claimed. Language detection based on the sequence of characters in a textual input is known in the art as evidenced by Vu. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention, to have modified the system taught by the combination of Beliga, Zhang and Huang to include detecting the input text language based on the sequence of characters in the input as taught by Vu as it merely constitutes the combination of known processes to achieve the predictable result of detecting the language of the textual input. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PENNY L CAUDLE whose telephone number is (703)756-1432. The examiner can normally be reached M-Th 8:00 am to 5:00 pm eastern. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PENNY L CAUDLE/Examiner, Art Unit 2657 /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Show 3 earlier events
Oct 28, 2025
Applicant Interview (Telephonic)
Nov 06, 2025
Response Filed
Feb 03, 2026
Final Rejection mailed — §103
May 04, 2026
Request for Continued Examination
May 06, 2026
Response after Non-Final Action
May 13, 2026
Non-Final Rejection mailed — §103
Aug 10, 2026
Applicant Interview (Telephonic)
Aug 10, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731587
VISUAL SPEECH RECOGNITION FOR DIGITAL VIDEOS UTILIZING GENERATIVE ADVERSARIAL LEARNING
4y 7m to grant Granted Sep 08, 2026
Patent 12724812
AUTOMATIC GENERATION OF HANDOUTS FROM MULTI-MODAL DOCUMENTS
2y 8m to grant Granted Sep 01, 2026
Patent 12718829
SYSTEMS AND METHODS FOR VOICE RECEPTION AND DETECTION
3y 10m to grant Granted Aug 25, 2026
Patent 12711313
HUMAN-MACHINE COLLABORATIVE CONVERSATION INTERACTION SYSTEM AND METHOD
3y 2m to grant Granted Aug 18, 2026
Patent 12706091
PRONUNCIATION-AWARE EMBEDDING GENERATION FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
2y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
70%
Grant Probability
85%
With Interview (+14.5%)
2y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 84 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month