DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 16 and 17 recite the limitation "the presentation". There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101.
Claims 1, 8 and 15 are rejected as being directed to an abstract idea. Under Step 2A, prong one, the claims recite the abstract mental process of evaluating linguistic information and converting it into normalized plain text. A person can read text, apply linguistic rules or judgments, and produce a standardized version. The claims are characterized as a mental process, and they do not improve the machine learning algorithm itself.
Under Step 2B, prong two, the claims do not integrate the abstract idea into a practical application. They merely require generic processors and broadly identified “rule-based” and “machine learning” algorithms. They do not specify any normalization rule, model architecture, training technique, order or interaction between the algorithms, specialized data structure, or improvement in accuracy, processing speed, memory use, or computer operation. Generating plain text output is simply the result of the abstract information processing task. The Office’s AI examples explain that invoking a trained neural network without describing how it performs the claimed result, followed by generic output, amounts to using a computer merely as a tool.
Under Step 2B, the claims contain no inventive concept. The generic processors perform ordinary data processing, while the rule-based and machine-learning components are recited only as result oriented labels. Even as an ordered combination, the claims do not state how those components cooperate to provide a nonconventional technical solution.
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device.
Dependent claims 2-7, 9-14 and 16-20 further recite an abstract idea performable by a human and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known.
Claims 2 and 9, merely transform information and add no technological improvement.
Claims 3 and 10, merely evaluate and organize information without improving computer technology.
Claims 4 and 11, not explaining how the model provides a technical improvement.
Claims 5 and 12, not an improvement to speech synthesis technology.
Claims 6 and 13, add no specific technological improvement or inventive concept.
Claims 7 and 14, merely used as a tool and do not provide an improvement in computer function.
Claim 16, not improving how synthetic speech is created.
Claim 17, not improving SR or display technology.
Claim 18, not improving the computer or ML model.
Claim 19, no technological improvement.
Claim 20, merely limits the used and adds no inventive concept.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-4, 8-11, 15 and 20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Sproat et al. (US 9,852,123).
Claims 1 and 8,
Sproat teaches a system comprising (Sproat states “a system, comprising a data processing apparatus, and a non-transitory computer readable storage medium in data communication with the data processing apparatus storing instructions executable by the data processing apparatus” [claim 1]; Sproat also identifies “a language processing system 120” that performs “text normalization for input strings of semiotic classes” [col. 2 lines 52-58]):
one or more processors (Sproat teaches that a data-processing apparatus includes “a programmable processor, a computer, a system on a chip or multiple ones or combinations of the foregoing” [col. 6 lines 8-24])
to generate a plain text output of an input (Sproat defines the output as ordinary words: “such a pronounceable version expressed in terms of ordinary words is referred to as a ‘verbalization’” [col. 1 lines 7-14]; the system “receives an input string w,” “generates a collection of possible verbalizations” for that input, and “selects an appropriate verbalization v” [col. 2 line 59 to col. 3 line 2]; for example, the system generates the alphabetic, ordinary-word output “ninety seven” from the numerical input “97” [col. 4 lines 4-13]; thus, the selected verbalization is a plain-text output corresponding to the input)
based at least on performing text normalization on at least a portion of the input (Sproat defines the operation: “the process of converting instances of semiotic classes to verbalizations is referred to as text normalization” [col. 1 lines 20-27]; Sproat further states that the generator and selector “perform text normalization for semiotic classes” where “the input string w includes text belonging to a semiotic class” [col. 3 lines 9-16]; Normalizing the semiotic-class text contained in the input teaching satisfies normalizing “at least portion” of the input)
using a processing sequence (Sproat teaches the ordered process of: 1) receiving an input string; 2) accessing a covering grammar and lexical maps; 3) generating a lattice of possible verbalizations; and 4) selecting a verbalization; specifically, the generator “generates a collection of possible verbalizations Hw” and then, “using a verbalization selector 124, selects an appropriate verbalization v” [col. 2 line 59 to col. 3 line 35] [col. 4 lines 4-35] [steps 202, 204, 208 and 210])
that includes one or more rule-based algorithms (Sproat states “for each semiotic class, the covering grammar G incorporates rules that will map the input string to the word-level components”; the covering grammar may include “factorization rules” consistent with the spoken form of a number [col. 3 lines 44-60]; the grammar and lexical maps may be implemented as finite-state transducers [col. 4 lines 14-20]; thus, the covering grammar is a rule-based algorithm that generates candidate verbalizations from the input) and
one or more machine learning algorithms (Sproat further teaches that after the rule-based grammar generates lattice Hw, Sproat teaches that the verbalization selector “may be based on a supervised sequence model, in which case v is the highest-scoring path through Hw”; the selector may specifically be “a MaxEnt ranker that is trained using parallel data” [col. 4 lines 48-61]; Sproat further identifies “the machine learner system 140” and explains that it “trains the scoring model to learn preferred lattice outputs” [col. 4 line 66 to col. 5 line 4] [claim 4]).
Claims 2 and 9,
Sproat further teaches the system of claim 1, wherein the one or more processors are further to: generate, using the one or more rule-based algorithms, one or more plain text representations from one or more tokens associated with the input (Sproat teaches that the input contains an instance of a semiotic class, such as “97”; the covering-grammar rules map that non-standard-word instance to word components, and the lexical map generates one or more ordinary-word verbalizations, include “ninety seven”; under BRI, the semiotic-class nonstandard word is the claimed token [col. 3 lines 9-30] [col. 3 line 44 to col. 4 line 20] [claim 1]).
Claims 3 and 10,
Sproat further teaches the system of claim 2, wherein the one or more processors are further to: select, using the one or more rule-based algorithms, a subset of the one or more plain text representations (Sproat teaches a grammatical ruleset is applied to every candidate; noncompliant verbalizations may be “filtered from being selected”; filtering candidates produces a subset [col. 4 lines 21-35] [claim 7]; Sproat present this rule filter and col. 4 lines 36-47 describes weighting as alternative selection processes) from a weighted set of the one or more plain text representations (Sproat generates a score for each possible verbalization and implements the scoring function as a “weighted finite state acceptor”; thus, its candidate verbalizations constitute a weighted set [col. 4 lines 36-47] [claims 2-3]).
Claims 4 and 11,
Sproat further teaches the system of claim 3, wherein the one or more processors are further to: select, using the one or more machine learning algorithms, a selected plain text representation from the subset of the one or more plain text representations (Sproat’s supervised sequence model selects the “highest-scoring path through Hw,” and its trained MaxEnt model learns preferred verbalizations; in the modified system of claim 3, Hw presented to the trained surviving subset [col. 4 line 48 to col. 5 line 27] [claims 3 and 4]).
Claim 15,
Sproat teaches a method comprising: causing presentation of an output (Sproat teaches a computer display “for displaying formation to the user” and visual, auditory, or tactile user feedback; Sproat also teaches transmitting data to a user device for display and permits separately described features to be combined [col. 7 line 12 to col. 8 line 6]),
the output generated based at least on performing at least one of text normalization or inverse text normalization (Sproat teaches that conversion of semiotic-class instances to verbalization is “referred to as text normalization” [col. 1 lines 20-27]; the generator produces candidate verbalizations and the selector selects verbalization [col. 2 line 59 to col. 3 line 2]; for input “97”, Sproat generates the ordinary-word verbalization “ninety seven” [col. 4 lines 4-13])
on at least a portion of an input (Sproat receives input string w, which contains text belonging to a “semiotic class,” and performs text normalization on that text [col. 3 lines 9-17])
using a processing sequence that includes one or more rule-based algorithms (Sproat’s covering grammar “incorporates rules” that map the input string to word-level components; the grammar and lexical maps generate the candidate-verbalization lattice [col. 3 line 44 to col. 4 line 20]) and
one or more machine learning algorithms (Sproat further teaches that after the rule-based grammar generates lattice Hw, Sproat teaches that the verbalization selector “may be based on a supervised sequence model, in which case v is the highest-scoring path through Hw”; the selector may specifically be “a MaxEnt ranker that is trained using parallel data” [col. 4 lines 48-61]; Sproat further identifies “the machine learner system 140” and explains that it “trains the scoring model to learn preferred lattice outputs” [col. 4 line 66 to col. 5 line 4] [claim 4]).
Claim 20,
Sproat further teaches the method of claim 15, wherein the method is performed using at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; an infotainment system of a machine; an entertainment system of a machine; a system for generating synthetic data; a system for collaborative content creation of multi-dimensional assets; a system for performing digital twin simulation; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system for rendering graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system incorporating one or more Virtual Machines (VMs) (Sproat teaches a “virtual machine” [col. 6 lines 8-24]); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 5 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sproat et al. (US 9,852,123) and further in view of Yamasaki et al. (US 2015/0269927).
Claims 5 and 12,
Sproat teaches all the limitations in claim 4. The difference between the prior art and the claimed invention is that Sproat does not explicitly teach provide the selected plain text representation to a text-to-speech system.
Yamasaki teaches provide the selected plain text representation to a text-to-speech system (Yamasaki’s selector “outputs, to the generator 31, the selected normalized text”; Generator 31 is part of synthesizer 30, which generates a speech waveform, and generates phonetic parameters for the selected normalized text [0023] [0037] [0065] [Fig. 1] [claim 1]).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Sproat with teachings of Yamasaki by modifying the semiotic class normalization as taught by Sproat to include provide the selected plain text representation to a text-to-speech system as taught by Yamasaki for the benefit of correctly analyzing the text containing peculiar expressions (Yamasaki [0004]).
Claim(s) 6 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sproat et al. (US 9,852,123) and further in view of Ebden et al. (“The Kestrel TTS text normalization system”; 2014).
Claims 6 and 13,
Sproat teaches all the limitations in claim 3. The difference between the prior art and the claimed invention is that Sproat does not explicitly teach wherein the weighted set of the one or more plain text representations is determined using a weighted finite-state transducer.
Ebden teaches wherein the weighted set of the one or more plain text representations is determined using a weighted finite-state transducer (Ebden’s normalization grammars are “complied into libraries of weighted finite-state transducers (WFSTs)”; WFST costs operate as “hand-assigned penalties that allow for ranking of different possible analyses,” and its number-verbalization “grammar produces all possible forms,” with weights attached to preferred alternatives [pgs. 1, 3 and 11]).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Sproat with teachings of Ebden by modifying the semiotic class normalization as taught by Sproat to include wherein the weighted set of the one or more plain text representations is determined using a weighted finite-state transducer as taught by Ebden for the benefit of separating the initial tokenization and classification phase of analysis from verbalization (Edben [Abstract]).
Claim(s) 7 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sproat et al. (US 9,852,123) and further in view of Zhang et al. (“Neural Models of Text Normalization for Speech Applications”; 2019).
Claims 7 and 14,
Sproat teaches all the limitations in claim 3. The difference between the prior art and the claimed invention is that Sproat does not explicitly teach wherein the one or more machine learning algorithms include one or more neural networks.
Zhang teaches wherein the one or more machine learning algorithms include one or more neural networks (Zhang states “we propose neural network models that treat text normalization for TTS as a sequence to sequence problem”; the disclosed architecture’s core is “an RNN based attentional sequence to sequence network” [Abstract] [5.3] [Fig. 2]; the network uses a bidirectional GRU encoder, attention based decoding and additional GRU context networks).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Sproat with teachings of Zhang by modifying the semiotic class normalization as taught by Sproat to include wherein the one or more machine learning algorithms include one or more neural networks as taught by Zhang for the benefit of finding the most effective model, in accuracy and efficiency to normalize text for speech applications using neural network models (Zhang [Abstract]).
Claim(s) 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sproat et al. (US 9,852,123) and further in view of Pusateri et al. (US 2019/0278841).
Claim 16,
Sproat teaches all the limitations in claim 15. The difference between the prior art and the claimed invention is that Sproat does not explicitly teach wherein the output includes a text to speech output, and the presentation includes an audible output of synthetic speech corresponding to the text to speech output.
Pusateri teaches wherein the output includes a text to speech output (Pusateri sends a generated response to “speech synthesis processing module 740” to synthesize the response “in speech form” [0243]; the module synthesizes speech outputs based on text [0245]), and
the presentation includes an audible output of synthetic speech corresponding to the text to speech output (Pusateri teaches that the speech synthesis module “converts the text string to an audible speech output” and synthesizes that output for presentation to the user [0245]; [0246] sends the synthesized speech to the user device for output).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Sproat with teachings of Pusateri by modifying the semiotic class normalization as taught by Sproat to include wherein the output includes a text to speech output, and the presentation includes an audible output of synthetic speech corresponding to the text to speech output as taught by Pusateri for the benefit of improving the battery life of the device (Pusateri [0006]).
Claim 17,
Sproat teaches all the limitations in claim 15. The difference between the prior art and the claimed invention is that Sproat does not explicitly teach wherein the output includes a speech to text output, and the presentation includes a visual display of the text to speech output.
Pusateri teaches wherein the output includes a speech to text output (Pusateri sends a generated response to “speech synthesis processing module 740” to synthesize the response “in speech form” [0243]; the module synthesizes speech outputs based on text [0245]), and
the presentation includes a visual display of the text to speech output (Pusateri teaches displaying the ITN-generated written-form speech-recognition result to the user [0273]; [0058], confirms that touch screen 212 displays visual output including text [0289]).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Sproat with teachings of Pusateri by modifying the semiotic class normalization as taught by Sproat to include wherein the output includes a speech to text output, and the presentation includes a visual display of the text to speech output as taught by Pusateri for the benefit of improving the battery life of the device (Pusateri [0006]).
Claim(s) 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sproat et al. (US 9,852,123) and further in view of Baldwin et al. (US 2015/0186355).
Claim 18,
Sproat further teaches the method of claim 15, further comprising: generating a weighted set of one or more representations (Sproat generates “a lattice of possible verbalizations”; Hw is “a lattice of all possible verbalizations”; under the supervised model, selectee v is the “the highest-scoring path through Hw” [col. 4 lines 4-61]); and
selecting, from the one or more representations, a selected representation using at least the one or more machine learning algorithms (Sproat teaches that the selected verbalization is “the highest-scoring path through Hw,” scored using a trained MaxEnt ranker; the machine learner trains the scoring model to learn preferred lattice outputs).
The difference between the prior art and the claimed invention is that Sproat does not explicitly teach for an input including a token and plain text.
Baldwin teaches for an input including a token and plain text (Baldwin teaches that “the original input text (un-normalized) may be represented as a sequence … of tokens,” illustrated by “ay woudent see ‘em” [0014-0015]; generators operate on that sequence and may replace woudent while leaving other ordinary words unchanged [0008]; discussing “non-standard word tokens” within the input text).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Sproat with teachings of Pusateri by modifying the semiotic class normalization as taught by Sproat to include for an input including a token and plain text as taught by Baldwin for the benefit of correcting grammar which includes modifying punctuation and capitalization as well as adding, removing and reordering words (Baldwin [0008]).
Claim 19,
Sproat further teaches the method of claim 18, wherein the token corresponds to one or more classes including at least one of a number, a letter, a symbol, a fraction, or a date (Sproat teaches standard words include “numbers (e.g. 97)” and “dates (e.g. 3/23)” [col. 1 lines 7-14]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Skinner et al. (US 2020/0042837) – Provided is a process, including: receiving a screen capture event from an operating system of a first client computing device of a first user, the screen capture event including, or being associated with, a bitmap image of at least part of a display of the first computing device; causing optical character recognition (OCRing) of text in the bitmap image; classifying each of the n-grams into two or more categories, the two or more categories including a category for confidential information; and for each of the n-grams classified in the category for confidential information, obfuscating the respective n-gram in the bitmap image to form a modified version of the bitmap image.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
SHREYANS A. PATEL
Primary Examiner
Art Unit 2653
/SHREYANS A PATEL/ Examiner, Art Unit 2659