DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1-20 are pending in this application.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 2A, Prong One: The independent claim 1 recites “obtaining query information of a query text corresponding to user speech using a generative language model; obtaining, based on the query information and a token sequence, a first representation sequence of the token sequence using the generative language model, wherein the token sequence is output in a streaming manner by a large language model based on the query text; encoding the token sequence using a language encoding model to obtain a second representation sequence; and combining the first representation sequence and the second representation sequence to generate an answer speech for the user speech”. The limitation of “obtaining…”, “obtaining…”, “encoding…” and “combining” is a process that, under its broadest reasonable interpretation, covers a human organizing of activities. More specifically, a human receives text and reads using a first and second representation sequence.
Step 2A, Prong Two: This judicial exception is not integrated into a practical application. The computer is recited at a high-level of generality (i.e., as performing a generic computer function and being used as an applying) such that it amounts no more than mere instructions to apply the exception using a generic computer. Accordingly, there additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B — Claims Do Not Recite an Inventive Concept That Transforms the Mental Process into Patent-Eligible Subject Matter. The claims add generic, well-understood computer components (memory, processor, and storage medium) and broadly recite use of “generative language model” and “language encoding model” without describing any specific, unconventional structure, algorithmic detail, data structure, or system architecture that provides a concrete technical improvement in computer functionality.
Applying Alice step two and relevant Federal Circuit precedent:
The mere invocation of “generative language model” and “language encoding model” without particularity does not demonstrate an unconventional machine or technique or a specific improvement in computer technology.
The claims recite high-level, result-oriented steps (e.g., “obtaining,” “encoding,” “combining”) that describe human activities/mental processes rather than specific technical means for performing those processes.
Because the claims lack limitations that tie the mental-process steps to a particular way of achieving a technological improvement (for example, a novel model architecture, specialized data representation, unique training regimen that yields demonstrable technical performance gains, a specialized streaming/decoding pipeline that reduces latency by a quantifiable amount, or hardware/software co-design), the additional elements do not transform the human activities/mental processes into significantly more.
With respect to claims 11 and 20, claims 11 and 20 recite additional element of “processing unit”, “memory” and “storage medium”. The processor and memory are recited at a high-level of generality (i.e., as a generic processor performing generic computer functions and being used as an applying) such that it amounts no more than mere instructions to apply the exception using a generic computer component as well. These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 2 and 12, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 3 and 13, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 4 and 14, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 5 and 15, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 6 and 16, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 7 and 17, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 8 and 18, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to dependent claims 9 and 19, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Conclusion — Rejection
Claims 1-20 are rejected under 35 U.S.C. § 101 as being directed to a judicial exception (mental processes/human activities) and failing to recite additional elements that amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-20 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Jia et al., (US 2022/0310059 A1) in view of Mukerjee et al., (US 2023/0099732 A1).
Regarding claim 1, Jia discloses a method comprising:
obtaining query information of a query text corresponding to user speech using a generative language model (Fig. 1, [0036][0037] obtaining a user query and identifying portions of the text to a search engine);
obtaining, based on the query information and a token sequence, a first representation sequence of the token sequence using the generative language model, wherein the token sequence is output in a streaming manner by a large language model based on the query text (Fig. 1, 2A and 2B, [0036]-[0041] obtaining a sequence of text 152 which will communicate to the user 10 as a response to a query of the spoken utterance 12 using a search engine);
encoding the token sequence using a language encoding model to obtain a second representation sequence (Fig. 1, 2A and 2B, [0037]-[0041] encoding the text input to obtain the input embedding 210 which is a sequence by augmented BERT encoder).
Jia does not explicitly teach however Mukerjee does explicitly teach:
combining the first representation sequence and the second representation sequence to generate an answer speech for the user speech (Mukerjee, [0043] concatenates the textual embedding 224 and the phoneme encoding 222 to generate the concatenation 226 and provides the concatenation 226 as input to the decoder 228 to generate speech).
Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate the method of synthesizing speech using encoding sequence as taught by Jia with the method of generating a concatenation of the textual embedding and the phoneme encoding as taught by Mukerjee to provide Advantages which do not need to process an additional label at run-time in order to synthesize the speech (Mukerjee, [0023]).
Regarding claim 2, Jia in view of Mukerjee discloses the method according to claim 1, and Jia further discloses:
wherein obtaining the first representation sequence of the token sequence comprises: generating, for a given token in the token sequence, a representation vector corresponding to the given token based on the query information and at least one token preceding the given token (Jia, Fig. 2, [0039] converting into a context vector corresponding to the received text which is in the input text sequence).
Regarding claim 3, Jia in view of Mukerjee discloses the method according to claim 1, and Jia further discloses:
wherein obtaining the first representation sequence of the token sequence comprises: generating the first representation sequence in a streaming manner (Jia, Figs.1 and 2, [0037]-[0039] generating the sequence, e.g., ‘Today is Sunny’, which is generated by speech-enabled device 110 in a streaming manner).
Regarding claim 4, Jia in view of Mukerjee discloses the method according to claim 1, and Jia further discloses:
wherein obtaining the second representation sequence comprises: determining at least one consecutive token in the token sequence; and generating a code of the at least one token using the language encoding model ([0029]-[0031] determining a text sequence and encoding using an encoder for a Bidirectional Encoder Representations from Transformers (BERT) model).
Regarding claim 5, Jia in view of Mukerjee discloses the method according to claim 4, and Mukerjee further discloses:
wherein combining the first representation sequence and the second representation sequence comprises: combining the code of the at least one token and a corresponding representation vector in the first representation sequence ([0030][0041]-[0045] combining the generated a textual embedding and identifiers for the domain which may be vectors) .
The previous motivation statement as in claim 1 is still applied.
Regarding claim 6, Jia in view of Mukerjee discloses the method according to claim 5, and Jia further discloses:
wherein there is an offset between a first index of the code in the first representation sequence and a second index of the corresponding representation vector in the second representation sequence ([0039]-[0041] “the word position embedding Ewp differs from the position embedding Ep in that the position embedding is an overall index of position for the plurality of tokens”).
Regarding claim 7, Jia in view of Mukerjee discloses the method according to claim 4, and Jia further discloses:
determining a clause boundary label for the token sequence based on the first representation sequence; and determining at least one consecutive token in the token sequence as a clause based on the clause boundary label ([0040] “the SEP token 212 SEP functions as a separator appended to each segment to indicate where one segment ends and another segment begins”).
Regarding claim 8, Jia in view of Mukerjee discloses the method according to claim 4, and Mukerjee further discloses:
wherein the code of the at least one token and the corresponding representation are in the form of vectors, and the combining comprises: performing addition and/or concatenation of the vectors (Mukerjee, [0043] concatenates the textual embedding 224 and the phoneme encoding 222 to generate the concatenation 226 and provides the concatenation 226 as input to the decoder 228 to generate speech).
The previous motivation statement as in claim 1 is still applied.
Regarding claim 9, Jia in view of Mukerjee discloses the method according to claim 1, and Jia further discloses:
providing a combined representation sequence to a text-to-speech (TTS) front-end task, wherein the TTS front-end task comprises at least one of text normalization, prosody labeling, or grapheme-to-phoneme ([0041]-[0048] providing the alignment context for the augmented encoder to incorporate graphemes in addition to phonemes).
Regarding claim 10, Jia in view of Mukerjee discloses the method according to claim 1, and Jia further discloses:
wherein a token of the token sequence comprises any of a phoneme, a grapheme, a morpheme, or a word ([0027][0041] tokens may comprise of a phoneme, a grapheme, or a word and the alignment between a phoneme and a grapheme is represented in the input embedding).
Regarding claims 11-19, Claims 11-19 relate to an electric device and have substantially the same technical features as method of claims 1-9. Accordingly, the same rational as in claims 1-9 can be applied to claims 11-19.
Regarding claim 20, Claim 20 relates to an electric device and has substantially the same technical features as method of claim 1. Accordingly, the same rational as in claim 1 can be applied to claim 20.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see attached form PTO-892.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEONG-AH A. SHIN whose telephone number is (571)272-5933. The examiner can normally be reached 9 AM-3PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Seong-ah A. Shin
Primary Examiner
Art Unit 2659
/SEONG-AH A SHIN/Primary Examiner, Art Unit 2659