Prosecution Insights
Last updated: August 16, 2026
Application No. 18/683,786

SPEECH SYNTHESIS APPARATUS, SPEECH SYNTHESIS METHOD, AND SPEECH SYNTHESIS PROGRAM

Non-Final OA §103
Filed
Feb 15, 2024
Priority
Aug 18, 2021 — JP 2021-133713 +1 more
Examiner
LAM, PHILIP HUNG FAI
Art Unit
2656
Tech Center
2600 — Communications
Assignee
The University of Tokyo
OA Round
3 (Non-Final)
85%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
127 granted / 150 resolved
+22.7% vs TC avg
Strong +48% interview lift
Without
With
+48.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
28 currently pending
Career history
170
Total Applications
across all art units

Statute-Specific Performance

§101
24.0%
-16.0% vs TC avg
§103
55.7%
+15.7% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
4.1%
-35.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 150 resolved cases

Office Action

§103
DETAILED ACTION A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 4/22/2026 has been entered. Response to Amendment and Arguments 35 U.S.C. 102 and 103 Rejections Applicant’s arguments are moot in view of the new or modified grounds of rejection that were necessitated by the amendments to the Claims. Applicant’s arguments are directed to material that is added by the most recent amendments to the independent Claims. Response, p. 9. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. Claims 1-2, and 6-8 are rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson (US 20220068296). Regarding Claim 1, Kingsbury discloses: A speech synthesis apparatus comprising: a memory; and a processor coupled to the memory and configured to: ([0009] In accordance with embodiments herein, a device is provided that is comprised of a processor and a memory storing program instructions accessible by the processor.) obtain (i) utterance information on texts contained in specific book data of a book; ([0009] Responsive to execution of the instructions, the processor analyzes textual content to identify narratives for associated content segments of the textual content; designates character models for the corresponding narratives; and generates an audio rendering of the textual content utilizing the character models in connection with the corresponding narratives for the associated content segments.) Also see para 0034. obtain (ii) image information on images that are contained in the specific book data of the book; ([0024] The term “textual content” refers to any and all textual, graphical, image or video information or data that may be converted to corresponding audio information or data. The textual content may represent various types of incoming and outgoing textual, graphical, image and video content including, but not limited to, electronic books, electronic stories, correspondence, email, webpages, technical material, text messages, social media content, alerts, advertisements and the like.) obtain (iii) speech data of a reading of the text contained in the specific book data of the book; ([0009] Responsive to execution of the instructions, the processor analyzes textual content to identify narratives for associated content segments of the textual content; designates character models for the corresponding narratives; and generates an audio rendering of the textual content utilizing the character models in connection with the corresponding narratives for the associated content segments.) Also see para 0027. and generate, based on the obtained (i) utterance information, the ([0009] Responsive to execution of the instructions, the processor analyzes textual content to identify narratives for associated content segments of the textual content; designates character models for the corresponding narratives; and generates an audio rendering of the textual content utilizing the character models in connection with the corresponding narratives for the associated content segments.) Also see para 0027 and 0034. Kingsbury does not appear to disclose convert the image information into visual feature vectors, each visual feature vector representing a visual feature of a corresponding image of the (ii) image information contained in the specific book data of the book from which the (i) utterance information and the (iii) speech data were obtained, and each visual feature representing at least one of an appearance of a character, an expression of a character, a background scenery, or weather depicted in the corresponding image; Wilson in the related art discloses: convert the image information into visual feature vectors, each visual feature vector representing a visual feature of a corresponding image of the (ii) image information contained in the specific book data of the book from which the (i) utterance information and the (iii) speech data were obtained, and each visual feature representing at least one of an appearance of a character, an expression of a character, a background scenery, or weather depicted in the corresponding image; ([0037] In various embodiments, program 150 incorporates determined associated weather (e.g., snowing, hot, raining, etc.) into the location model resulting in an accurate generated backdrop. In an embodiment, program 150 inputs (e.g., feeds/fed) one or more generated avatars, an identified utterance sentiment, an identified topic, and location into the trained image generator model generating an image representation of the utterance. … (e.g., hair color, unique facial features, sentiment ranges (e.g., do not allow anger or negative emotions), etc.), etc.). For example, the user designates that no personal information be present in the generated image, thus program 150 generates an image representation only containing a genericized avatar with little to no expressed sentiment, generalized backdrop with no weather additions, and an identified topic. … In various embodiments, program 150 creates an image representation for each utterance, each topic, or each user. In these embodiments, program 150 chains a plurality of generated image representations (e.g., comic panels) forming a cohesive (i.e., unified) “storyline” or aggregation of similar topics and recipients. … In the depicted embodiment, program 150 generates the image representation stylized in a comic (e.g., comic book/strip) fashion.) visual feature vectors representing the visual features of the corresponding images of the (ii) image information contained in the specific book data of the book from which the (i) utterance information and the (iii) speech data were obtained, ([0037] In various embodiments, program 150 incorporates determined associated weather (e.g., snowing, hot, raining, etc.) into the location model resulting in an accurate generated backdrop. In an embodiment, program 150 inputs (e.g., feeds/fed) one or more generated avatars, an identified utterance sentiment, an identified topic, and location into the trained image generator model generating an image representation of the utterance. … (e.g., hair color, unique facial features, sentiment ranges (e.g., do not allow anger or negative emotions), etc.), etc.). For example, the user designates that no personal information be present in the generated image, thus program 150 generates an image representation only containing a genericized avatar with little to no expressed sentiment, generalized backdrop with no weather additions, and an identified topic. … In various embodiments, program 150 creates an image representation for each utterance, each topic, or each user. In these embodiments, program 150 chains a plurality of generated image representations (e.g., comic panels) forming a cohesive (i.e., unified) “storyline” or aggregation of similar topics and recipients. … In the depicted embodiment, program 150 generates the image representation stylized in a comic (e.g., comic book/strip) fashion.) Kingsbury and Wilson are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury to combine with the teaching of Wilson to disclose the above feature, because visual narrative can make storytelling more interesting and thus enhance communication (Wilson, [0010]). Regarding Claim 2, Kingsbury and Wilson disclose all the limitation of Claim 1 (see detailed mapping from above) Kingsbury further discloses: wherein the processor configured to obtain, as the image information, information on an image that is contained in a specific page of the book and that is associated with one of the texts that is contained in the specific page of the book. ([0024] The term “textual content” refers to any and all textual, graphical, image or video information or data that may be converted to corresponding audio information or data. The textual content may represent various types of incoming and outgoing textual, graphical, image and video content including, but not limited to, electronic books, electronic stories, correspondence, email, webpages, technical material, text messages, social media content, alerts, advertisements and the like. [0085] The camera unit 616 may capture one or more frames of image data.) Claim 6 recites a method that corresponds to the apparatus of claim 1 and is therefore rejected under the same grounds as claim 1 above. Claim 7 recites a non-transitory computer readable storage medium that corresponds to the apparatus of claim 1 and is therefore rejected under the same grounds as claim 1 above. Kingsbury discloses ([0097] implemented as hardware with associated instructions (for example, software stored on a tangible and non-transitory computer readable storage medium, such as a computer hard drive, ROM, RAM, or the like) that perform the operations described herein.) Regarding Claim 8, Kingsbury discloses: A speech synthesis apparatus comprising: a memory; and a processor coupled to the memory and configured to: ([0009] In accordance with embodiments herein, a device is provided that is comprised of a processor and a memory storing program instructions accessible by the processor.) Obtain (i) utterance information on a text contained in data of a book; ([0009] Responsive to execution of the instructions, the processor analyzes textual content to identify narratives for associated content segments of the textual content; designates character models for the corresponding narratives; and generates an audio rendering of the textual content utilizing the character models in connection with the corresponding narratives for the associated content segments.) Also see para 0034. Obtain (ii) image information on an image contained in the data of the book, wherein the image information corresponds to the text contained in the data of the book; ([0024] The term “textual content” refers to any and all textual, graphical, image or video information or data that may be converted to corresponding audio information or data. The textual content may represent various types of incoming and outgoing textual, graphical, image and video content including, but not limited to, electronic books, electronic stories, correspondence, email, webpages, technical material, text messages, social media content, alerts, advertisements and the like.) acquire a synthesized speech corresponding to the text by inputting the obtained (i) utterance information and the ([0009] Responsive to execution of the instructions, the processor analyzes textual content to identify narratives for associated content segments of the textual content; designates character models for the corresponding narratives; and generates an audio rendering of the textual content utilizing the character models in connection with the corresponding narratives for the associated content segments.) Also see para 0027 and 0034. Kingsbury does not appear to discloses convert the image information into a visual feature vector, the visual feature vector representing a visual feature of the image of the (ii) image information contained in the data of the book from which the (i) utterance information was obtained, and the visual feature representing at least one of an appearance of a character, an expression of a character, a background scenery, or weather depicted in the image; Wilson in the related art discloses: convert the image information into a visual feature vector, the visual feature vector representing a visual feature of the image of the (ii) image information contained in the data of the book from which the (i) utterance information was obtained, and the visual feature representing at least one of an appearance of a character, an expression of a character, a background scenery, or weather depicted in the image; ([0037] In various embodiments, program 150 incorporates determined associated weather (e.g., snowing, hot, raining, etc.) into the location model resulting in an accurate generated backdrop. In an embodiment, program 150 inputs (e.g., feeds/fed) one or more generated avatars, an identified utterance sentiment, an identified topic, and location into the trained image generator model generating an image representation of the utterance. … (e.g., hair color, unique facial features, sentiment ranges (e.g., do not allow anger or negative emotions), etc.), etc.). For example, the user designates that no personal information be present in the generated image, thus program 150 generates an image representation only containing a genericized avatar with little to no expressed sentiment, generalized backdrop with no weather additions, and an identified topic. … In various embodiments, program 150 creates an image representation for each utterance, each topic, or each user. In these embodiments, program 150 chains a plurality of generated image representations (e.g., comic panels) forming a cohesive (i.e., unified) “storyline” or aggregation of similar topics and recipients. … In the depicted embodiment, program 150 generates the image representation stylized in a comic (e.g., comic book/strip) fashion.) visual feature vector representing the visual feature of the image of the (ii) image information contained in the data of the book from which the (i) utterance information was obtained ([0037] In various embodiments, program 150 incorporates determined associated weather (e.g., snowing, hot, raining, etc.) into the location model resulting in an accurate generated backdrop. In an embodiment, program 150 inputs (e.g., feeds/fed) one or more generated avatars, an identified utterance sentiment, an identified topic, and location into the trained image generator model generating an image representation of the utterance. … (e.g., hair color, unique facial features, sentiment ranges (e.g., do not allow anger or negative emotions), etc.), etc.). For example, the user designates that no personal information be present in the generated image, thus program 150 generates an image representation only containing a genericized avatar with little to no expressed sentiment, generalized backdrop with no weather additions, and an identified topic. … In various embodiments, program 150 creates an image representation for each utterance, each topic, or each user. In these embodiments, program 150 chains a plurality of generated image representations (e.g., comic panels) forming a cohesive (i.e., unified) “storyline” or aggregation of similar topics and recipients. … In the depicted embodiment, program 150 generates the image representation stylized in a comic (e.g., comic book/strip) fashion.) Kingsbury and Wilson are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury to combine with the teaching of Wilson to disclose the above feature, because visual narrative can make storytelling more interesting and thus enhance communication (Wilson, [0010]). Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson and further view of Li (US 20200027454). Regarding Claim 3, Kingsbury and Wilson disclose all the limitation of Claim 1 (see detailed mapping from above). Kingsbury and Wilson do not explicitly disclose the feature recited below. Li (in the related field of text and voice information processing) discloses: wherein the processor configured to obtain, as the speech data, data of speech reading out one of the texts that is contained in a specific page of the book and that is associated with an image contained in the specific page of the book. ([0048] When attending a conference, to better record meeting content, the user may not only obtain the target picture by using the terminal, but also obtain the voice file corresponding to the text information in the target picture. The voice file is in one-to-one correspondence with the text information.) [Also see fig. 4, where it is obvious that the text described the image is on the same page] Kingsbury, Wilson and Li are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury and Wilson to combine with the teaching of Li to disclose the above feature, because voice associating text with image will aid reader/viewer understand content, especially those that may be visually impaired (Li, [0048]). Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson, and further in view of Mitsui (US 20170345412). Regarding Claim 4, Kingsbury and Wilson disclose all the limitation of Claim 1 (see detailed mapping from above). Wilson further discloses: wherein the processor is configured to obtain the utterance information presenting at least one of ([0027] Responsive to program 150 detecting a conversational utterance, program 150 utilizes natural language processing (NLP) techniques and corpus linguistic analysis techniques (e.g., syntactic analysis, etc.) to identify parts of speech and syntactic relations between various portions of the utterance. Program 150 utilizes corpus linguistic analysis techniques, such as part-of-speech tagging, statistical evaluations, optimization of rule-bases, and knowledge discovery methods, to parse, identify, and evaluate portions of the utterance. In an embodiment, program 150 utilizes part-of-speech tagging to identify the particular part of speech of one or more words in an utterance based on its relationship with adjacent and related words. For example, program 150 utilizes the aforementioned techniques to identity the nouns, adjectives, adverbs, and verbs in the example sentence: “Henry, I believe this link will solve your issue”. In this example, program 150 identifies “Henry”, “link”, and “issue” as nouns, “solve” and “believe” as verbs.) Kingsbury and Wilson do not explicitly disclose the feature recited below. Mitsui (in the related field of processing speech) discloses: wherein the processor is configured to obtain the utterance information presenting at least one of accents, ([0081] The applicable segment searching unit 108 may compare an accent phrase obtained by dividing input utterance information (hereinafter referred to as an input accent phrase) with an accent phrase obtained by dividing original-speech utterance information (hereinafter referred to as an original-speech accent phrase).) [the claim only required one of the recited elements] Kingsbury, Wilson and Mitsui are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury and Wilson to combine with the teaching of Mitsui to disclose the above feature, because when generating synthetic speech, matching to a pre-segmented accent phrase library allows the system to produce more natural-sounding output, as it can select segments that best match the desired prosodic features (Mitsui, [0081]). Claims 5 is rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson and further in view of Kolluru (US 20150052084). Regarding Claim 5, Kingsbury and Wilson disclose all the limitation of Claim 1 (see detailed mapping from above) Wilson further discloses: wherein the processor further configured to: convert the utterance information into a language vectors, wherein each language vector represents linguistic information on a corresponding text of the texts; ([0029] Responsive to program 150 processes the utterance, program 150 vectorizes the processed utterance. In an embodiment, program 150 utilizes one-hot encoding techniques to vectorize categorical or string-based feature sets. For example, when vectorizing feature sets of individual words, program 150 creates a one-hot vector comprising a 1×N matrix, where N symbolizes the number of distinguishable words. In another embodiment, program 150 utilizes one-of-c coding to recode categorical data into a vectorized form. For example, when vectorizing an example categorical feature set consisting of [sunshine, happy, vacation], program 150 encodes the corresponding feature set into [[1,0,0], [0,1,0], [0,0,1]]. In another embodiment, program 150 utilizes featuring scaling techniques (e.g., rescaling, mean normalization, etc.) to vectorize and normalize numerical feature sets. In various, program 150 utilizes lda2vec (e.g., word embedding) to convert the aforementioned latent Dirichlet allocation (LDA) and biterm topic results, documents, and matrices into vectorized representations.) Kingsbury and Wilson do not appear to disclose the following. Kolluru in the related art discloses: and generate the speech synthesis model using training data containing the speech data that is associated with the language vectors and the visual feature vectors. ([0074] produce a statistical model, said statistical model comprising a plurality of model parameters describing probability distributions which relate an acoustic unit to an image vector and speech vector, said image vector comprising a plurality of parameters which define the subject's face and said speech vector comprising a plurality of parameters which define the subject's voice, [0075] the processor being configured to train said statistical model such that a sequence of speech vectors and image vectors which are synchronised when outputted cause the generated head to appear to talk.) Kingsbury, Wilson and Kolluru are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury and Wilson to combine with the teaching of Kolluru to disclose the above feature, because the technique described by Kolluru enable the ability to be able to emulate a subject such that a subject's voice, face and dialogue intelligence are emulated, has a wide variety of uses such a providing a human-like interface to a query system through to providing a personalised avatar which can represent the subject (Kolluru, [Background]). Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson, and further in view of Ma (already of record). Regarding Claim 9, Kingsbury and Wilson discloses all the limitation of Claim 1 (see detailed mapping from above) Kingsbury and Wilson do not explicitly disclose the below cited features. Ma in the related art discloses: wherein the speech synthesis model includes: a first input configured to receive a language vector derived from the utterance information; ([sect 3.1 Modality-specific encoders] Text encoder: We process text into a sequence of 128-D character-level embeddings via a 66-symbol trainable lookup table. We then feed each of the embeddings into two fully-connected (FC) layers.) [text reads on utterance information, as utterance information comes from text] Also see fig. 3 for the architecture of the modality specific encoders and decoders. – reproduced below for convenience of viewing. an encoder layer configured to encode the language vector; (see sect 3.1 and fig. 3 below) PNG media_image1.png 426 975 media_image1.png Greyscale a second input configured to receive the visual feature vectors; ([sect 3.1 Modality-specific encoders] Image encoder: We feed images to a three-layer CNN and perform max-pooling to obtain the output eimg ∈ R512.) Also see fig. 3 for the architecture of the modality specific encoders and decoders. a visual information extraction layer configured to extract features from the visual feature vector; ([sect 3.1 Modality-specific encoders] Image encoder: We feed images to a three-layer CNN and perform max-pooling to obtain the output eimg ∈ R512.) Also see fig. 3 for the architecture of the modality specific encoders and decoders. and a decoder layer configured to receive an output from the encoder layer and an output from the visual information extraction layer to generate speech features. (see fig. 3 from above) Kingsbury, Wilson and Ma are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury and Wilson to combine with the teaching of Ma to disclose the above feature, because the technique described by Ma improves performance on traditional cross-modal generation, suggesting that it improves data efficiency in solving individual tasks (Ma, [Abstract]). Claims 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson, and further in view of Krishnamurthy (US 20210319321). Regarding Claim 12, Kingsbury and Wilson disclose all the limitation of Claim 1 (see detailed mapping from above) Kingsbury and Wilson do not explicitly disclose the below cited features. Krishnamurthy discloses: wherein the processor is configured to convert the image information into the visual feature vectors using a neural network for image identification that has been trained previously. ([0054] The multimodal correlation neural network after being trained can generate a representation (embedding) for a given audio. The multimodal correlation neural network after being trained can generate a representation (embedding) for a given image/video. For a pair of correlated image/video and audio, the visual representation generated in and audio representation generated in are likely to be close (that is, distance between them is small).) Kingsbury, Wilson and Krishnamurthy are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury and Wilson to combine with the teaching of Krishnamurthy to disclose the above feature, because the technique described is a self-supervised training which does not required manual human annotations to train (Krishnamurthy, [0054]). Regarding Claim 13, Kingsbury/Wilson/Krishnamurthy disclose all the limitation of Claim 12 (see detailed mapping from above) Krishnamurthy further discloses: wherein the processor is configured to use an output of an intermediate layer of the neural network as the visual feature vectors. [0031] One or more subnetwork layers that are part of the NNs 204 and 205 may be chosen suitable to create a representation, or feature vector of the training data. In some implementations the audio and image input subnetworks may produce embeddings in the form of feature vectors having 128 components, though aspects of the present disclosure are not limited to 128 component feature vectors and may encompass other feature vector configurations and embedding configurations.) Where the rationale for the combination would be similar to the one already provided. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson, and further in view of Ahmed (US 20240127832). Regarding Claim 10, Kingsbury and Wilson disclose all the limitation of Claim 1 (see detailed mapping from above) Kingsbury and Wilson do not explicitly disclose the below cited features. Ahmed discloses: wherein the speech data includes speech parameters including a basic frequency and spectral parameters including a mel-spectrogram. ([0538] These linguistics features (bitstream 3) may then be converted, e.g. through an acoustic model, to acoustics features, like MFCCs, fundamental frequency, mel-spectrogram for example, or a combinations of those. This operation may be performed by a preconditioning layer 710c, which may be either deterministic or learnable.) Kingsbury, Wilson and Ahmed are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury and Wilson to combine with the teaching of Ahmed to disclose the above feature, because the technique described by Ahmed improves human perception by using mel-spectrogram, leading to higher quality and more natural sound (Ahmed, [0538]). Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Kingsbury, in view of Wilson, and further in view of Lee (US 20210027168). Regarding Claim 11, Kingsbury and Wilson disclose all the limitation of Claim 1 (see detailed mapping from above) Kingsbury and Wilson do not explicitly disclose the below cited features. Lee discloses: wherein the processor is configured to convert the utterance information into a language vector using a one-hot expression, and a number of dimensions of the one-hot expression corresponds to a number of characters contained in the utterance information. ([0043] When input data is input, the processor 120 may represent the input data as a vector (a matrix or tensor). Here, the method of representing the input data as a vector (a matrix or tensor) may vary depending on the type of the input data. For example, if a text (or a text converted from a user voice) is input as input data, the processor 120 may represent the text as a vector through One hot Encoding, or represent the text as a vector through Word Embedding. Here, the One hot encoding is a method in which only the value of the index of a specific word is represented as 1 and the value of the remaining index is represented as 0, and the Word Embedding is a method in which a word is represented as a real number in the dimension (e.g., 128 dimensions) of a vector set by a user. As an example of the Word Embedding method, Word2Vec, FastText, Glove, etc. may be used.) Kingsbury, Wilson and Lee are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Kingsbury and Wilson to combine with the teaching of Lee to disclose the above feature, because the technique described by Lee has the advantage of speed and simplicity as one hot vector encoding use low computational resource for simple task (Lee, [0043]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Scott,II US 20210209365 – discloses “The speech generator 245 can use a language model to generate text for the objects and actions identified by the object identifier 240. In one embodiment, the language model is a long short-term memory (LSTM) recurrent neural network that is trained on encoded video frames images and word-embedding that describe the objects and corresponding actions occurring in the images. The speech generator 245 can then translate the text into the audio descriptions 120.” See Abstract and para 0021 for additional details. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Phillip H Lam whose telephone number is (571)272-1721. The examiner can normally be reached 9 AM-3 PM Pacific Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on (571) 272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHILIP H LAM/ Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Show 5 earlier events
Dec 17, 2025
Response Filed
Jan 27, 2026
Final Rejection mailed — §103
Mar 31, 2026
Interview Requested
Apr 10, 2026
Examiner Interview Summary
Apr 10, 2026
Applicant Interview (Telephonic)
Apr 22, 2026
Request for Continued Examination
May 04, 2026
Response after Non-Final Action
Jul 14, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688847
ERROR-CORRECTION AND EXTRACTION IN REQUEST DIALOGS
4y 1m to grant Granted Jul 21, 2026
Patent 12682164
CHAT SUPPORT PLATFORM HAVING AUTOMATIC KEYWORD CORRECTION
3y 3m to grant Granted Jul 14, 2026
Patent 12670519
CONTENT RECOMMENDATION USING RETRIEVAL AUGMENTED ARTIFICIAL INTELLIGENCE
3y 2m to grant Granted Jun 30, 2026
Patent 12657395
METHODS AND SYSTEMS FOR AVOIDING OFFENSIVE LANGUAGE BASED ON PERSONAS
2y 9m to grant Granted Jun 16, 2026
Patent 12639529
ENHANCING LARGE LANGUAGE MODELS USING IN-CONTEXT LEARNING AND ONLINE KNOWLEDGE
2y 6m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
85%
Grant Probability
99%
With Interview (+48.0%)
2y 6m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 150 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month