Prosecution Insights
Last updated: August 06, 2026
Application No. 19/042,988

SYSTEM AND METHOD FOR AI-ASSISTED RESPONSES TO USER QUERIES

Non-Final OA §103
Filed
Jan 31, 2025
Priority
Feb 27, 2024 — provisional 63/558,590
Examiner
CAUDLE, PENNY LOUISE
Art Unit
Tech Center
Assignee
The Pamoja Institute For Community Engagement And Action
OA Round
1 (Non-Final)
68%
Grant Probability
Favorable
1-2
OA Rounds
1y 5m
Est. Remaining
85%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
53 granted / 78 resolved
+7.9% vs TC avg
Strong +17% interview lift
Without
With
+16.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
17 currently pending
Career history
95
Total Applications
across all art units

Statute-Specific Performance

§101
22.2%
-17.8% vs TC avg
§103
45.1%
+5.1% vs TC avg
§102
15.8%
-24.2% vs TC avg
§112
16.8%
-23.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 78 resolved cases

Office Action

§103
DETAILED ACTION This examination is in response to the communication filed on 01/31/2025. Claims 1-10 are currently pending, where claims 1 and 2 are independent. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 09/17/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “input module”, “voice analysis module”, “NLP module”, “trained model module” and “response module” in claim 1. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. Claims 1-10 are rejected under 35 U.S.C. 103 as being unpatentable over Song et al. (KR 20230120787 A; herein “Song”)1 further in view of Bissell et al. (US 11,551,663 B1; herein “Bissell”). Regarding claim 1, Song teaches a system for providing audio responses to audio queries, the system comprising: an input module for receiving an audio input from a user (Fig. 2, input device 120 and page 6 teaches “The processor 140 receives a user request input through an input interface (text input or voice input)…the processor 140 may receive a user’s voice and perform a voice recognition operation…”); a voice analysis module for analyzing said audio input for indicators of a mental state of said user, said voice analysis module producing state data indicative of said possible mental state based on an analysis of said audio input (Page 5 teaches “The intelligent agent module 172 can perform various natural language processing processes including…emotion analysis…”; page 6 teaches “When the voice recognition result matches a specific emotional tag among emotional tags, a guide comment according to a specific emotional state can be configured as a chatbot response”; and page 8 teaches “The processor 140 may perform a third user intention…when it is determined that the tertiary user intention determination result and voice recognition result include a keyword capable of inferring the user’s emotional state (S440: YES), the processor 140 reflects the user’s emotional state and generates an emotion guidance voice” ); an NLP module for analyzing said audio input for indicators regarding a query contained in said audio input, said NLP module producing query data indicative of said query ( page 7 teaches “the processor 140 may generate a chatbot response in which the user’s intention is reflected…operations of 1) command processing, 2) query content processing…and 4) emotion processing”); a database of prerecorded audio responses (Page 3 teaches “The database 160 may be divided into a first database for voice recognition and a second database for voice synthesis…The second database may include a voice unit corresponding to the utterance content” ); a trained model module receiving said query data and said state data, said trained model module selecting multiple of said prerecorded audio responses based on said query data and said state data (Fig. 2, AI agent 172; page 5 teaches “The intelligent agent 172 may be designed to perform at least some of the functions performed by the aforementioned ASR module 171, NLU module 173, and/or TTS module 173…The intelligent agent module 172 can perform various natural language processing processes including machine translation, emotion analysis, and information retrieval by using a deep artificial neural network structure in the field of natural language processing” ); a response module receiving prerecorded audio responses selected by said trained model module, said response module arranging and adjusting said selected audio responses to produce a final audio response based on said selected audio responses such that said final audio response approximates regular human speech (Fig. 1, response generator 14; page 3 teaches “The chatbot response may be decrypted through the NLU 11 and the DMM 12, and then an API call may be transmitted to the database 13 for necessary information…The information obtained from the database 13 through the API call is made into a chatbot response sentence in the DMM12 and sent to the message generator 14 to output the response sentence”; and page 5 teaches “speech synthesis module 174 synthesizes speech output based on the provided text…uses any suitable speech synthesis technique…including concatenative synthesis, unit selection synthesis…”). wherein said prerecorded audio responses include: short segments of audio that are interjections; long response segments of audio that are responses to specific queries; and short segments of audio that are expository in nature (Page 3 teaches “The second database may include a voice unit corresponding to the utterance content” and page 7 teaches that the utterance content, i.e., responses may include “1) command processing, 2) query content processing), 3) daily conversation processing, and 4) emotion processing” one skilled in the art would appreciate that the “utterance content” may include interjections, i.e., emotional processing, query responses, and expository, i.e., daily conversation processing). However, Song fails to disclose that the response outputs corresponding to said regular human speech has speech patterns specific to at least one of: a specific geographic region; a specific culture; and a specific ethnic group. Bissell teaches a natural language processing system that uses system response configuration data to determine customized output data forms when outputting data for a user. For example, components such as natural language generation, text-to-speech (TTS), or the like can us the customized system response configuration data to determine the form, timing, etc. of output data to be presented to a user (Bissell, Abstract) More specifically, Bissell teaches a TTS component 280 for generating responses using a unit selection engine 1230, wherein the TTS component 280 selects matching unit of recorded speech and concatenates the units together to form audio data (col. 12, lines 9-13). Column 63, lines 33-37 teaches “The TTS unit storage 1272 may include, among other things, voice inventories 1278a-1278n that may include pre-recorded audio segments (called units) to be used by the unit selection engine 1230 when performing unit selection synthesis…” Further, column 64, lines 117 teaches “…a symbolic linguistic representation, which may include linguistic context features such as…emotion, speaker, accent, or other features for processing by the speech synthesis engine 1218…The speaker feature may include data corresponding to a type of speaker, such as sex, age, or profession. The accent features may include data corresponding to an accent associated with the speaker, such as Southern, Boston, English, French, or other such accent” The accent information is interpreted as corresponding/representing a specific geographic region. Thus, Bissell teaches the TTS response output corresponding to said regular human speech has speech patterns specific to at least one of: a specific geographic region; a specific culture (the “at least one of” makes this element optional); and a specific ethnic group (the “at least one of” makes this element optional). Song differs from the invention, as defined in claim 1, in that Song fails to explicitly disclose that the human speech response output includes specific geographic region speech patterns. Synthesized speech which include specific geographic region speech patterns, i.e., accents, are known in the art as evidenced by Bissell. Therefore, it would have been obvious to one having ordinary skill in the art to have modified that unit selection synthesis taught by Song to include personality attributes such as region specific accents as taught by Bissell in order to offer more customized responses to a user in a manner that is more tailored to the particular user and/or user’s current situation (Bissell, col. 3, lines 4-7). Regarding claim 2, Song teaches a method for providing responses to user queries, the method comprising: a) receiving user input (page 6 teaches “The processor 140 receives a user request input through an input interface (text input or voice input)…the processor 140 may receive a user’s voice and perform a voice recognition operation…”); b) analyzing user input using a trained AI-based NLP based model to determine a user query in said user input (Abstract teaches “…a processor for extracting the intention of the user by analyzing a text…”; page 3 teaches “The NLU 11 may infer a user’s intention and specific requirements (entities) among the intentions based on the result of the speech language processing model (ASR)”; and page 7 teaches “the processor 140 may generate a chatbot response in which the user’s intention is reflected…operations of 1) command processing, 2) query content processing…and 4) emotion processing” ); c) producing query data relating to said user query based on results of step b) (page 3 teaches “The NLU 11 may infer a user’s intention and specific requirements (entities) among the intentions based on the result of the speech language processing model (ASR)”; and page 7 teaches “the processor 140 may generate a chatbot response in which the user’s intention is reflected…operations of 1) command processing, 2) query content processing…and 4) emotion processing” ); d) analyzing query data using a trained AI-based model to determine response data, said response data being suitable for said user query (abstract teaches “a processor for extracting the intention of the user by analyzing a text extracted through the voice recognition part, and generating a chatbot response having the intention of the user reflected thereon”; and page 3 teaches “The chatbot response may be decrypted through the NLU 11 and the DMM 12, and then an API call may be transmitted to the database 13 for necessary information…” ); e) based on said response data, selecting one or more prerecorded pieces of spoken audio, at least one of said one or more pieces of spoken audio being related to said user query (page 3 teaches “The chatbot response may be decrypted through the NLU 11 and the DMM 12, and then an API call may be transmitted to the database 13 for necessary information… The information obtained from the database 13 through the API call is made into a chatbot response sentence in the DMM 12 and sent to the message generator 14 to output the response sentence” and page 5 teaches “Speech synthesis module 174 uses any suitable speech synthesis technique… including concatenative synthesis, unit selection synthesis…” One skill in the art would appreciate the “unit selection synthesis” is a known process which includes concatenating together pre-recorded segments of spoken audio.); f) arranging and adjusting said one or more prerecorded pieces of spoken audio to result in an audio response that, when played, approximates regular human speech (page 5 teaches “Speech synthesis module 174 uses any suitable speech synthesis technique… including concatenative synthesis, unit selection synthesis…” One skill in the art would appreciate the “unit selection synthesis” is a known process which includes concatenating together pre-recorded segments of spoken audio and page 2 teaches “One of the most important tasks in this emotion-sharing chatbot technology is make the user feel as if they are having a natural conversation with a real person…”); g) providing said audio response to said user such that said user hears said audio response (page 3 teaches “…the chatbot system 10 may include a chatbot response output interface, and the chatbot response output interface may include an audio output unit and a display unit. Corresponding information (conversation) can be extracted through the Dialogue Management Model (DMM), and a chatbot response can be output to the speaker with voice or text”). However, Song fails to disclose that the response outputs corresponding to said regular human speech has speech patterns specific to at least one of: a specific geographic region; a specific culture; and a specific ethnic group. Bissell teaches a natural language processing system that uses system response configuration data to determine customized output data forms when outputting data for a user. For example, components such as natural language generation, text-to-speech (TTS), or the like can us the customized system response configuration data to determine the form, timing, etc. of output data to be presented to a user (Bissell, Abstract) More specifically, Bissell teaches a TTS component 280 for generating responses using a unit selection engine 1230, wherein the TTS component 280 selects matching unit of recorded speech and concatenates the units together to form audio data (col. 12, lines 9-13). Column 63, lines 33-37 teaches “The TTS unit storage 1272 may include, among other things, voice inventories 1278a-1278n that may include pre-recorded audio segments (called units) to be used by the unit selection engine 1230 when performing unit selection synthesis…” Further, column 64, lines 117 teaches “…a symbolic linguistic representation, which may include linguistic context features such as…emotion, speaker, accent, or other features for processing by the speech synthesis engine 1218…The speaker feature may include data corresponding to a type of speaker, such as sex, age, or profession. The accent features may include data corresponding to an accent associated with the speaker, such as Southern, Boston, English, French, or other such accent” The accent information is interpreted as corresponding/representing a specific geographic region. Thus, Bissell teaches the TTS response output corresponding to said regular human speech has speech patterns specific to at least one of: a specific geographic region; a specific culture (the “at least one of” makes this element optional); and a specific ethnic group (the “at least one of” makes this element optional). Song differs from the invention, as defined in claim 2, in that Song fails to explicitly disclose that the human speech response output includes specific geographic region speech patterns. Synthesized speech which include specific geographic region speech patterns, i.e., accents, are known in the art as evidenced by Bissell. Therefore, it would have been obvious to one having ordinary skill in the art to have modified that unit selection synthesis taught by Song to include personality attributes such as region specific accents as taught by Bissell in order to offer more customized responses to a user in a manner that is more tailored to the particular user and/or user’s current situation (Bissell, col. 3, lines 4-7). Regarding claim 3, the combination of Song and Bissell teaches all of the elements of claim 2 (see detailed element mapping above). In addition, Song further teaches said user input is spoken audio (page 6 teaches “The processor 140 receives a user request input through an input interface (text input or voice input)…the processor 140 may receive a user’s voice and perform a voice recognition operation…”). Regarding claim 4, the combination of Song and Bissell teaches all of the elements of claim 2 (see detailed element mapping above). In addition, Song further teaches said user input is textual input (page 6 teaches “The processor 140 receives a user request input through an input interface (text input or voice input)…”). Regarding claim 5, the combination of Song and Bissell teaches all of the elements of claim 2 (see detailed element mapping above). In addition, Song further teaches step d) includes analyzing query data for other user inputs from a same user and producing response data based on query data for a current user input and based on query data from previous user inputs from said same user (Page 3 teaches “The chatbot response may be decrypted through the NLU 11 and the DMM 12, and then an API call may be transmitted to the database 13 for necessary information. In this case, when deciphering the user's request sentence, the past conversation information, the processed response of the chatbot, and the API search information result may be additionally utilized”). Regarding claim 6, the combination of Song and Bissell teaches all of the elements of claim 2 (see detailed element mapping above). In addition, Song further teaches wherein said response data determined in step d) is also based on previous response data generated for said previous user inputs from said same user (Page 3 teaches “The chatbot response may be decrypted through the NLU 11 and the DMM 12, and then an API call may be transmitted to the database 13 for necessary information. In this case, when deciphering the user's request sentence, the past conversation information, the processed response of the chatbot, and the API search information result may be additionally utilized”). Regarding claim 7, the combination of Song and Bissell teaches all of the elements of claim 3 (see detailed element mapping above). In addition, Song further teaches prior to step c), said user input is analyzed by a voice analysis module for indicators of a possible mental state of said user, said voice analysis module producing state data indicative of said possible mental state based on an analysis of said spoken audio (Page 5 teaches “The intelligent agent module 172 can perform various natural language processing processes including…emotion analysis…”; page 6 teaches “When the voice recognition result matches a specific emotional tag among emotional tags, a guide comment according to a specific emotional state can be configured as a chatbot response”; and page 8 teaches “The processor 140 may perform a third user intention…when it is determined that the tertiary user intention determination result and voice recognition result include a keyword capable of inferring the user’s emotional state (S440: YES), the processor 140 reflects the user’s emotional state and generates an emotion guidance voice”). Regarding claim 8, the combination of Song and Bissell teaches all of the elements of claim 7 (see detailed element mapping above). In addition, Song further teaches step d) includes analyzing said state data with said query data to produce said response data (Page 7 teaches “When a keyword representing an emotional state is included in the voice recognition result and text, the processor 140 recognizes the user's intention as outputting a chatbot response corresponding to the emotional state of a specific state, and prepares a previously prepared (or based on the voice recognition result) Corresponding) can generate a chatbot response” In other words, the response is generated based on both the emotional state and the request). Regarding claim 9, the combination of Song and Bissell teaches all of the elements of claim 8 (see detailed element mapping above). In addition, Song further teaches wherein for step e), at least one of said one or more pieces of spoken audio is related to said possible mental state (page 8 teaches “The processor 140 may perform a third user intention…when it is determined that the tertiary user intention determination result and voice recognition result include a keyword capable of inferring the user’s emotional state (S440: YES), the processor 140 reflects the user’s emotional state and generates an emotion guidance voice” The voice recognition is based on the spoken audio input, thus the mental state is based on the spoken audio). Regarding claim 10, the combination of Song and Bissell teaches all of the elements of claim 1 (see detailed element mapping above). In addition, Song further teaches said system executes a method for providing responses to user queries, the method comprising: a) receiving user input (page 6 teaches “The processor 140 receives a user request input through an input interface (text input or voice input)…the processor 140 may receive a user’s voice and perform a voice recognition operation…”); b) analyzing user input using a trained AI-based NLP based model to determine a user query in said user input (Abstract teaches “…a processor for extracting the intention of the user by analyzing a text…”; page 3 teaches “The NLU 11 may infer a user’s intention and specific requirements (entities) among the intentions based on the result of the speech language processing model (ASR)”; and page 7 teaches “the processor 140 may generate a chatbot response in which the user’s intention is reflected…operations of 1) command processing, 2) query content processing…and 4) emotion processing”); c) producing query data relating to said user query based on results of step b) (page 3 teaches “The NLU 11 may infer a user’s intention and specific requirements (entities) among the intentions based on the result of the speech language processing model (ASR)”; and page 7 teaches “the processor 140 may generate a chatbot response in which the user’s intention is reflected…operations of 1) command processing, 2) query content processing…and 4) emotion processing” ); d) analyzing query data using a trained AI-based model to determine response data, said response data being suitable for said user query (abstract teaches “a processor for extracting the intention of the user by analyzing a text extracted through the voice recognition part, and generating a chatbot response having the intention of the user reflected thereon”; and page 3 teaches “The chatbot response may be decrypted through the NLU 11 and the DMM 12, and then an API call may be transmitted to the database 13 for necessary information…” ); e) based on said response data, selecting one or more prerecorded pieces of spoken audio, at least one of said one or more pieces of spoken audio being related to said user query (page 3 teaches “The chatbot response may be decrypted through the NLU 11 and the DMM 12, and then an API call may be transmitted to the database 13 for necessary information… The information obtained from the database 13 through the API call is made into a chatbot response sentence in the DMM 12 and sent to the message generator 14 to output the response sentence” and page 5 teaches “Speech synthesis module 174 uses any suitable speech synthesis technique… including concatenative synthesis, unit selection synthesis…” One skill in the art would appreciate the “unit selection synthesis” is a known process which includes concatenating together pre-recorded segments of spoken audio.); f) arranging and adjusting said one or more prerecorded pieces of spoken audio to result in an audio response that, when played, approximates regular human speech (page 5 teaches “Speech synthesis module 174 uses any suitable speech synthesis technique… including concatenative synthesis, unit selection synthesis…” One skill in the art would appreciate the “unit selection synthesis” is a known process which includes concatenating together pre-recorded segments of spoken audio and page 2 teaches “One of the most important tasks in this emotion-sharing chatbot technology is make the user feel as if they are having a natural conversation with a real person…” ); g) providing said audio response to said user such that said user hears said audio response (page 3 teaches “…the chatbot system 10 may include a chatbot response output interface, and the chatbot response output interface may include an audio output unit and a display unit. Corresponding information (conversation) can be extracted through the Dialogue Management Model (DMM), and a chatbot response can be output to the speaker with voice or text” ). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Smith et al. (US 11,893,984 B1) teaches a speech processing system that utilizes pre-recorded audio segments for TTS; Ingel et al. (US 2020/0169591 A1) teaches a system and method for artificial dubbing which, among other things, utilizes preferred language characteristics such as accent or specific dialects; and Frost et al. (US 2022/0382988) teaches a system and method for generating automated conversation responses. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PENNY L CAUDLE whose telephone number is (703)756-1432. The examiner can normally be reached M-Th 8:00 am to 5:00 pm eastern. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PENNY L CAUDLE/Examiner, Art Unit 2657 /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657 1 All citations to Song reference the machine English Translation provided here with.
Read full office action

Prosecution Timeline

Jan 31, 2025
Application Filed
Jul 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688859
Coding and Decoding of Spherical Coordinates Using an Optimized Spherical Quantization Dictionary
1y 11m to grant Granted Jul 21, 2026
Patent 12682905
VOICE DATA TRANSMISSION METHOD AND APPARATUS
2y 6m to grant Granted Jul 14, 2026
Patent 12682166
AUTO-SUGGESTION WITH RICH OBJECTS
2y 7m to grant Granted Jul 14, 2026
Patent 12682915
System and Method for Generating Brand Standards from Voice Input
2y 3m to grant Granted Jul 14, 2026
Patent 12675646
Optimal Subword Tokenization and Vocabulary Creation
2y 5m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
68%
Grant Probability
85%
With Interview (+16.7%)
2y 11m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 78 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month