Prosecution Insights
Last updated: August 17, 2026
Application No. 18/969,979

Apparatus and Method and for Correcting Result of Speech Recognition by Using Camera

Non-Final OA §102§103
Filed
Dec 05, 2024
Priority
Apr 05, 2024 — RE 10-2024-0046339
Examiner
SONIFRANK, RICHA MISHRA
Art Unit
Tech Center
Assignee
Kia Corporation
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
261 granted / 391 resolved
+6.8% vs TC avg
Strong +25% interview lift
Without
With
+24.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
25 currently pending
Career history
416
Total Applications
across all art units

Statute-Specific Performance

§101
15.7%
-24.3% vs TC avg
§103
62.6%
+22.6% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
8.4%
-31.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 391 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Detailed Action The office action sent in response to Applicant’s communication received on 12/5/2024 for the application number 18969979. The office hereby acknowledges receipt of the following placed of record in the file: Specification, Abstract, Oath/Declaration and claims. Status of Claims Claims 1-14 are presented for examination. Examiner’s Remark Rejection under 101 is not applicable since the claims are directed at operating a vehicle using technological steps which are not well-known, routine or conventional. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1 and 8 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Lenke (US 20200216086) Regarding claim 1, Lenke teaches a speech recognition apparatus comprising: memory storing instructions; and at least one processor, wherein the at least one processor, by executing the instructions, is configured to: receive, via a microphone in a vehicle, an utterance spoken by a user of the vehicle ( microphone to receive voice command, Para 0007) ; identify, based on one or more images received from a camera of the vehicle, context information indicating: an action of the user while speaking the utterance ( for e.g. user can say , the occupant may spot a building or other point of interest outside the vehicle, and either describe it verbally (e.g., “the red building on the left-hand side”), or by looking at it (e.g., “that building”), and ask the vehicle to let him/her out there. The vehicle may employ gaze detection (e.g., using camera/gaze detection 114) and a two- or three-dimensional representation of the vehicle's vicinity (e.g., using vehicle position manager 126 of vehicle autonomous operation manager 120) to identify the point of interest identified by the occupant and translate it to an address which may then be used (e.g., by navigation manager 124) to establish a new waypoint or destination., Para 0068) , and an object associated with the action ( building, Para 0068) ; identify, based on performing speech recognition on the utterance, an intent of the utterance ( for e.g. drop of location, Para 0068, Fig 6) ; identify, based on the intent and based on a sentence type associated with the utterance ( sentence type can be a new location or a question or a full sentence or sentence containing ambiguity for e.g. “that building”, Para 0068, Fig 6) , an ambiguity associated with the utterance ( for e.g. that building, Para 0068) ; adjust, based on the ambiguity and the context information, a result of the speech recognition ( results is based on user’s gaze and the speech, Para 0068) ; and control, based on the adjusted result of the speech recognition, an operation of the vehicle ( for e.g. stopping at a particular location based on gaze and speech etc., Para 0068) Regarding claim 8, arguments analogous to claim 1, are applicable. In addition, Lenke teaches speech recognition method (abstract, Fig 1) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. And KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Exemplary rationales that may support a conclusion of obviousness include: (A) Combining prior art elements according to known methods to yield predictable results; (B) Simple substitution of one known element for another to obtain predictable results; (C) Use of known technique to improve similar devices (methods, or products) in the same way; (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results; (E) "Obvious to try" – choosing from a finite number of identified, predictable solutions, with a reasonable expectation of success; (F) Known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art; (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention; See MPEP § 2143 for a discussion of the rationales listed above along with examples illustrating how the cited rationales may be used to support a finding of obviousness. See also MPEP § 2144 - § 2144.09 for additional guidance regarding support for obviousness determination. Claims 2 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Lenke (US 20200216086) and further in view of Zhao (US 11507346) Regarding claim 2, Lenke as above in claim 1, teach wherein the at least one processor is configured to identify the context information by: identifying the context information further based on action database ( determine intent or location based on vehicle operation manager, Fig 1, Para 0035; The occupant context model 162, and information received by autonomous vehicle operation manager 120, may be used by voice assistant 160 to define voice output to the occupant (e.g., intent options, dialogue content, voice quality, intonation, speech style, etc.) via TTS 106. For example, voice assistant 160 may have access to a plurality of intent, dialogue, voice quality, intonation and/or speech style options, select certain options based on current vehicle context information, current occupant context automation, and/or associations between the two, and use the selected options to produce voice output to the occupant. ….. The information gathered by voice assistant 160 during interaction with the occupant may be provided to autonomous vehicle operation manager 120 for use in controlling operation of the vehicle, such as to complete an action or task specified via voice input., Para 0028) Lenke does not explicitly teach identifying the context information further based on a core action priority database and a core action-free database However, Zhao teaches identifying the context information further based on a core action priority database and a core action-free database (At block 116, the controller 34 determines consolidated confidence scores for each of the multiple probable speech recognition results. To do so, the controller 34 accesses a database storing a priori information at block 118. The a priori information may include, but is not limited, to navigation information (e.g., frequent points of interest (POIs) or nearby POIs); phone information (e.g., frequent calls and/or text recipients), and/or controls information (frequent controls such as climate, liftgate, screen brightness, etc.)., Col 7, line 10-20) It would have been obvious to POSITA having the teachings of Lenke to further include the concept of Zhao before effective filing date to have different separate databases so to access the database based on the command (col 7, line 15-30, Zhao) Regarding claim 9, arguments analogous to claim 2, are applicable. Claims 3 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Lenke (US 20200216086) and further in view of Zhao (US 11507346) and further in view of Garcia (US 20210027780) Regarding claim 3, Lenke modified by Zhao as above in claim 2, does not teach wherein the at least one processor is configured to identify the context information by: identifying, based on the user performing a plurality of actions, the action according to the core action priority database, and wherein the core action priority database indicates a higher priority for a more specific action of the plurality of actions However, Garcia teaches wherein the at least one processor is configured to identify the context information by: identifying, based on the user performing a plurality of actions, the action according to the core action priority database, and wherein the core action priority database indicates a higher priority for a more specific action of the plurality of actions ( if there is a plurality of actionable intents, the virtual assistant selects, from a plurality of tasks associated with the plurality of actionable intents, a single task to perform based on a priority associated with each of the plurality of tasks, Para 0304) It would have been obvious to POSITA before effective filing date having the teachings of Lenke and Zhao to further include the concept of Garcia before effective filing date so to perform the more important and time sensitive task first. Regarding claim 10, arguments analogous to claim 3, are applicable. Claims 4-6 and 11-13 are rejected under 35 U.S.C. 103 as being unpatentable over Lenke (US 20200216086) and further in view of Moniri (US 20200219320) Regarding claim 4, Lenke as above in claim 1, teach wherein the at least one processor is configured to identify the ambiguity by: identify the ambiguity further based on the intent being out-of-domain and the utterance comprising a demonstrative word (for e.g. “that building”, Para 0068) Lenke does not explicitly teach ambiguity includes demonstrative pronoun However, Moniri teach ambiguity includes demonstrative pronoun ( When the natural language processor 212 determines (e.g., from detection of a term that indicates a position or object, such as a demonstrative pronoun or a place adverb) that the speech input contains an ambiguity that may be resolved with input received via another modality, the multimodal interface 206 will obtain from the gaze detection interface 216 and/or the gesture recognition interface 218 geometric input that was provided via one (or both) of those interfaces., Para 0058) It would have been obvious to POSITA having the teachings of Lenke to further include the concept of Moniri before effective filing date to make communication system more convenient for the user since user can speak short phrases. Regarding claim 5, Lenke as above in claim 1, teach wherein the at least one processor is configured to identify the ambiguity by: identify the ambiguity further based on the intent being out-of-domain and the utterance only containing an adverb or a predicate (and ask the vehicle to let him/her out there. The vehicle may employ gaze detection (e.g., using camera/gaze detection 114) and a two- or three-dimensional representation of the vehicle's vicinity (e.g., using vehicle position manager 126 of vehicle autonomous operation manager 120) to identify the point of interest identified by the occupant and translate it to an address which may then be used (e.g., by navigation manager 124) to establish a new waypoint or destination. For example, the occupant may specify that the point of interest is to be the new destination by saying “please let me out there” (e.g., received by voice assistant via ASR/NLU 104), and the vehicle may identify (e.g., using navigation manager 124) the closest safe place for letting the occupant leave the vehicle., Para 0068- where “there” is an adverb) Lenke does not explicitly teach utterance containing only adverbs or predicate However, Moniri teaches utterance containing only adverbs or predicate ( the other input modalities of the multimodal interface 206 include a gaze detection interface 216 and a gesture recognition interface 218. When the natural language processor 212 determines (e.g., from detection of a term that indicates a position or object, such as a demonstrative pronoun or a place adverb) that the speech input contains an ambiguity that may be resolved with input received via another modality, the multimodal interface 206 will obtain from the gaze detection interface 216 and/or the gesture recognition interface 218 geometric input that was provided via one (or both) of those interfaces, Para 0058- place adverb) It would have been obvious to POSITA having the teachings of Lenke to further include the concept of Moniri before effective filing date to make communication system more convenient for the user. Regarding claim 6, Lenke as above in claim 1, teach wherein the at least one processor is configured to adjust the result of the speech recognition by: receiving, based on the user performing the plurality of actions, additional context information indicating a second object associated with a second action; and adjusting the result of the speech recognition based on the additional context information ( system is able to respond to multiple questions and use gaze as an additional context for questions, 0027-0028) Lenke does not teach the concept of determining whether the user is performing a plurality of actions However, Moniri teach determining whether the user is performing a plurality of actions ( a user could speak “Send that building's address to Kevin” from which the multimodal interface could determine that a message is to be sent to “Kevin” (who might be identified from the user's contact list), and then determine from the non-speech input (e.g., as described above, such as using a direction of a user's gaze or gesture) that the user is referring to a building that was in the environment of the automobile at the time of the speech input. When such a task is performed, the system may not generate any output to be displayed., Para 0033); receiving, based on the user performing the plurality of actions, additional context information indicating a second object associated with a second action ( for e.g. address of “that” building, Para 0033) ; and adjusting the result of the speech recognition based on the additional context information ( result is based on the finding the contact list and determining the address of the building, Para 0033) It would have been obvious for POSITA having the teachings Lenke to further include the concept of Moniri before effective filing date because the it reduce an amount of time and/or attention needed for the interaction for the user (Para 0027, Moniri) Regarding claim 11, arguments analogous to claim 4, are applicable. Regarding claim 12, arguments analogous to claim 5, are applicable. Regarding claim 13, arguments analogous to claim 6, are applicable. Claims 7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Lenke (US 20200216086) and further in view of Moniri (US 20200219320) and further in view of Kim (US 20210139036) Regarding claim 7, Moniri as above in claim 6, teach wherein the at least one processor is configured to: output, based on the intent of the utterance being out-of-domain with respect to the result of the speech recognition and the adjusted result of the speech recognition, a notification that the intent of the utterance corresponds to an missing feature ( the task facility determines, based on output from the processing of the speech input, whether the speech input includes ambiguities regarding the task to be performed, such that the task is unclear or cannot be performed. If so, then in block 308 the task facility interacts with the user in block 308 regarding the ambiguities to obtain additional information, Para 0070) Lenke modified by Moniri does not teach a notification that the intent of the utterance corresponds to an unsupported feature However, Kim teaches a notification that the intent of the utterance corresponds to an unsupported feature (When any command and/or information on a function associated with the any command is not stored in memory 140, the first and second voice recognition engines 810, 820 may not output any information in response to the any command. In this case, the first and second voice recognition engines 810, 820 may output guide information indicating that a function (or information) corresponding to the command cannot be executed through the output unit 850., Para 0377) It would have been obvious to POSITA having the teaching of Lenke and Moniri to further include the concept of Kim before effective filing so to inform the user. Regarding claim 14, arguments analogous to claim 7, are applicable. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Park (US 20220198151) discloses The dialogue system includes a storage configured to store target information about a target and a target value for ambiguous language; a first input device configured to receive speech signals; and a dialogue manager configured to: convert the speech signals received in the first input device into text; determine a user's intention based on the received speech signals; and based on determining that the determined user's intention corresponds to a request intention and the converted text corresponds to the ambiguous language, obtain the target and the target value corresponding to the ambiguous language from the target information stored in the storage. The dialogue system also includes a result processor configured to generate a response based on the target and the target value obtained from the dialogue manager, and to control an output of the generated response. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Richa Sonifrank whose telephone number is (571)272-5357. The examiner can normally be reached M-T 7AM - 5:30PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Phan Hai can be reached at (571)272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Richa Sonifrank/Primary Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Dec 05, 2024
Application Filed
Aug 04, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705432
NATURAL LANGUAGE GOAL DETERMINATION
3y 1m to grant Granted Aug 11, 2026
Patent 12676156
ENCODING DEVICE AND ENCODING METHOD, DECODING DEVICE AND DECODING METHOD, AND PROGRAM
2y 8m to grant Granted Jul 07, 2026
Patent 12664183
ONLINE QUESTION ANSWERING, USING READING COMPREHENSION WITH AN ENSEMBLE OF MODELS
4y 11m to grant Granted Jun 23, 2026
Patent 12664973
VOICE DIALOGUE PROCESSING METHOD AND APPARATUS
4y 4m to grant Granted Jun 23, 2026
Patent 12645879
ENTITY RECOGNITION METHODS AND APPARATUSES, ELECTRONIC DEVICES AND STORAGE MEDIA
3y 4m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
92%
With Interview (+24.7%)
3y 0m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 391 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month