Prosecution Insights
Last updated: August 17, 2026
Application No. 18/958,655

ACCELEROMETER-BASED ENDPOINTING MEASURE(S) AND /OR GAZE-BASED ENDPOINTING MEASURE(S) FOR SPEECH PROCESSING

Non-Final OA §103§112§DP
Filed
Nov 25, 2024
Priority
Dec 17, 2021 — continuation of 12/154,561
Examiner
ZHANG, LESHUI
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
740 granted / 950 resolved
+17.9% vs TC avg
Strong +35% interview lift
Without
With
+35.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
28 currently pending
Career history
984
Total Applications
across all art units

Statute-Specific Performance

§101
5.7%
-34.3% vs TC avg
§103
44.5%
+4.5% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
29.3%
-10.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 950 resolved cases

Office Action

§103 §112 §DP
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the response to this office action, the Examiner respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Examiner in prosecuting this application. Claim Objections Claims 12, 14-19 are objected to because of the following informalities: Claim 12 recites “rendering content that is …” which should be -- rendering a content that is …--. Claim 14 recites “one or more microphones of device” which should be -- one or more microphones of a device --. Claims 15-19 are objected due to the dependencies to claim 14. Claim 18 is objected for the at least similar reason as described in claim 12 above because claim 18 recites the similar deficient limitation as recited in claim 12. Appropriate correction is required. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/ patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/ patents/apply/applying-online/eterminal-disclaimer. Claims 1-20 rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-16 of U.S. Patent No. 12,154,561 B2 in view of reference Konzelmann et al. (US 20200349966 A1) and Piernot et al. (US 20150348548 A1). The conflicting claims 1-16 of U.S. Patent No. 12,154,561 B2) teaches all the limitations of claims 1-20 of the instant application, except explicitly teaching an order of the implementation among processing audio-data, processing the stream of accelerometer data, and processing the stream of the image data, e.g., the audio data processing is performed before processing the accelerometer data and processing the gaze data as recited in claim 2, or processing the stream of accelerometer data/image data in response to determining the audio based endpointing measure with an initial threshold as recited in claims 3-4, 7, and rep-fetching the content responsive to determining that the audio-based endpointing measure as recited in claim 13, etc., However, the combination of Konzelmann (above) and Piernot (above) teaches these limitations of the claims 2-7, 13, etc. above, as discussed in prior art rejection as set forth below, for the benefits discussed in prior art rejection as set forth below. Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to applied these limitations above, as taught by the combination of Konzelmann and Piernot above, to processing audio data, processing the stream of accelerometer data, and processing the gaze data in the method, as taught by the conflicting claims 1-16 of U.S. Patent No. 12,154,561 B2), for the benefits discussed above. The following is the comparison between the claims 1-20 of the instant application and the conflicting claim 1-16 of U.S. Patent No. 12,154,561 B2 for reference: Claims in the current application Conflicting claims 1-10 in US Patent No. 12,154,561 B2 1. A method implemented by one or more processors, the method comprising: processing an audio data stream, using a machine learning model, to generate an audio-based endpointing measure, wherein the audio data stream captures a spoken utterance of a user and is detected via one or more microphones of a client device; processing a stream of accelerometer data, using the machine learning model, to generate an accelerometer-based endpointing measure and/or processing a stream of image data, using the machine learning model, to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device; determining an overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure; determining whether the overall endpointing measure satisfies a threshold; and in response to determining the overall endpointing measure satisfies the threshold: performing one or more actions based on the spoken utterance. 2. The method of claim 1, further comprising: prior to processing the stream of accelerometer data to generate the accelerometer-based endpointing measure and/or prior to processing the stream of image data to generate the gaze-based endpointing measure, determining whether the audio-based endpointing measure satisfies an initial threshold indicating a candidate endpoint in the audio data stream. 3. The method of claim 2, wherein processing the stream of accelerometer data, using the machine learning model, to generate the accelerometer-based endpointing measure further comprises: processing the stream of accelerometer data in response to determining the audio-based endpointing measure satisfies the initial threshold indicating the candidate endpoint in the audio data stream. 4. The method of claim 2, wherein processing the stream of image data, using the machine learning model, to generate the gaze-based endpointing measure further comprises: processing the stream of image data in response to determining the audio-based endpointing measure satisfies the initial threshold indicating the candidate endpoint in the audio data stream. 5. The method of claim 1, wherein the accelerometer-based endpointing measure classifies movement of the client device. 6. The method of claim 5, wherein processing the stream of accelerometer data, using the machine learning model, to generate the accelerometer-based endpointing measure comprises: processing, using the machine learning model model, (i) a portion of the stream of accelerometer data captured prior to the user speaking the spoken utterance and (ii) a portion of the stream of accelerometer data captured subsequent to the user beginning to speak the spoken utterance, to generate the accelerometer-based endpointing measure. 7. The method of claim 2, wherein processing the stream of image data, using the machine learning model, to generate the gaze-based endpointing measure further comprises: processing the stream of accelerometer data in response to determining the audio-based endpointing measure satisfies the initial threshold indicating the candidate endpoint in the audio data stream. 8. The method of claim 1, wherein determining the overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure comprises: boosting, based on the accelerometer-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold and/or boosting, based on the gaze-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold. 9. The method of claim 1, wherein the accelerometer-based endpointing measure indicates no movement of the client device, and wherein determining the overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure comprises: decreasing, based on the accelerometer-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold. 10. The method of claim 1, wherein the gaze-based endpointing measure indicates the user is not looking at the client device, and wherein determining the overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure comprises: decreasing, based on the gaze-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold. 11. The method of claim 1, wherein performing the one or more actions based on the spoken utterance comprises: processing the spoken utterance using an automatic speech recognition model to generate a text representation of the spoken utterance. 12. The method of claim 1, wherein performing the one or more actions based on the spoken utterance comprises: rendering content that is responsive to the spoken utterance. 13. The method of claim 1, further comprising: determining, prior to determining the overall endpointing measure satisfies the threshold, that the audio-based endpointing measure satisfies the threshold or an alternate threshold; and pre-fetching the content responsive to determining that the audio-based endpointing measure satisfies the threshold or the alternate threshold. 14. A client device, comprising: one or more processors, and memory configured to store instructions that, when executed by the one or more processors, cause the one or more processors to perform a method that includes: processing, using a machine learning model: an audio data stream, where the audio data stream captures a spoken utterance of a user and is detected via one or more microphones of device; and a stream of accelerometer data and/or a stream of image data, where the stream of accelerometer data classifies movement of the client device, and where the stream of image data classifies the gaze of the user of the client device; determining, based on processing the audio data stream and the stream of accelerometer data and/or the stream of image data, an endpointing measure indicating the likelihood of a candidate endpoint in the audio data stream; determining whether the endpointing measure satisfies a threshold; and in response to determining the endpointing measure satisfies a threshold, performing one or more actions based on the spoken utterance. 15. The client device of claim 14, wherein processing, using the machine learning model, the audio data stream and the stream of accelerometer data and/or stream of image data comprises: processing the audio data stream, using the machine learning model, to generate an audio-based endpointing measure; and processing the stream of accelerometer data, using the machine learning model, to generate an accelerometer-based endpointing measure and/or processing the stream of image data, using the machine learning model, to generate a gaze-based endpointing measure. 16. The client device of claim 15, determining whether the endpointing measure satisfies a threshold comprises: determining whether the audio-based endpointing measure satisfies an initial threshold; in response to determining the audio-based endpointing measure satisfies the initial threshold: boosting, based on the accelerometer-based endpointing measure, the likelihood the endpointing measure satisfies the threshold and/or boosting, based on the gaze-based endpointing measure, the likelihood the endpointing measure satisfies the threshold. 17. The client device of claim 14, wherein performing the one or more actions based on the spoken utterance comprises: processing the spoken utterance using an automatic speech recognition model to generate a text representation of the spoken utterance. 18. The client device of claim 14, wherein performing the one or more actions based on the spoken utterance comprises: rendering content that is responsive to the spoken utterance. 19. The client device of claim 14, wherein the instructions cause the one or more processors to perform the method that includes: determining, prior to determining the endpointing measure satisfies the threshold, that the audio-based endpointing measure satisfies the threshold or an alternate threshold; and pre-fetching the content responsive to determining that the audio-based endpointing measure satisfies the threshold or the alternate threshold. 20. A system comprising: one or more processors; and memory configured to store instructions that, when executed by the one or more processors cause the one or more processors to perform operations that include: processing an audio data stream, using a machine learning model, to generate an audio-based endpointing measure, wherein the audio data stream captures a spoken utterance of a user and is detected via one or more microphones of a client device; processing a stream of accelerometer data, using the machine learning model, to generate an accelerometer-based endpointing measure and/or processing a stream of image data, using the machine learning model, to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device; determining an overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure; determining whether the overall endpointing measure satisfies a threshold; and in response to determining the overall endpointing measure satisfies the threshold: performing one or more actions based on the spoken utterance. 1. A method implemented by one or more processors, the method comprising: processing an audio data stream, using an audio-based endpointer model, to generate an audio-based endpointing measure, wherein the audio data stream captures a spoken utterance of a user and is detected via one or more microphones of a client device; determining whether the audio-based endpointing measure satisfies an initial threshold indicating a candidate endpoint in the audio data stream, wherein the initial threshold, when satisfied, is indicative of a passage of a threshold duration of time since the user stopped speaking; in response to determining the audio-based endpointing measure satisfies the initial threshold: processing a stream of accelerometer data, using an accelerometer model, to generate an accelerometer-based endpointing measure; determining an overall endpointing measure as a function of both the audio-based endpointing measure and the accelerometer-based endpointing measure; determining whether the overall endpointing measure satisfies a threshold; and in response to determining the overall endpointing measure satisfies the threshold: performing one or more actions based on the spoken utterance. 6. The method of claim 1, further comprising: processing a stream of image data, using a gaze model, to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device. 6. The method of claim 1, further comprising: processing a stream of image data, using a gaze model, to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device. 2. The method of claim 1, wherein the accelerometer-based endpointing measure classifies movement of the client device. 3. The method of claim 2, wherein processing the stream of accelerometer data, using the accelerometer model, to generate the accelerometer-based endpointing measure comprises: processing, using the accelerometer model, (i) a portion of the stream of accelerometer data captured prior to the user speaking the spoken utterance and (ii) a portion of the stream of accelerometer data captured subsequent to the user beginning to speak the spoken utterance, to generate the accelerometer-based endpointing measure. 6. The method of claim 1, further comprising: processing a stream of image data, using a gaze model, to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device. 4. The method of claim 1, wherein determining the overall endpointing measure as a function of both the audio-based endpointing measure and the accelerometer-based endpointing measure comprises: boosting, based on the accelerometer-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold. 5. The method of claim 1, wherein the accelerometer-based endpointing measure indicates no movement of the client device, and wherein determining the overall endpointing measure as a function of both the audio-based endpointing measure and the accelerometer-based endpointing measure comprises: decreasing, based on the accelerometer-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold. 6. The method of claim 1, further comprising: processing a stream of image data, using a gaze model, to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device. 7. The method of claim 6, wherein determining the overall endpointing measure is further a function of the gaze-based endpointing measure. 8. The method of claim 1, wherein performing the one or more actions based on the spoken utterance comprises: processing the spoken utterance using an automatic speech recognition model to generate a text representation of the spoken utterance. 9. The method of claim 1, wherein performing the one or more actions based on the spoken utterance comprises: rendering content that is responsive to the spoken utterance. 10. The method of claim 9, further comprising: determining, prior to determining the overall endpointing measure satisfies the threshold, that the audio-based endpointing measure satisfies the threshold or an alternate threshold; and pre-fetching the content responsive to determining that the audio-based endpointing measure satisfies the threshold or the alternate threshold. 11. A method implemented by one or more processors, the method comprising: processing an audio data stream, using an audio-based endpointer model, to generate an audio-based endpointing measure, wherein the audio data stream captures a spoken utterance of a user and is detected via one or more microphones of a client device; determining whether the audio-based endpointing measure satisfies an initial threshold indicating a candidate endpoint in the audio data stream, wherein the initial threshold, when satisfied, is indicative of a passage of a threshold duration of time since the user stopped speaking; in response to determining the audio-based endpointing measure satisfies the initial threshold: processing a stream of image data, using a gaze model, to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device; determining an overall endpointing measure as a function of both the audio-based endpointing measure and the gaze-based endpointing measure; determining whether the overall endpointing measure satisfies a threshold; and in response to determining the overall endpointing measure satisfies the threshold: performing one or more actions based on the spoken utterance. 12. The method of claim 11, wherein determining the overall endpointing measure as a function of both the audio-based endpointing measure and the gaze-based endpointing measure comprises: boosting, based on the gaze-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold. 13. The method of claim 11, wherein the gaze-based endpointing measure indicates the user is not looking at the client device, and wherein determining the overall endpointing measure as a function of both the audio-based endpointing measure and the gaze-based endpointing measure comprises: decreasing, based on the gaze-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold. 14. The method of claim 11, further comprising: processing a stream of accelerometer data, using an accelerometer model, to generate an accelerometer-based endpointing measure. 15. The method of claim 14, wherein determining the overall endpointing measure is further a function of the accelerometer-based endpointing measure. 16. The method of claim 11, wherein performing the one or more actions based on the spoken utterance comprises: processing the spoken utterance using an automatic speech recognition model to generate a text representation of the spoken utterance. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (B) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 13, 19 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention. Claim 13 recites “pre-fetching the content responsive to …” and wherein “the content” herein has an insufficient antecedent basis for the limitation in claim 13 and causes confusing because it is unclear what it is referred to and it is unclear what it is and thus, renders claim indefinite. Note: claim 12 recites “rendering a content” and it might be that claim 13 depends on claim 12, other than currently recited dependency to claim 1. Claim 19 is rejected for the at least similar reason as described in claim 13 above because claim 19 recites the similar deficient limitation “the content” as recited in claim 13. Claim 19 further recites “the audio-based endpointing measure satisfies the threshold …” and wherein “the audio-based endpointing measure” in line 4 and line 5 of claim 19 has an insufficient antecedent basis for the limitation in claim 19, and further causes confusing because it is unclear what it is referred to and it is unclear what it is and thus, renders claim indefinite. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Konzelmann et al. (US 20200349966 A1, hereinafter Konzelmann) and in view of Piernot et al. (US 20150348548 A1, hereinafter Piernot) Claim 1: Konzelmann teaches a method implemented by one or more processors (title and abstract, ln 1-14, method steps in figs. 4-6 and implemented by a computing device 710 in fig. 7, including processors 716, para 122), the method comprising: processing an audio data stream (processing audio data detected by microphones, para 130), using a machine learning model (adaptation engine 115 having voice activity module 1155), to generate an audio-based endpointing measure (represented by voice activity occurrence detected by voice module 1155, para 41), wherein the audio data stream captures a spoken utterance of a user (user’s speech and/or other audio data captured via microphones 109 in fig. 1, para 37) and is detected via one or more microphones of a client device (the microphones 109 on a client device 101 in fig. 1, para 37); processing a stream of sensor data (sensor data such as image frames from cameras of the client device, para 22), using the machine learning model (the adaptation engine 115 having mouth/voice machine learning models 1163, distance module 1152, mouth/voice module 1153, and other modules 1156 in fig. 1, para 40), to generate an sensor-based endpointing measure (occurrence of mouth movement determined via the models 1163 in fig. 1 and 1153 in fig. 2A, para 48, and a satisfied distance of the user relative to the client device, para 22, e.g., within a threshold distance, para 21) and/or processing a stream of image data (vision frame outputted from visual capture 114 with presence sensor 105 in fig. 2A), using the machine learning model (the adaptation engine 115 having a gaze machine learning model 1161, para 23, 42), to generate a gaze-based endpointing measure that indicates whether the user is looking at the client device (detected occurrence of a directed gaze of the user to the client device, para 22, 42); determining an overall endpointing measure (determining a degree of confidence of various attributes of the user with respect to intending to interact with the assistant, para 94) as a function of (1) the audio-based endpointing measure (the attributes including detected voice activity, para 94) and (2) the sensor-based endpointing measure (the attributes also including detected user’s distance within a distance threshold, para 94 and occurrence of mouth movement, para 91, or co-occurrence of mouth movement, para 94) and/or the gaze-based endpointing measure (the attributes also including detected gaze, para 93-94); determining whether the overall endpointing measure satisfies a threshold (satisfied criteria for rendering 3rd cue 516, e.g., at least a threshold confidence metric met for directed gaze, at least a threshold confidence metrics met for measured distance, voice activity, and mouth movement co-occurrence, para 107-109 and with the completion of user input, para 110-111); and in response to determining the overall endpointing measure satisfies the threshold: performing one or more actions based on the spoken utterance (performing rendering of 3rd cue 516, performing further local processing of vision data and audio data at 518, and determining a provision of response, etc., at 520, para 111-112). However, Konzelmann does not explicitly teach where the sensor data is accelerometer data and an accelerometer-based endpointing measure and thus, Konzelmann does not explicitly teach that the determined overall endpointing measure is also a function of the accelerometer-based endpointing measure. Piernot teaches an analogous field of endeavor by disclosing a method implemented by one or more processors (title and abstract, ln 1-18, method steps in figs. 3-4 and implemented by a system 102 in fig. 2 and wherein the system includes processors 204) and wherein an accelerometer-based endpointing is disclosed (via accelerometer to detect movement of the user device, by which a likelihood that user is intended to speak is determined, e.g., positive contribution to the likelihood if detected movement is about closing to the user’s mouth, para 116, 117, and vise verse, para 115) by processing a stream of accelerometer data (a motion data sensed by the accelerometer of the other sensors 216, para 115), using a machine learning model (using a probabilistic system including machine learning system, para 38, and through processing contextual information, para 56 and the contextual information including the accelerometer data, para 115) and the determination of the overall endpointing measure is also a function of the accelerometer-based endpointing measure (the likelihood or confidence score = sum of the N different types of contextual information including the data contributed by measurement of the accelerometer above, para 38-39) for benefits of improving quality and effectiveness of speech recognition (by simplifying user input, avoiding false detect of user’s speech, para 3-4 and by combining probabilistic data from variety of contributions, para 39) and adapted to individual preference and needs (para 161). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the accelerometer data and the accelerometer-based endpoining measure and wherein the determination of the overall endpointing measure is also the function of the accelerometer-based endpointing measure, as taught by Piernot, to the sensor data and the sensor-based endpointing measure in the method implemented by one or more processors, as taught by Konzelmann, for the benefits discussed above. Claim 14 has been analyzed and rejected according to claim 1 above and the combination of Konzelmann and Piernot further teaches, a client device (Konzelmann, the client device 101 in fig. 1 and Piernot, user device 102 in fig. 1), comprising one or more processors (Konzelmann, processors 716, para 122, and Piernot, processors 204 in the user device 102 in fig. 2), and memory configured to store instructions (Konzelmann, storage subsystem 724 in fig. 7, and storing programming and data constructs as provision of modules, para 236 and Piernot, memory 250 coupled to memory interface 202 in the user device 102 in fig. 1 and storing instructions, para 24) that, when executed by the one or more processors, cause the one or more processors to perform a method that includes (Konzelmann, the modules implemented by the processors 714, para 126, and Piernot, for processing method steps 300 and 400 in figs. 3-4, para 24): processing, using a machine learning model (Konzelmann, adaptation engine 115, including modules gaze 1151 as machine learning models 1161 in fig. 1, para 42, voice activity module 1155 as machine learning models 1165, para 55, and other modules 1156 as other machine learning models, para 56 and Piernot, by using machine learning system, para 12): an audio data stream, where the audio data stream captures a spoken utterance of a user (Konzelmann, discussed in claim 1 above, and Piernot, detecting an endpoint of the spoken by using algorithms, energy calculation, as contextual information, para 29) and is detected via one or more microphones of device (Konzelmann, microphones 109 used for capturing user’s speech, para 37, and Piernot, speech detected by using microphone 230, para 33); and a stream of accelerometer data and/or a stream of image data (Konzelmman, other sensor data, para 56, and image data captured by cameras, para 37, and Piernot, the accelerometer data as discussed in claim 1 above), where the stream of accelerometer data classifies movement of the client device (Konzelmman, by using classification model to detect gesture movement, para 56, and Piernot, using classifier, para 56, and determining the device movement such as shaking, toward or away from the user, para 115), and where the stream of image data classifies the gaze of the user of the client device (Konzelmman, discussed in claim 1 above, and Piernot, using eye-tracking technique to determine whether the user is looking at the user device or not, i.e., gaze at, para 82, 134); determining, based on processing the audio data stream and the stream of accelerometer data and/or the stream of image data, an endpointing measure indicating the likelihood of a candidate endpoint in the audio data stream (Konzelmman, determining a degree of confidence of various attributes of the user with respect to intending to interact with the assistant, para 94, and discussed in claim 1 above, and Piernot, likelihood or confidence score determined to include consideration of audio input and contextual information, para 39); determining whether the endpointing measure satisfies a threshold (Konzelmman, satisfied criteria for rendering 3rd cue 516, at least a threshold confidence metrics met for measured distance, voice activity, and mouth movement co-occurrence, para 107-109, para 110-111, the discussed in claim 1 above, and Piernot, determining that likelihood or confidence score is greater than the threshold value, para 38); and in response to determining the endpointing measure satisfies a threshold, performing one or more actions based on the spoken utterance (Konzelmman, discussed in claim 1 above and Piernot, e.g., in rule-based system, if the threshold is satisfied, a command is expected to be given by the user and executed, para 98). Claim 20 has been analyzed and rejected according to claims 1, 14 above. Claim 2: the combination of Konzelmann and Piernot teaches, the method further comprising: prior to processing the stream of accelerometer data to generate the accelerometer-based endpointing measure and/or prior to processing the stream of image data to generate the gaze-based endpointing measure (Konzelmann, voice activity detected by voice module 1155 in fig. 2A, and occurrence of directed gaze by the gaze module 1151, para 41, and Piernot, voice identified by identifying user’s spoken including user’s question, para 45 and accelerometer data as contextual information to detect movement of the user, para 115), determining whether the audio-based endpointing measure satisfies an initial threshold indicating a candidate endpoint in the audio data stream (Konzelmann, the voice activity occurrence is detected at processing 404, prior to processing gaze data at processing 406 by using vision data in fig. 4, para 102 and Piernot, the spoken user input has been identified at 306, para 36, 45, which is prior to calculating contribution of the contextual data including accelerometer data at 308 in fig. 3, para 36-37, 115, e.g., the distance of the user to the device measured by using the accelerometer, para 131). Claim 3: the combination of Konzelmann and Piernot further teaches, according to claim 2 above, wherein processing the stream of accelerometer data, using the machine learning model, to generate the accelerometer-based endpointing measure further comprises: processing the stream of accelerometer data in response to determining the audio-based endpointing measure satisfies the initial threshold indicating the candidate endpoint in the audio data stream (Konzelmann, other sensor data used for endpointing the distance of the user, or face recognition, etc., at 406, 410, para 46, which is performed after the voice activity is determined at 404, para 102, and the voice activity meeting a second human perceptible cue, para 15and discussed in claim 2 above, and Piernot, the user’s voice input is identified by identifying a user’s question included in the user’s input, para 45, as discussed in claim 2 above). Claim 4: the combination of Konzelmann and Piernot further teaches, according to claim 2 above, wherein processing the stream of image data, using the machine learning model, to generate the gaze-based endpointing measure further comprises: processing the stream of image data in response to determining the audio-based endpointing measure satisfies the initial threshold indicating the candidate endpoint in the audio data stream. Claim 5: the combination of Konzelmann and Piernot further teaches, according to claim 1 above, wherein the accelerometer-based endpointing measure classifies movement of the client device (Konzelmann, detecting a distance threshold of the user device relative to the user’s position, para 21, and Piernot, the motion data from the accelerometer and the motion data representing movement of the user device and caused by the user shaking the device, movement of the device toward or away from the user’s mouth, para 115). Claim 6: the combination of Konzelmann and Piernot further teaches, according to claim 5 above, wherein processing the stream of accelerometer data, using the machine learning model, to generate the accelerometer-based endpointing measure comprises: processing, using the machine learning model model, (i) a portion of the stream of accelerometer data captured prior to the user speaking the spoken utterance (Piernot, contextual information including movement of user device being close to away from the user, and one example, the motion data indicating that the user device was not moved toward the user’s mouth before the spoken user input was received, para 116-117) and (ii) a portion of the stream of accelerometer data captured subsequent to the user beginning to speak the spoken utterance, to generate the accelerometer-based endpointing measure (Piernot, the step 308 is behind the step 306 in fig. 3, and continuation of receiving audio data, para 54, and the discussed in claim 2 above). Claim 7: the combination of Konzelmann and Piernot further teaches, according to claim 2 above, wherein processing the stream of image data, using the machine learning model, to generate the gaze-based endpointing measure further comprises: processing the stream of accelerometer data in response to determining the audio-based endpointing measure satisfies the initial threshold indicating the candidate endpoint in the audio data stream (Konzelmann, the audio activity cooccurred with mouth movement as the initial threshold to endpoint the voice activity occurred, para 15, and Piernot, e.g., voice volume level is greater than a threshold as the indication of the user’s speech input, para 70). Claim 8: the combination of Konzelmann and Piernot further teaches, according to claim 1 above, wherein determining the overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure comprises: boosting, based on the accelerometer-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold and/or boosting, based on the gaze-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold (Konzelmann, one or more confidence measures satisfying one or more thresholds, para 146, and e.g., the likelihood that the automatied assistant properly interprets the spoken utterance is increased, para 18, and Piernot, overall likelihood P = C1 + C2 + … + CN or sum of variety of contributions, i.e., positive contributions from accelerometer, para 115, would boosting the likelihood by increasing the value of P above, para 38-39). Claim 9: the combination of Konzelmann and Piernot further teaches, according to claim 1 above, wherein the accelerometer-based endpointing measure indicates no movement of the client device, and wherein determining the overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure comprises: decreasing, based on the accelerometer-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold (Konzelmann, as discussed in claim 8 above, and an opposite to boosting would be obvious if one or more confidence is not satisfied, and Piernot, if the confidence score is not greater than the threshold value, the likelihood is down to that spoken user input was not intended for the virtual assistant, para 38 and, e.g., detecting the user’s device is close to or away from the user, shaking the device, etc., para 115). Claim 10: the combination of Konzelmann and Piernot further teaches, according to claim 1 above, wherein the gaze-based endpointing measure indicates the user is not looking at the client device, and wherein determining the overall endpointing measure as a function of (1) the audio-based endpointing measure and (2) the accelerometer-based endpointing measure and/or the gaze-based endpointing measure comprises: decreasing, based on the gaze-based endpointing measure, the likelihood the overall endpointing measure satisfies the threshold (Konzelmann, the user’s gaze is redirected, halting is performed, para 13, and e.g., user’s gaze is redirected, the gaze module is no longer detect a gaze, para 94). Claim 11: the combination of Konzelmann and Piernot further teaches, according to claim 1 above, wherein performing the one or more actions based on the spoken utterance comprises: processing the spoken utterance using an automatic speech recognition model to generate a text representation of the spoken utterance (Konzelmann, automatic speech recognition of the audio data is performed for the action, para 132 and Piernot, ASR is applied, para 103). Claim 12: the combination of Konzelmann and Piernot further teaches, according to claim 1 above, wherein performing the one or more actions based on the spoken utterance comprises: rendering content that is responsive to the spoken utterance (Konzelmann, the human perceptible cues can be displayed, para 94, and Piernot, the action can be invoking an output to the user in responses to the user in an audible form, para 13, e.g., the output can be voice, sound, alerts, text messages, menus, graphics, videos, etc., para 27). Claim 13: the combination of Konzelmann and Piernot further teaches, according to claim 1 above, the method further comprising: determining, prior to determining the overall endpointing measure satisfies the threshold, that the audio-based endpointing measure satisfies the threshold or an alternate threshold (Konzelmann, based on the detection of voice activity and directed gaze detection at 502, but prior to the processing on attributes and confidence metrics to determining the final decision of the user speech input at step 514, the first and second human perceptible cue are rendered, e.g., displayed on a screen, para 101, 108, and 136 and Piernot, generating a response to the first spoken user input prior to the determination of the likelihood that the user is intended to have virtual assistant at step 308-310); and pre-fetching the content responsive to determining that the audio-based endpointing measure satisfies the threshold or the alternate threshold (Konzelmann, outputting the first and the second human perceptual cues, e.g., by color, shape, etc., para 136 and Piernot, the response to the first spoken user input is performed, as pre-fetching action in fig. 4). Claim 15 has been analyzed and rejected according to claims 14, 1 above. Claim 16 has been analyzed and rejected according to claims 15, 8 above. Claim 17 has been analyzed and rejected according to claims 14, 11 above. Claim 18 has been analyzed and rejected according to claims 14, 12 above. Claim 19 has been analyzed and rejected according to claims 14, 13 above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to LESHUI ZHANG whose telephone number is (571)270-5589. The examiner can normally be reached Monday-Friday 6:30amp-4:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LESHUI ZHANG/ Primary Examiner, Art Unit 2695
Read full office action

Prosecution Timeline

Nov 25, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103, §112, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707227
METHOD FOR PROCESSING AUDIO INPUT DATA AND A DEVICE THEREOF
2y 6m to grant Granted Aug 11, 2026
Patent 12706102
SEPARATING SPATIAL AUDIO OBJECTS
2y 10m to grant Granted Aug 11, 2026
Patent 12694886
AREA REPRODUCTION SYSTEM AND AREA REPRODUCTION METHOD
2y 6m to grant Granted Jul 28, 2026
Patent 12682897
MOTOR VEHICLE AND METHOD FOR SUMMARIZING A CONVERSATION IN A MOTOR VEHICLE
2y 9m to grant Granted Jul 14, 2026
Patent 12676162
APPARATUS AND METHOD FOR MULTICHANNEL INTERFERENCE CANCELLATION
6y 8m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+35.2%)
2y 9m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 950 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month