Prosecution Insights
Last updated: August 18, 2026
Application No. 18/565,041

IMAGE ANALYSIS TO SWITCH AUDIO DEVICES

Non-Final OA §103
Filed
Nov 28, 2023
Priority
Jun 18, 2021 — nonprovisional of PCTUS2021038090
Examiner
DUFFY, CAROLINE TABANCAY
Art Unit
2662
Tech Center
2600 — Communications
Assignee
Hewlett-Packard Development Company, L.P.
OA Round
2 (Non-Final)
80%
Grant Probability
Favorable
2-3
OA Rounds
2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
74 granted / 92 resolved
+18.4% vs TC avg
Strong +18% interview lift
Without
With
+18.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
13 currently pending
Career history
100
Total Applications
across all art units

Statute-Specific Performance

§101
13.9%
-26.1% vs TC avg
§103
58.6%
+18.6% vs TC avg
§102
8.1%
-31.9% vs TC avg
§112
16.6%
-23.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 92 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed 05/04/2026 has been entered. Claims 1-15 remain pending. Claims 16-20 are new. In response to the amendment to Claim 6, the objection of record is withdrawn. In response to the amendments to Claims 6 and 10, the rejections under 35 U.S.C. 112(b) are withdrawn. In response to the amendments to Claims 1 and 11, the rejections under 35 U.S.C. 101 are withdrawn. In response to the amendments to Claim 1, the provisional double patenting rejection is withdrawn because the reference application 18/560793 does not recite donning or doffing, and these features are patentably distinct. Furthermore, reference application 18/560793 is no longer pending. Response to Arguments Applicant's arguments filed 05/04/2026 regarding Claims 1-5 have been fully considered but they are not persuasive. Regarding Claim 1, Applicant argues that Smith does not teach “automatically select the audio endpoint of the computing device based on the sequence of motion” and that “muting or unmuting the same audio channel within an HMD is fundamentally different from selecting or switching among audio endpoints.” However, under the broadest reasonable interpretation, “select” does not require any audio outputting action, but merely that a particular audio output is identified, or selected, and may or may not be then used for output or some other action. Additionally, the argument relies on selection or switching “among audio endpoints,” however Claim 1 recites only “an audio endpoint,” indicating a single audio endpoint. Thus the selection is a selection of a single audio endpoint from only one available audio endpoint, not “among audio endpoints.” The gestures of Smith of covering or uncovering a left or right ear are followed by muting or adjusting volume of an audio of the left or right audio channel: Smith, [0109] discloses “In response to detecting gesture sequence 900A, a previously muted left audio channel (e.g., reproduced by a left speaker in the HMD) is unmuted; without changes to the operation of the right audio channel. Additionally, or alternatively, a volume of the audio being reproduced in the left audio channel may be increased, also without affecting the right audio channel.” Thus Smith teaches detecting a gesture sequence and automatically selecting a left speaker, in this case, for the subsequent action of unmuting or volume increase. In the absence of a limitation requiring turning on the automatically selected audio endpoint, muting or unmuting the selected audio endpoint, or otherwise using the audio endpoint in some manner, the limitation “automatically select the audio endpoint of the computing device based on the sequence of motion” is taught by the selection of a left or right audio channel for subsequent actions of unmuting based on a gesture sequence, as taught by Smith. Applicant’s arguments filed 05/04/2026, with respect to Claims 6-15 have been fully considered and are persuasive. The rejections of record under 35 U.S.C. 103 have been withdrawn. In particular, per the response above, Claims 6 and 11 require at least two audio endpoints; Claim 6 recites “automatically switch from a second audio endpoint to the first audio endpoint” and Claim 11 recites “automatically switch between the speaker and the wearable audio endpoint.” Thus, the arguments regarding selecting or switching “among audio endpoints” are considered persuasive in Claims 6 and 11 because these claims require multiple audio endpoints. Furthermore, Claims 6 and 11 recite “automatically switch” audio endpoints where Claim 1 recites “automatically select” an audio endpoint. Even under the broadest reasonable interpretation, a switching step between more than one audio endpoints requires that one audio endpoint is engaged, selected, or identified for some further action, while an other audio endpoint is disengaged, deselected, or no longer identified for some further action. Thus, although Smith teaches muting or unmuting a selected audio channel, Smith explicitly teaches, for example in [0109], “a previously muted left audio channel (e.g., reproduced by a left speaker in the HMD) is unmuted; without changes to the operation of the right audio channel.” Thus, because Smith teaches no changes to the operation of the other audio channel, even under the broadest reasonable interpretation, Smith does not explicitly teach “automatically switch” audio endpoints because the other audio endpoint is not disengaged, deselected, or otherwise no longer identified or used for further action. Thus, the arguments with respect to Claims 6-15 are considered persuasive and the rejections under 35 U.S.C. 103 are withdrawn. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claims 1-5 are rejected under 35 U.S.C. 103 as being unpatentable over Xiang et al. (US 2020/0077193 A1) in view of Smith et al. (US 2019/0384406 A1). Regarding Claim 1, Xiang teaches “A non-transitory machine-readable medium comprising instructions that, when executed by a processor, cause the processor to” (Xiang, [0170] discloses “A processor or other means for processing as disclosed herein may also be embodied as one or more computers (e.g., machines including one or more arrays programmed to execute one or more sets or sequences of instructions) or other processors”): “analyze a video captured by a computing device to detect a sequence of motion in the video (Xiang, [0147] discloses “FIG. 28C shows a block diagram of an implementation A120 of apparatus A110 that includes a capture device CD10 which captures the scene that includes the gesture”; where an apparatus A110 is a computing device; where a scene that includes a gesture is a video captured by a computing device; where a gesture is a sequence of motion in the video); “and Xiang does not explicitly teach “indicating a donning gesture or a doffing gesture associated with an audio endpoint” and “automatically select an audio endpoint of the computing device based on the sequence of motion.” However, in an analogous field of endeavor, Smith teaches “indicating a donning gesture or a doffing gesture associated with an audio endpoint” (Smith, Fig. 9A sequence 900A shows motions corresponding to unmuting: Smith, [0109] discloses “In gesture sequence 900A, starting position 901 shows a user covering her left ear 910 with the palm of her left hand 909, followed by motion 902, where the left hand 909 moves out and away from the user, uncovering her left ear 910”; where gesture sequence 900A is a donning gesture. Smith, Fig. 9B sequence 900B shows motions corresponding to muting: Smith, [0110] discloses “In gesture sequence 900B, starting position 903 shows a user with her left hand 909 out and away by a selected distance, followed by motion 904 where the user covers her left ear 910 with the palm of her left hand 909”; where gesture sequence 900B is a doffing gesture); and “automatically select an audio endpoint of the computing device based on the sequence of motion” (Smith [0109] discloses “In response to detecting gesture sequence 900A, a previously muted left audio channel (e.g., reproduced by a left speaker in the HMD) is unmuted; without changes to the operation of the right audio channel”; where performing a muting action on a left audio channel is automatically selecting an audio endpoint for the subsequent process of muting). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xiang to incorporate the teachings of Smith by detecting donning and doffing sequences and muting a particular audio channel in response to a detected gesture. The prior art contained a ‘base’ device upon which the claimed invention can be seen as an ‘improvement.’ That is, Xiang teaches a device that detects gestures from a captured scene. The prior art contained a known technique that is applicable to the base device. Smith teaches a technique of adjusting audio channels based on detected gestures, and specific gestures covering or uncovering the ears (donning or doffing). One of ordinary skill in the art would have recognized that applying the known technique would have yielded predictable results and resulted in an improved system. Additionally, Xiang broadly teaches a “muting gesture” that may be interpreted by the “gesture interpreter” (see Xiang, [0143]). Smith teaches a particular audio control of audio channels. Thus, one of ordinary skill in the art would be motivated to combine the Xiang and Smith references by applying a known technique to a known device ready for improvement to yield predictable results. Accordingly, the combination of Xiang and Smith discloses the invention of Claim 1. Regarding Claim 2, the combination of Xiang and Smith teaches “The non-transitory machine-readable medium of claim 1, wherein: the audio endpoint is a first audio endpoint” (Smith [0110] discloses “In response to detecting gesture sequence 900B, the left audio channel may be muted, or its volume may be reduced, without changes to the operation of the right audio channel”; where a right audio channel is a first audio endpoint); “and based on the sequence of motion, the instructions are further to deselect a second audio endpoint” (Smith [0110] discloses “In response to detecting gesture sequence 900B, the left audio channel may be muted, or its volume may be reduced, without changes to the operation of the right audio channel”; where a left audio channel is a second audio endpoint; where muting a left audio channel is deselecting a second audio endpoint). The proposed combination as well as the motivation for combining the Xiang and Smith references presented in the rejection of Claim 1, apply to Claim 2 and are incorporated herein by reference. Thus, the apparatus recited in Claim 2 is met by Xiang and Smith. PNG media_image1.png 530 711 media_image1.png Greyscale Fig. 9A and 9B of Smith Regarding Claim 3, the combination of Xiang and Smith teaches “The non-transitory machine-readable medium of claim 1, wherein the instructions are further to: analyze the video to detect a representation of the audio endpoint in the video” (Smith, 110 discloses “In gesture sequence 900B, starting position 903 shows a user with her left hand 909 out and away by a selected distance, followed by motion 904 where the user covers her left ear 910 with the palm of her left hand 909. In response to detecting gesture sequence 900B, the left audio channel may be muted, or its volume may be reduced, without changes to the operation of the right audio channel”; where a gesture covering a left ear is a representation of the audio end point in the video); “and select the audio endpoint further based on detection of the representation of the audio endpoint” (Smith, 110 discloses “In response to detecting gesture sequence 900B, the left audio channel may be muted, or its volume may be reduced, without changes to the operation of the right audio channel”; where muting a left audio channel is selecting an audio endpoint). The proposed combination as well as the motivation for combining the Xiang and Smith references presented in the rejection of Claim 1, apply to Claim 3 and are incorporated herein by reference. Thus, the apparatus recited in Claim 3 is met by Xiang and Smith. Regarding Claim 4, the combination of Xiang and Smith teaches “The non-transitory machine-readable medium of claim 1, wherein the instructions are further to: classify the sequence of motion as a donning gesture” (Smith, Fig. 9A sequence 900A shows motions corresponding to unmuting: Smith, [0109] discloses “In gesture sequence 900A, starting position 901 shows a user covering her left ear 910 with the palm of her left hand 909, followed by motion 902, where the left hand 909 moves out and away from the user, uncovering her left ear 910”; where gesture sequence 900A is a donning gesture); “and select the audio endpoint based on the donning gesture.” (Smith, Fig. 9A discloses “In response to detecting gesture sequence 900A, a previously muted left audio channel (e.g., reproduced by a left speaker in the HMD) is unmuted”). The proposed combination as well as the motivation for combining the Xiang and Smith references presented in the rejection of Claim 1, apply to Claim 4 and are incorporated herein by reference. Thus, the apparatus recited in Claim 4 is met by Xiang and Smith. Regarding Claim 5, the combination of Xiang and Smith discloses “The non-transitory machine-readable medium of claim 1, wherein the instructions are further to: classify the sequence of motion as a doffing gesture” (Smith, Fig. 9B sequence 900B shows motions corresponding to muting: Smith, [0110] discloses “In gesture sequence 900B, starting position 903 shows a user with her left hand 909 out and away by a selected distance, followed by motion 904 where the user covers her left ear 910 with the palm of her left hand 909”; where gesture sequence 900B is a doffing gesture); “and select the audio endpoint based on the doffing gesture” (Smith, [0110] discloses “In response to detecting gesture sequence 900B, the left audio channel may be muted”). The proposed combination as well as the motivation for combining the Xiang and Smith references presented in the rejection of Claim 1, apply to Claim 5 and are incorporated herein by reference. Thus, the apparatus recited in Claim 5 is met by Xiang and Smith. Allowable Subject Matter Claims 6-20 are allowed. The following is an examiner’s statement of reasons for allowance: As detailed in the response to arguments above, Claims 6 and 11 require at least two audio endpoints; Claim 6 recites “automatically switch from a second audio endpoint to the first audio endpoint” and Claim 11 recites “automatically switch between the speaker and the wearable audio endpoint.” Thus, the arguments regarding automatically selecting or switching “among audio endpoints” are considered persuasive in Claims 6 and 11 because these claims require multiple audio endpoints. Although Xiang and Smith teach multiple audio endpoints (Xiang, [0054] discloses a “loudspeaker array” and Smith teaches left and right audio channels of a head mounted device), neither Xiang nor Smith explicitly teaches “automatically switch” between the audio endpoints. That is, the teaching of Smith of muting or unmuting a particular left or right audio channel while not altering the function of the other audio channel does not explicitly teach an automatic switching step. Thus, none of the previously cited prior art, alone or in combination, provides a motivation to teach the ordered combination of “A non-transitory machine-readable medium comprising instructions that, when executed by a processor, cause the processor to: apply images captured by a computing device to a trained machine-learning system; detect with the trained machine-learning system a visual indication indicative of an engagement of a first audio endpoint; in response to detection of the visual indication, detect with the trained machine-learning system a representation of the first audio endpoint; and in response to detection of the representation of the first audio endpoint, automatically switch from a second audio endpoint to the first audio endpoint of the computing device.” Thus, none of the previously cited prior art, alone or in combination, provides a motivation to teach the ordered combination of “A computing device comprising: a camera; a speaker; a machine-learning system; and a processor connected to the camera and the speaker, the processor further connectable to a wearable audio endpoint, the processor to: apply the machine-learning system to perform image analysis on images captured by the camera; and automatically switch between the speaker and the wearable audio endpoint based on the image analysis.” Dependent Claims 7-10 and 12-20 contain all allowable subject matter of Claims 6 and 11 and thus are also allowable. Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.” Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAROLINE TABANCAY DUFFY whose telephone number is (703)756-1859. The examiner can normally be reached Monday - Friday 8:00 am - 5:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at 5712723382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CAROLINE TABANCAY DUFFY/Examiner, Art Unit 2662 /AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Show 2 earlier events
Mar 25, 2026
Interview Requested
Apr 02, 2026
Applicant Interview (Telephonic)
Apr 02, 2026
Examiner Interview Summary
May 05, 2026
Response Filed
May 27, 2026
Final Rejection mailed — §103
Jul 24, 2026
Response after Non-Final Action
Aug 13, 2026
Applicant Interview (Telephonic)
Aug 13, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705743
MEDICAL IMAGE PROCESSING DEVICE AND ENDOSCOPE SYSTEM
2y 8m to grant Granted Aug 11, 2026
Patent 12700219
TRAINING OF MACHINE LEARNING MODELS FOR VIDEO ANALYTICS
3y 4m to grant Granted Aug 04, 2026
Patent 12700225
Channel Fusion for Vision-Language Representation Learning
2y 10m to grant Granted Aug 04, 2026
Patent 12694666
METHOD FOR STABILIZING LINE-OF-SIGHT FALLING POINT
3y 7m to grant Granted Jul 28, 2026
Patent 12682750
TRAFFIC CONTROL APPARATUS, SYSTEM, METHOD, AND COMPUTER READABLE MEDIUM
2y 8m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+18.5%)
2y 11m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 92 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month