Prosecution Insights
Last updated: October 02, 2026
Application No. 18/249,913

VOCAL TRACK REMOVAL BY CONVOLUTIONAL NEURAL NETWORK EMBEDDED VOICE FINGER PRINTING ON STANDARD ARM EMBEDDED PLATFORM

Final Rejection §102§103
Filed
Apr 20, 2023
Priority
Oct 22, 2020 — nonprovisional of PCTCN2020122852
Examiner
SCOLES, PHILIP GRANT
Art Unit
2837
Tech Center
2800 — Semiconductors & Electrical Systems
Assignee
Harman International Industries Incorporated
OA Round
2 (Final)
57%
Grant Probability
Moderate
3-4
OA Rounds
1m
Est. Remaining
72%
With Interview

Examiner Intelligence

Grants 57% of resolved cases
57%
Career Allowance Rate
40 granted / 70 resolved
-10.9% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
33 currently pending
Career history
100
Total Applications
across all art units

Statute-Specific Performance

§101
1.5%
-38.5% vs TC avg
§103
59.3%
+19.3% vs TC avg
§102
19.6%
-20.4% vs TC avg
§112
16.5%
-23.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 70 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments, see page 5, line 23 – page 6, line 13, filed 6/22/2026, have been fully considered and are persuasive. The claim objections and 35 U.S.C. 112 rejections of record are withdrawn in response to Applicant’s amendments. Applicant's arguments, see page 6, line 14 – page 8, line 12, filed 6/22/2026, have been fully considered but they are not persuasive. Applicant has amended claim 1 to recite processing the input music “to obtain, based on voice fingerprints associated with an input music spectrogram magnitude image, voice spectrogram mask and accompaniment spectrogram mask, separately” and contends that “Jansson is completely silent in this regard.” Applicant’s arguments are incommensurate with the scope of the claims and with terms that are defined in the instant specification, as discussed below. Claim 1 scope The BRI of a claim term is the interpretation that a POSITA would reach when reading the term in light of the specification. (MPEP § 2111.01(I).) However, an Applicant may be his or her own lexicographer. (MPEP § 2111.1(IV).) Here, Applicant’s own specification supplies the definition: “The voice fingerprints can be considered as the summarized features of the voice separation model. In the example, the voice fingerprints reflect the weight of each layer in the 2D convolutional neural network.” (Instant Specification ¶0029.) The instant specification further identifies the “model features” that training modifies as “the weight and bias of the convolution kernel, and the batch normalization matrix parameters.” (Instant Specification ¶0031.) Conversely, by these arguments, Applicant is attempting to disavow claim scope that is not only encompassed within the plain and ordinary meaning of the amended language, but which is also expressly defined with Applicant’s own written description. Although plain and ordinary meaning preempts a narrower meaning argued by an applicant, express definition, when supplied in a specification, controls. (MPEP § 2111.01(V).) Here, because Applicant has expressly defined the term in the instant specification, this Examiner is bound to that definition rather than to its plain and ordinary meaning or to the narrower scope subsequently argued by Applicant. Claim 1 anticipation Applicant’s argument that Jansson would have to expressly disclose obtaining masks “based on voice fingerprints” fails to consider the full breadth of an anticipation inquiry. A claim limitation that is necessarily present in a reference’s disclosed method may be met inherently, whether or not the reference recognizes or describes the feature. The absence of express recitation does not defeat the rejection if the feature is inherently present. (MPEP § 2112.02(I).) Here, Jansson discloses a method of applying the magnitude patch of the input signal to the trained U-Net, whose final decoder layer output “is a soft mask that is multiplied element-wise with the mixed spectrogram to obtain the final estimate” (Jansson § 3.1). This necessarily computes the network’s summarized feature representations from the input spectrogram magnitude image in generating masks. Under the construction compelled by Applicant’s own written description (Instant Specification ¶0029), Jansson’s masks are necessarily “based on voice fingerprints associated with an input music spectrogram magnitude image.” Applicant may not import a term from the instant specification, define it as the network’s summarized features, and then argue that the reference must recite different, narrower subject matter to meet it. Additional references of record MPEP § 2145 states that “arguments presented by applicant cannot take the place of factually supported objective evidence.” Here, Applicant’s assertion that “[a] careful review of the other references cited by the Examiner shows that those references also fail to teach or suggest the above limitations of amended claim 1” is a conclusory argument unaccompanied by claim mapping, record evidence, declaration, or affidavit. Therefore, it is unconvincing, and particularly so to the extent that any of the other references of record may inherently disclose or teach the limitation in question in a manner similar to Jansson as discussed above. Claims 2-20 The above rationale applies equally to independent claim 11, as well as to dependent claims 2-10 and 12-20 by incorporation. Furthermore, the amendment does not disturb any other mapping presented in the rejection of record. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless –(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 3-5, 7-8, 11, 13-15, and 17-18 are rejected under 35 U.S.C. 102 as anticipated by Jansson et. al ("Singing Voice Separation with Deep U-Net Convolutional Networks," 2017, retrieved on 3/29/2026 from https://archives.ismir.net/ismir2017/paper/000171.pdf), hereinafter Jansson. Regarding claim 1, Janssen discloses a vocal removal method (Jansson § 1: "The task of automatic singing voice separation consists of estimating what the sung melody and accompaniment would sound like in isolation."), comprising the steps of: training, by a machine learning module (Jansson § 3.1.2: "The model is trained using the ADAM [12] optimizer."), a voice separation model (Jansson § 3: "This work adapts the U-Net [24] architecture to the task of vocal separation"); extracting, by a feature extraction module, music signal processing features of input music (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."); processing the input music, by the voice separation model, to obtain, based on voice fingerprints associated with an input music spectrogram magnitude image (Jansson § 3.1: "the output of the final decoder layer is a soft mask that is multiplied element-wise with the mixed spectrogram to obtain the final estimate." Jansson § 3.1.1: "Two U-Nets, Θv and Θi, are trained to predict vocal and instrumental spectrogram masks, respectively." Jansson § 3.1.2: "The magnitude spectrograms are normalized to the range [0,1]." The instant specification in ¶0029 defines "voice fingerprints" as "the summarized features of the voice separation model" that "reflect the weight of each layer in the 2E convolutional neural network." A trained convolutional neural network such as Jansson's U-Net necessarily computes its summarized feature representations from the input music spectrogram magnitude image in generating the soft masks, and the per-pixel voice and accompaniment probabilities from which the masks are formed are necessarily obtained based thereon. The instant specification in ¶0033 states that fingerprints are obtained after processing by the 2D convolutional neural network, and probabilities for each frequency bin of each pixel are obtained therefrom. Because Jansson's U-Net cannot produce the claimed masks other than by computing such summarized feature representations from the input spectrogram magnitude image, this limitation is inherently met. (MPEP § 2112.02(I).)), voice spectrogram mask and accompaniment spectrogram mask (Jansson § 3.1: "The goal of the neural network architecture is to predict the vocal and instrumental components of its input indirectly: the output of the final decoder layer is a soft mask that is multiplied element-wise with the mixed spectrogram to obtain the final estimate. Figure 1 outlines the network architecture. In this work, we choose to train two separate models for the extraction of the instrumental and vocal components of a signal"), separately (Jansson § 3.1.1: "Two U-Nets, Θv and Θi, are trained to predict vocal and instrumental spectrogram masks, respectively"); and reconstructing, by a feature reconstruction module, voice minimized music (Jansson § 3.1.3: " The audio signal for an individual (vocal/instrumental) component is rendered by constructing a spectrogram: the output magnitude is given by applying the mask predicted by the U-Net to the magnitude of the original spectrum, while the output phase is that of the original spectrum, unaltered."). Regarding claim 3, Jansson discloses a vocal removal method comprising the features of claim 1 as discussed above. Jansson further discloses that the voice separation model comprises a convolutional neural network (Jansson § 3: "This work adapts the U-Net [24] architecture to the task of vocal separation. The architecture was introduced in biomedical imaging, to improve precision and localization of microscopic images of neuronal structures. The architecture builds upon the fully convolutional network [14]"). Regarding claim 4, Jansson discloses a vocal removal method comprising the features of claim 3 as discussed above. Jansson further discloses that training the voice separation model comprises modifying features of the voice separation model via machine learning (Jansson § 3.1.1: "The loss function used to train the model is the L1,1 norm1 of the difference of the target spectrogram and the masked input spectrogram: L(X,Y;Θ) = ||f(X,Θ) X −Y||1,1 (1) where f(X,Θ) is the output of the network model applied to the input X with parameters Θ– that is the mask generated by the model." The features of the model are modified via a loss function.). Regarding claim 5, Jansson discloses a vocal removal method comprising the features of claim 1 as discussed above. Jansson further discloses extracting the music signal processing features of the input music comprises composing spectrogram images of the input music (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."). Regarding claim 7, Jansson discloses a vocal removal method comprising the features of claim 5 as discussed above. Jansson further discloses that the spectrogram images of the input music are composed using the music signal processing features (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."). Regarding claim 8, Jansson discloses a vocal removal method comprising the features of claim 1 as discussed above. Jansson further discloses that processing the input music comprises inputting spectrogram magnitude of the input music into the voice separation model (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."). Regarding claim 11, Jansson discloses a vocal removal system (Jansson § 1: "The task of automatic singing voice separation consists of estimating what the sung melody and accompaniment would sound like in isolation."), comprising: a machine learning module for training (Jansson § 3.1.2: "The model is trained using the ADAM [12] optimizer.") a voice separation model (Jansson § 3: "This work adapts the U-Net [24] architecture to the task of vocal separation"); a feature extraction module for extracting music signal processing features of input music (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."), wherein the voice separation model processes the input music to obtain, based on voice fingerprints associated with an input music spectrogram magnitude image (Jansson § 3.1: "the output of the final decoder layer is a soft mask that is multiplied element-wise with the mixed spectrogram to obtain the final estimate." Jansson § 3.1.1: "Two U-Nets, Θv and Θi, are trained to predict vocal and instrumental spectrogram masks, respectively." Jansson § 3.1.2: "The magnitude spectrograms are normalized to the range [0,1]." The instant specification in ¶0029 defines "voice fingerprints" as "the summarized features of the voice separation model" that "reflect the weight of each layer in the 2E convolutional neural network." A trained convolutional neural network such as Jansson's U-Net necessarily computes its summarized feature representations from the input music spectrogram magnitude image in generating the soft masks, and the per-pixel voice and accompaniment probabilities from which the masks are formed are necessarily obtained based thereon. The instant specification in ¶0033 states that fingerprints are obtained after processing by the 2D convolutional neural network, and probabilities for each frequency bin of each pixel are obtained therefrom. Because Jansson's U-Net cannot produce the claimed masks other than by computing such summarized feature representations from the input spectrogram magnitude image, this limitation is inherently met. (MPEP § 2112.02(I).)), voice spectrogram mask and accompaniment spectrogram mask (Jansson § 3.1: "The goal of the neural network architecture is to predict the vocal and instrumental components of its input indirectly: the output of the final decoder layer is a soft mask that is multiplied element-wise with the mixed spectrogram to obtain the final estimate. Figure 1 outlines the network architecture. In this work, we choose to train two separate models for the extraction of the instrumental and vocal components of a signal"), separately (Jansson § 3.1.1: "Two U-Nets, Θv and Θi, are trained to predict vocal and instrumental spectrogram masks, respectively"); and a feature reconstruction module for reconstructing voice minimized music (Jansson § 3.1.3: "The audio signal for an individual (vocal/instrumental) component is rendered by constructing a spectrogram: the output magnitude is given by applying the mask predicted by the U-Net to the magnitude of the original spectrum, while the output phase is that of the original spectrum, unaltered."). Regarding claim 13, Jansson discloses a vocal removal system comprising the features of claim 11 as discussed above. Jansson further discloses that the voice separation model comprises a convolutional neural network (Jansson § 3: "This work adapts the U-Net [24] architecture to the task of vocal separation. The architecture was introduced in biomedical imaging, to improve precision and localization of microscopic images of neuronal structures. The architecture builds upon the fully convolutional network [14]"). Regarding claim 14, Jansson discloses a vocal removal system comprising the features of claim 13 as discussed above. Jansson further discloses that training the voice separation model comprises modifying features of the voice separation model via machine learning (Jansson § 3.1.1: "The loss function used to train the model is the L1,1 norm1 of the difference of the target spectrogram and the masked input spectrogram: L(X,Y;Θ) = ||f(X,Θ) X −Y||1,1 (1) where f(X,Θ) is the output of the network model applied to the input X with parameters Θ– that is the mask generated by the model." The features of the model are modified via a loss function.). Regarding claim 15, Jansson discloses a vocal removal system comprising the features of claim 11 as discussed above. Jansson further discloses extracting the music signal processing features of the input music comprises composing spectrogram images of the input music (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."). Regarding claim 17, Jansson discloses a vocal removal system comprising the features of claim 15 as discussed above. Jansson further discloses that the spectrogram images of the input music are composed using the music signal processing features (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."). Regarding claim 18, Jansson discloses a vocal removal system comprising the features of claim 11 as discussed above. Jansson further discloses that processing the input music comprises inputting spectrogram magnitude of the input music into the voice separation model (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network. The magnitude spectrograms are normalized to the range [0,1]."). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 2 and 12 are rejected under 35 U.S.C. 103 as unpatentable over Jansson in view Choi (US 20200035256 A1, January 30, 2020), hereinafter Choi. Regarding claim 2, Jansson discloses a vocal removal method comprising the features of claim 1 as discussed above. Jansson does not explicitly disclose that the voice separation model is generated and put on an embedded platform. However, Choi suggests that the voice separation model is generated and put on an embedded platform (Choi ¶0020-0022: "Another aspect of the present disclosure relates to an electronic instrument having a keyboard wherein respective keys are luminescent, comprising: a memory that stores a trained model generated with machine learning; and at least one processor configured to transform a first audio type of audio data into a first image type of image data, wherein a first audio component and a second audio component are mixed in the first audio type of audio data."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the vocal removal method of Jansson by adding embedded platform of Choi to separate audio component data from audio data for instructional purposes (Choi ¶0002-0006). Regarding claim 12, Jansson discloses a vocal removal system comprising the features of claim 11 as discussed above. Jansson does not explicitly disclose that the voice separation model is generated and put on an embedded platform. However, Choi suggests that the voice separation model is generated and put on an embedded platform (Choi ¶0020-0022: "Another aspect of the present disclosure relates to an electronic instrument having a keyboard wherein respective keys are luminescent, comprising: a memory that stores a trained model generated with machine learning; and at least one processor configured to transform a first audio type of audio data into a first image type of image data, wherein a first audio component and a second audio component are mixed in the first audio type of audio data."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the vocal removal method of Jansson by adding embedded platform of Choi to separate audio component data from audio data for instructional purposes (Choi ¶0002-0006). Claims 6, 9, 16, and 19 are rejected under 35 U.S.C. 103 as unpatentable over Jansson in view of Covell et al. (US 20210166731 A1, filed September 1, 2020), hereinafter Covell. Regarding claim 6, Jansson discloses a vocal removal method comprising the features of claim 1 as discussed above. Jansson further teaches that the music signal processing features comprise window shape (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames"), frequency resolution (Jansson § 3.1.2: "Given the heavy computational requirements of training such a model, we first downsample the input audio to 8192 Hz in order to speed up processing."), and time buffer (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network."). Jansson does not explicitly disclose that the music signal processing features comprise overlap percentage. However, Covell teaches that the music signal processing features comprise overlap percentage (Covell ¶0056: "As another example, in some embodiments, process 200 can calculate the spectrograms with any suitable percentage overlap between slices (e.g., 50% overlap, 75% overlap, 80% overlap, and/or any other suitable percentage overlap)."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the music processing features of Jansson by adding the overlap percentages of Covell to improve transitions between songs (Covell ¶0068). Regarding claim 9, Jansson discloses a vocal removal method comprising the features of claim 1 as discussed above. Jansson does not explicitly disclose modifying the music signal processing features. However, Covell teaches modifying the music signal processing features (Covell ¶0056: "As another example, in some embodiments, process 200 can calculate the spectrograms with any suitable percentage overlap between slices (e.g., 50% overlap, 75% overlap, 80% overlap, and/or any other suitable percentage overlap)." Modifying percentage overlap comprises modifying the music signal processing features.). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the music processing features of Jansson by adding the overlap percentages of Covell to improve transitions between songs (Covell ¶0068). Regarding claim 16, Jansson discloses a vocal removal system comprising the features of claim 11 as discussed above. Jansson further teaches that the music signal processing features comprise window shape (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames"), frequency resolution (Jansson § 3.1.2: "Given the heavy computational requirements of training such a model, we first downsample the input audio to 8192 Hz in order to speed up processing."), and time buffer (Jansson § 3.1.2: "We then compute the Short Time Fourier Transform with a window size of 1024 and hop length of 768 frames, and extract patches of 128 frames (roughly 11 seconds) that we feed as input and targets to the network."). Jansson does not explicitly disclose that the music signal processing features comprise overlap percentage. However, Covelle teaches that the music signal processing features comprise overlap percentage (Covelle ¶0056: "As another example, in some embodiments, process 200 can calculate the spectrograms with any suitable percentage overlap between slices (e.g., 50% overlap, 75% overlap, 80% overlap, and/or any other suitable percentage overlap)."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the music processing features of Jansson by adding the overlap percentages of Covell to improve transitions between songs (Covell ¶0068). Regarding claim 19, Jansson discloses a vocal removal system comprising the features of claim 11 as discussed above. Jansson does not explicitly disclose that the machine learning module modifies the music signal processing features. However, Covell teaches that the machine learning module modifies the music signal processing features (Covell ¶0056: "As another example, in some embodiments, process 200 can calculate the spectrograms with any suitable percentage overlap between slices (e.g., 50% overlap, 75% overlap, 80% overlap, and/or any other suitable percentage overlap)." Modifying percentage overlap comprises modifying the music signal processing features.). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the music processing features of Jansson by adding the overlap percentages of Covelle to improve transitions between songs (Covell ¶0068). Claims 10 and 20 are rejected under 35 U.S.C. 103 as unpatentable over Jansson in view Liu (CN 111128211 A, May 8, 2020), hereinafter Liu. Regarding claim 10, Jansson discloses a vocal removal method comprising the features of claim 1 as discussed above. Jansson does not explicitly disclose reinforce learning the voice separation model. However, Liu teaches training, via reinforcement learning, the voice separation model (Liu ¶0109: "This method proposes a speech separation method based on reinforcement learning, similar to the actor-critic network structure."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the music processing features of Jansson by adding the reinforcement learning of Liu to avoid the problem of inconsistency between training loss and test index and improve the performance of voice separation (Liu ¶0109). Regarding claim 20, Jansson discloses a vocal removal system comprising the features of claim 11 as discussed above. Jansson does not explicitly disclose that the voice separation model is further trained via reinforcement learning. However, Liu teaches that the voice separation model is further trained via reinforcement learning (Liu ¶0109: "This method proposes a speech separation method based on reinforcement learning, similar to the actor-critic network structure."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the music processing features of Jansson by adding the reinforcement learning of Liu to avoid the problem of inconsistency between training loss and test index and improve the performance of voice separation (Liu ¶0109). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHILIP SCOLES whose telephone number is (703)756-1831. The examiner can normally be reached Monday-Friday 8:30-4:30 ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei Hammond can be reached on 571-270-7938. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHILIP G SCOLES/ Examiner, Art Unit 2837 /DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837
Read full office action

Prosecution Timeline

Apr 20, 2023
Application Filed
Apr 07, 2026
Non-Final Rejection mailed — §102, §103
Jun 22, 2026
Response Filed
Sep 25, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731566
Electronic Cymbal
4y 3m to grant Granted Sep 08, 2026
Patent 12711933
KEY FOR KEYBOARD DEVICE
3y 11m to grant Granted Aug 18, 2026
Patent 12670888
ELECTRONIC MUSICAL INSTRUMENT, ELECTRONIC MUSICAL INSTRUMENT CONTROLLING METHOD AND NON-TRANSITORY COMPUTER-READABLE STORAGE MEDIUM
4y 2m to grant Granted Jun 30, 2026
Patent 12646494
Attaching Hand-Actuated Music Controllers to a Saxophone
4y 0m to grant Granted Jun 02, 2026
Patent 12620381
SWITCH LOCK APPARATUS AND METHOD THEREOF
1y 9m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
57%
Grant Probability
72%
With Interview (+14.5%)
3y 7m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 70 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month