Prosecution Insights
Last updated: October 02, 2026
Application No. 18/994,678

SPEECH DATA PROCESSING METHOD, SPEECH DATA PROCESSING DEVICE AND SPEECH CONTROL SYSTEM

Non-Final OA §103
Filed
Jan 15, 2025
Priority
May 31, 2023 — CN 202310639508.5 +1 more
Examiner
HUTCHESON, CODY DOUGLAS
Art Unit
Tech Center
Assignee
BOE Technology Group Co., Ltd.
OA Round
1 (Non-Final)
62%
Grant Probability
Moderate
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
20 granted / 32 resolved
+2.5% vs TC avg
Strong +38% interview lift
Without
With
+37.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
28 currently pending
Career history
66
Total Applications
across all art units

Statute-Specific Performance

§101
31.4%
-8.6% vs TC avg
§103
45.3%
+5.3% vs TC avg
§102
14.1%
-25.9% vs TC avg
§112
5.4%
-34.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 32 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed with the application. Information Disclosure Statement The information disclosure statement (IDS) submitted on 07/03/2025 was filed in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Election/Restriction 1. Restriction to one of the following inventions is required under 35 U.S.C. 121: I. Claims 1-7, 12-14, and 17-21, drawn to a method of keyword spotting utilizing a compute-in-memory chip and a programmable logic unit, classified in G10L 15/22. II. Claims 8-10 and 15-16, drawn to a method of training a keyword spotting model using data augmentation, classified in G10L 15/063. The inventions are independent or distinct, each from the other because: Inventions I and II are related as subcombinations disclosed as usable together in a single combination. The subcombinations are distinct if they do not overlap in scope and are not obvious variants, and if it is shown that at least one subcombination is separately usable. In the instant case, subcombination II has separate utility, as training a keyword spotting model utilizing data augmentation does not require the implementation of the system using a programmable logic unit and computer-in-memory chip architecture as claim in subcombination I. See MPEP § 806.05(d). The examiner has required restriction between subcombinations usable together. Where applicant elects a subcombination and claims thereto are subsequently found allowable, any claim(s) depending from or otherwise requiring all the limitations of the allowable subcombination will be examined for patentability in accordance with 37 CFR 1.104. See MPEP § 821.04(a). Applicant is advised that if any claim presented in a divisional application is anticipated by, or includes all the limitations of, a claim that is allowable in the present application, such claim may be subject to provisional statutory and/or nonstatutory double patenting rejections over the claims of the instant application. Restriction for examination purposes as indicated is proper because all the inventions listed in this action are independent or distinct for the reasons given above and there would be a serious search and/or examination burden if restriction were not required because one or more of the following reasons apply: The inventions have different fields of search as they would require employing different search queries to find pertinent prior art. For example, Invention I would require search queries involving terms/concepts such as “programmable logic unit” and “compute-in-memory” architectures, whereas Invention II would require search queries involving terms/concepts such as “data augmentation”. Applicant is advised that the reply to this requirement to be complete must include (i) an election of an invention to be examined even though the requirement may be traversed (37 CFR 1.143) and (ii) identification of the claims encompassing the elected invention. The election of an invention may be made with or without traverse. To reserve a right to petition, the election must be made with traverse. If the reply does not distinctly and specifically point out supposed errors in the restriction requirement, the election shall be treated as an election without traverse. Traversal must be presented at the time of election in order to be considered timely. Failure to timely traverse the requirement will result in the loss of right to petition under 37 CFR 1.144. If claims are added after the election, applicant must indicate which of these claims are readable upon the elected invention. Should applicant traverse on the ground that the inventions are not patentably distinct, applicant should submit evidence or identify such evidence now of record showing the inventions to be obvious variants or clearly admit on the record that this is the case. In either instance, if the examiner finds one of the inventions unpatentable over the prior art, the evidence or admission may be used in a rejection under 35 U.S.C. 103 or pre-AIA 35 U.S.C. 103(a) of the other invention. During a telephone conversation with Attorney Scott Houtteman on 08/28/2026 a provisional election was made with traverse to prosecute Invention I, claims 1-7, 12-14, and 17-21. Affirmation of this election must be made by applicant in replying to this Office action. Claims 8-10 and 15-16 are withdrawn from further consideration by the examiner, 37 CFR 1.142(b), as being drawn to a non-elected invention. Applicant is reminded that upon the cancelation of claims to a non-elected invention, the inventorship must be corrected in compliance with 37 CFR 1.48(a) if one or more of the currently named inventors is no longer an inventor of at least one claim remaining in the application. A request to correct inventorship under 37 CFR 1.48(a) must be accompanied by an application data sheet in accordance with 37 CFR 1.76 that identifies each inventor by his or her legal name and by the processing fee required under 37 CFR 1.17(i). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 2. Claims 1-2, 12-14, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wu & Huang (US 2023/0022800 A1, hereinafter Wu) in view of Dbouk et al. (NPL “KeyRAM: A 0.34 uJ/decision 18k decision/s Recurrent Attention In-memory Processor for Keyword Spotting”, hereinafter Dbouk). Regarding claim 1, Wu discloses A speech data processing method comprising: acquiring speech data to be processed (para. 0034 “The user device may include (or be in communication with) two or more microphones 107, 107a-n to capture the utterance 120 from the user 10. Each microphone 107 may separately record the utterance 120 on a separate dedicated channel 119 of the multi-channel streaming audio 118.”); inputting the speech data to be processed into a programmable logic unit (Fig. 2B, input speech data input to end to end model 300; para. 0066 “The processes and logic flows described in this specification can be performed by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).”) and performing a speech feature extraction on the speech data to be processed so as to obtain a speech feature vector corresponding to the speech data to be processed (para. 0046 “In some examples, the audio features 510 for each input frame 210 are converted from raw audio signals 502 of a channel 119 of the multi-channel audio stream 118 during a pre-processing stage 504 (i.e., a feature extraction or feature generation stage). The audio features 510 may include one or more log-filterbanks. Thus, the pre-processing stage may segment the audio stream channel 119 into the sequence of input frames 210 (e.g., 30 ms each), and generate separate log-filterbanks for each frame 210. For example, each frame 210 may be represented by forty log-filterbanks.”; Fig. 5A, 510); converting the speech feature vector into a multi-channel input feature map in the programmable logic unit (para. 0046 “Each SVDF processing cell 304 of the input 3D SVDF layer 302 receives the sequence of input frames 210 from a respective channel 119. Moreover, each successive layer (e.g., SVDF layers 350) receives, as input, the concatenated filtered audio features 510 (i.e., output 344) with respect to time.”; para. 0055 “Also for each input frame, the method 800, at step 806, includes generating, by the data processing hardware 103, using an intermediate layer 410 of the memorized neural network 300, a corresponding multi-channel audio feature representation 420 based on a concatenation of the respective audio features 344 of each channel 119 of the streaming multi-channel audio 118.”);…performing speech keyword spotting on the input feature map by using a speech keyword spotting model…to obtain a speech keyword spotting result (para. 0044 “The sequentially-stacked SVDF layers 350 may generate a probability score 360 indicating a presence of a hotword in the streaming multi-channel audio 118 based on the corresponding multi-channel audio feature representation 420 of each input frame 210. The sequentially-stacked SVDF layers 350 include an initial SVDF layer 350a configured to receive the corresponding multi-channel audio feature representation 420. …In some examples, the final layer 350n of the memorized neural network 300 outputs a probability score 360 indicating the probability that the utterance 120 includes the hotword. The system 100 may determine that the utterance 120 includes the hotword when the probability score satisfies a hotword detection threshold and initiate a wake-up process on the user device 102.”;para. 0056 “The sequentially-stacked SVDF layers 350 include an initial SVDF layer 350 configured to receive the corresponding multi-channel audio feature representation 420 of each input frame 210 in sequence. At step 810, the method 800 includes determining, by the data processing hardware 103, whether the probability score 360 satisfies a hotword detection threshold…”); and feedback back the speech keyword spotting result to a main control system, wherein the main control system is configured to execute a response operation corresponding to the speech keyword spotting result according to the speech keyword spotting result (para. 0032 “Optionally, the trained memorized neural network 300 may additionally or alternatively reside in an automatic speech recognizer (ASR) 108 of the user device 102 and/or the remote system 110 to confirm that the multi-channel hotword detector 106 correctly detected the presence of a hotword in the multi-channel streaming audio 118.”; para. 0033 “In the example shown, when the user 10 speaks an utterance 120 including a hotword (e.g., “Hey Google”) captured as multi-channel streaming audio 118 by the user device 102, the memorized neural network 300 executing on the user device 102 is configured to detect the presence of the hotword in the utterance 120 to initiate a wake-up process on the user device 102 for processing the hotword and/or one or more other terms (e.g., query or command) following the hotword in the utterance 120.”). Wu does not specifically disclose inputting the input feature map into a computer-in-memory chip, [and performing speech keyword spotting on the input feature map by using a speech keyword spotting model], which is configured on the computer-in-memory chip in advance, [to obtain a speech keyword spotting result]. Dbouk teaches inputting the input feature map into a computer-in-memory chip (input features (MFCC) input to Recurrent Attention Model realized on im-memory computing chips; pg. 2, 1st para. “To enable audio classification using RAM, we use Mel-frequency Cepstral Coefficient (MFCC) features as RAM inputs…”; see Fig. 3a for RAM model, see Fig. 3b for compute-in-memory chip architecture), [and performing speech keyword spotting on the input feature map by using a speech keyword spotting model], which is configured on the computer-in-memory chip in advance, [to obtain a speech keyword spotting result] (KeyRAM architecture stores weights on chip in advance for keyword spotting, see pg. 2, section III, ‘Chip Architecture’ discussing IMC blocks; KeyRAM model implemented on IMC chips outputs class scores predicting a keyword spoken). Wu and Dbouk are considered to be analogous to the claimed invention as they both are in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu to incorporate the teachings of Dbouk in order to input the input feature map to a computer-in-memory chip and to perform keyword spotting using a model configured on the computer-in-memory chip in advance. Doing so would be beneficial, as this would provide significant savings in enery/decision while achieving high throughput for decisions (Dbouk, Abstract). Regarding claim 2, Wu discloses wherein the performing the speech feature extraction on the speech data to be processed so as to obtain the speech feature vector corresponding to the speech data to be processed, comprises: performing the speech feature extraction on the speech data to be processed with a preset speech feature extraction algorithm so as to obtain the speech feature vector corresponding to the speech data to be processed (Wu, see Fig. 5A, “Feature Generation” 504; para. 0046 “In some examples, the audio features 510 for each input frame 210 are converted from raw audio signals 502 of a channel 119 of the multi-channel audio stream 118 during a pre-processing stage 504 (i.e., a feature extraction or feature generation stage). The audio features 510 may include one or more log-filterbanks. Thus, the pre-processing stage may segment the audio stream channel 119 into the sequence of input frames 210 (e.g., 30 ms each), and generate separate log-filterbanks for each frame 210.”). Regarding claim 12, Wu discloses A speech control system, comprising: a sound acquisition device configured to acquire speech data to be processed in an environment where the sound acquisition device is located, and send the speech data to be processed to a main control system (para. 0034 “The user device may include (or be in communication with) two or more microphones 107, 107a-n to capture the utterance 120 from the user 10. Each microphone 107 may separately record the utterance 120 on a separate dedicated channel 119 of the multi-channel streaming audio 118.”); the main control system configured to receive the speech data to be processed acquired by the sound acquisition device, and send the speech data to be processed to a programmable logic unit (para. 0034 “For example, the user device 102 may include two microphones 107 that each record the utterance 120, and the recordings from the two microphones 107 may be combined into two-channel streaming audio 118 (i.e., stereophonic audio or stereo). In some examples, the user device 102 may include more than two microphones. That is, the two microphones reside on the user device 102. Additionally or alternatively, the user device 102 may be in communication with two or more microphones separate/remote from the user device 102. For example, the user device 102 may be a mobile device disposed within a vehicle and in wired or wireless communication (e.g., Bluetooth) with two or more microphones of the vehicle. In some configurations, the user device 102 is in communication with least one microphone 107 residing on a separate device 101, which may include, without limitation, an in-vehicle audio system, a computing device, a speaker, or another user device. In these configurations, the user device 102 may also be in communication with one or more microphones residing on the user device 102.”; Fig. 2B, input speech data input to end to end model 300; para. 0066 “The processes and logic flows described in this specification can be performed by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).”); the programmable logic unit configured to perform speech feature extraction on the speech data to be processed to obtain a speech feature vector corresponding to the speech data to be processed (para. 0046 “In some examples, the audio features 510 for each input frame 210 are converted from raw audio signals 502 of a channel 119 of the multi-channel audio stream 118 during a pre-processing stage 504 (i.e., a feature extraction or feature generation stage). The audio features 510 may include one or more log-filterbanks. Thus, the pre-processing stage may segment the audio stream channel 119 into the sequence of input frames 210 (e.g., 30 ms each), and generate separate log-filterbanks for each frame 210. For example, each frame 210 may be represented by forty log-filterbanks.”; Fig. 5A, 510), and convert the speech feature vector into a multi-channel input feature map (para. 0046 “Each SVDF processing cell 304 of the input 3D SVDF layer 302 receives the sequence of input frames 210 from a respective channel 119. Moreover, each successive layer (e.g., SVDF layers 350) receives, as input, the concatenated filtered audio features 510 (i.e., output 344) with respect to time.”; para. 0055 “Also for each input frame, the method 800, at step 806, includes generating, by the data processing hardware 103, using an intermediate layer 410 of the memorized neural network 300, a corresponding multi-channel audio feature representation 420 based on a concatenation of the respective audio features 344 of each channel 119 of the streaming multi-channel audio 118.”); …a speech keyword spotting model …, and configured to perform speech keyword spotting on the input feature map by using a speech keyword spotting model … to obtain a speech keyword spotting result (para. 0044 “The sequentially-stacked SVDF layers 350 may generate a probability score 360 indicating a presence of a hotword in the streaming multi-channel audio 118 based on the corresponding multi-channel audio feature representation 420 of each input frame 210. The sequentially-stacked SVDF layers 350 include an initial SVDF layer 350a configured to receive the corresponding multi-channel audio feature representation 420. …In some examples, the final layer 350n of the memorized neural network 300 outputs a probability score 360 indicating the probability that the utterance 120 includes the hotword. The system 100 may determine that the utterance 120 includes the hotword when the probability score satisfies a hotword detection threshold and initiate a wake-up process on the user device 102.”;para. 0056 “The sequentially-stacked SVDF layers 350 include an initial SVDF layer 350 configured to receive the corresponding multi-channel audio feature representation 420 of each input frame 210 in sequence. At step 810, the method 800 includes determining, by the data processing hardware 103, whether the probability score 360 satisfies a hotword detection threshold…”); the main control system further configured to receive the speech keyword spotting result …and control a terminal device to execute a corresponding response operation according to the speech keyword spotting result (para. 0032 “Optionally, the trained memorized neural network 300 may additionally or alternatively reside in an automatic speech recognizer (ASR) 108 of the user device 102 and/or the remote system 110 to confirm that the multi-channel hotword detector 106 correctly detected the presence of a hotword in the multi-channel streaming audio 118.”; para. 0033 “In the example shown, when the user 10 speaks an utterance 120 including a hotword (e.g., “Hey Google”) captured as multi-channel streaming audio 118 by the user device 102, the memorized neural network 300 executing on the user device 102 is configured to detect the presence of the hotword in the utterance 120 to initiate a wake-up process on the user device 102 for processing the hotword and/or one or more other terms (e.g., query or command) following the hotword in the utterance 120.”); and the terminal device configured to execute the response operation corresponding to the speech keyword spotting result (para. 0033 “In the example shown, when the user 10 speaks an utterance 120 including a hotword (e.g., “Hey Google”) captured as multi-channel streaming audio 118 by the user device 102, the memorized neural network 300 executing on the user device 102 is configured to detect the presence of the hotword in the utterance 120 to initiate a wake-up process on the user device 102 for processing the hotword and/or one or more other terms (e.g., query or command) following the hotword in the utterance 120.”). Regarding claim 13, claim 13 is a device claim with limitations similar to those recited in method claim 1 and thus is rejected under similar rationale. Additionally, Wu discloses An electronic device (Fig. 9), comprising: at least one processor (Fig. 9, 910); and a memory communicatively connected to the at least one processor (Fig. 9, 930; para. 0058 “Each of the components 910, 920, 930, 940, 950, and 960, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate.”); wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to cause the at least one processor to perform the speech data processing method according to claim 1 (para. 0060 “The storage device 930 is capable of providing mass storage for the computing device 900. In some implementations, the storage device 930 is a computer-readable medium. …The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 920, the storage device 930, or memory on processor 910.”; see claim mapping for claim 1). Regarding claim 14, claim 14 is a CRM claim with limitations similar to those recited in method claim 1, and thus is rejected under similar rationale. Additionally, Wu discloses A non-transitory computer-readable storage medium, on which a computer program is stored, wherein the computer program, when being executed by a processor, executes the speech data processing method according to claim 1 (para. 0060 “The storage device 930 is capable of providing mass storage for the computing device 900. In some implementations, the storage device 930 is a computer-readable medium. In various different implementations, the storage device 930 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 920, the storage device 930, or memory on processor 910.”; see claim mapping for claim 1). Regarding claim 17, claim 17 is rejected for analogous reasons to claim 2. 3. Claims 3 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Dbouk, and further in view of Wang et al. (NPL “Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for M2MET Challenge”, hereinafter Wang). Regarding claim 3, Wu in view of Dbouk discloses wherein the speech vector is a two-dimensional speech feature vector (Wu, see Fig. 5A, each frame has a corresponding one-dimensional tensor audio feature (i.e. log-filterbanks); para. 0046 “In some examples, the audio features 510 for each input frame 210 are converted from raw audio signals 502 of a channel 119 of the multi-channel audio stream 118 during a pre-processing stage 504 (i.e., a feature extraction or feature generation stage). The audio features 510 may include one or more log-filterbanks. Thus, the pre-processing stage may segment the audio stream channel 119 into the sequence of input frames 210 (e.g., 30 ms each), and generate separate log-filterbanks for each frame 210.”; thus, the speech feature vector can be interpreted as two-dimensional (i.e. F number of frames x D number of log-filterbanks per frame)) but does not specifically disclose and the converting the speech feature vector into the multi-channel input feature map comprises: performing feature dimension expansion on the speech feature vector to obtain a three-dimensional speech feature vector; copying the three-dimensional speech feature vector to obtain a plurality of three-dimensional speech feature vectors; and splicing the plurality of three-dimensional speech feature vectors according to a channel dimension to obtain the multi-channel input feature map, wherein the channel dimension refers to a feature channel of the speech feature vectors. Wang teaches the converting the speech feature vector into the multi-channel input feature map comprises: performing feature dimension expansion on the speech feature vector to obtain a three-dimensional speech feature vector (pg. 3, 2nd para., frame-level speaker embedding is three dimensional feature generated from a speech feature vector (i.e. Fbank sequence X): “…Next, the front-end ResNet takes the C-channels of Fbank sequence as input and produce a frame-level speaker embedding sequence…”; the frame level speaker embedding is a sequence of two dimensional vectors TxD; thus the sequence as a whole is interpreted as three dimensional (C number of channels x T x D)); copying the three-dimensional speech feature vector to obtain a plurality of three-dimensional speech feature vectors (pg. 3 2nd para. “…Later, the target-speaker embedding is repeated T times and the frame-level speaker embedding is repeated N times…”); and splicing the plurality of three-dimensional speech feature vectors according to a channel dimension to obtain the multi-channel input feature map, wherein the channel dimension refers to a feature channel of the speech feature vectors (pg. 3, 2nd para., embedding sequence Sin generated from N repeated frame-level speaker embeddings; this embedding splices multiple three dimensional vectors (N number of frame-level speaker embeddings) according to a channel dimension (C number of these vectors concatenated in embeddings sequence) to obtain a four dimensional input feature map Sin, see Eq. 1). Wu, Dbouk, and Wang are considered to be analogous to the claimed invention as they are in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu in view of Dbouk to incorporate the teachings of Wang in order to generate the input feature map by performing feature dimension expansion on the speech feature evector to obtain a three-dimensional speech feature vector, copying the three-dimensional speech feature vector to obtain a plurality of three-dimensional speech feature vectors, and splicing the plurality of vector according to a channel dimension to obtain the multi-channel input feature map. Doing so would be beneficial, as this would allow for non-linear spatial correlations between different channels to be learned and fused, improving performance on speech processing tasks (Wang, see Abstract, Introduction). Regarding claim 18, claim 18 is rejected for analogous reasons to claim 3. 4. Claims 4-6 and 19-21 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Dbouk, and further in view of Ng et al. (NPL “ConvMixer: Feature Interactive Convolution With Curriculum Learning for Small Footprint and Noisy Far-Field Keyword Spotting”, hereinafter Ng) and Trockman & Kolter (NPL “Patches Are All You Need?”, hereinafter Trockman). Regarding claim 4, Wu in view of Dbouk discloses the performing the speech keyword spotting on the input feature map by using the speech keyword spotting model, which is configured on the computer- in-memory chip in advance, to obtain the speech keyword spotting result (see above mapping for claim 1), but does not specifically disclose wherein the speech keyword spotting model is a model constructed based on a convolution mixer network ConvMixer architecture, and comprises a feature sampling layer, the convolution mixer network and a classification output layer; [and the performing the speech keyword spotting on the input feature map by using the speech keyword spotting model, which is configured on the computer- in-memory chip in advance, to obtain the speech keyword spotting result], comprises: …obtain a … feature map;performing a convolution mixer process on the … feature map by utilizing the convolution mixer network to obtain a speech recognition feature map; and performing a keyword classification prediction process on the speech recognition feature map by using the classification output layer to obtain the speech keyword spotting result. Ng teaches wherein the speech keyword spotting model is a model constructed based on a convolution mixer network ConvMixer architecture (Fig. 1, see ConvMixer Block), and comprises a feature sampling layer (Fig. 1, see ‘Pre-Block 1-3’), the convolution mixer network (Fig. 1, see ‘ConvMixer Block 1-4’) and a classification output layer (Fig. 1, see ‘Linear - Sigmoid’ layer) and further teaches the process of …obtain a … feature map (Fig. 1, Pre-Blocks 1-4 obtain a feature map to be input to ConvMixer layers); performing a convolution mixer process on the … feature map by utilizing the convolution mixer network to obtain a speech recognition feature map (Fig. 1, ConvMixer Blocks 1-4 process the input from the Pre-Convolutional Blocks to obtain output to Post-Block and classifier; see also pg. 2, section 3.1 for implementation); and performing a keyword classification prediction process on the speech recognition feature map by using the classification output layer to obtain the speech keyword spotting result (Fig. 1, Post-Block takes output of ConvMixer layers and process with Linear layer to output class prediction). Wu, Dbouk, and Ng are considered to be analogous to the claimed invention as they are in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu in view of Dbouk to incorporate the teachings of Ng in order to utilize a ConvMixer architecture and to obtain the keyword spotting result by obtaining a feature map, performing a convolutional mixer process on the feature map by utilizing the convolutional mixer network to obtain a speech recognition feature map, and to perform a keyword classification prediction based on the speech recognition feature map by using a classification output layer to obtain the speech keyword spotting result. Doing so would be beneficial, as such architecture would yield temporally and frequency rich embeddings for use in keyword spotting (pg. 2, section 3.1 2nd para.) and would match performance of other KWS models while requiring significantly less memory consumption and computing resources (pg. 4, section 5). Wu in view of Dbouk and Ng does not specifically disclose performing a down-sampling process on the input feature map by using the feature sampling layer to obtain a down-sampled feature map; [performing a convolution mixer process on the ]down-sampled feature map [by utilizing the convolution mixer network to obtain a speech recognition feature map]. Trockman teaches performing a down-sampling process on the input feature map by using the feature sampling layer to obtain a down-sampled feature map (Fig. 2, patch embedding generated which downsamples input feature map; pg. 5, section 5 , 2nd para. “…Patch embeddings allow all the downsampling to happen at once, immediately decreasing the internal resolution…”); [performing a convolution mixer process on the ]down-sampled feature map [by utilizing the convolution mixer network to obtain a speech recognition feature map] (Fig. 2, downsampled feature map (output of patch embedding + GELU + BatchNorm) is input to the ConvMixer layer to obtain a feature map). Wu, Dbouk, Ng, and Trockman are considered to be analogous to the claimed invention as Wu, Dbouk, and Ng are in the same field of speech processing, and Trockman is in the same field of machine learning architectures. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu in view of Dbouk and Ng to incorporate the teachings of Trockman in order to perform a downsampling on the input feature map by using the feature sampling layer to obtain a down-sampled feature map, and to perform the convolutional mixer process on the down-sampled feature map. Doing so would be beneficial, as patch embeddings allow for all the downsampling to happen at once, immediately decreasing the internal resolution and thus increasing the effective receptive field size, making it easier to mix distant spatial information (see pg. 5, section 5, 2nd para.). Regarding claim 5, Wu in view of Dbouk, Ng, and Trockman discloses wherein the feature sampling layer comprises: a patch embedding layer (Trockman, Fig. 2, patch embedding layer used), a first batch normalization layer (Trockman, Fig. 2, ‘BatchNorm’ layer succeeding Patch Embedding layer) and a first activation function layer (pg. 2, Patch Embedding comprises batch normalization of Convolution with activation function applied ‘σ’); and wherein the patch embedding layer is realized based on a convolution layer (pg. 2, section 2, 1st para. “…Patch embeddings with patch size and embedding dimension h can be implemented as convolution with cin input channels, h output channels, kernel size p, and stride p…”), and the activation function layer adopts a Swish activation function (Ng teaches use of specifically a Swish activation function in this layer: see Ng, Fig. 1, Pre-Block use of ‘Swish’; see also pg. 2, section 3.1, 1st para. for implementation). Wu, Dbouk, Ng, and Trockman are considered to be analogous to the claimed invention as Wu, Dbouk, and Ng are in the same field of speech processing, and Trockman is in the same field of machine learning architectures. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Trockman in order to have the feature sampling layer comprise a patch embedding layer, a first batch normalization layer and a first activation function layer; and wherein the patch embedding layer is realized based on a convolution layer. Doing so would be beneficial, given the same rationale as given for claim 4. Further, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Ng in order to have the activation function layer adopt specifically a Swish activation layer. Doing so would be beneficial, as this activation function has been found to work better than ReLU on deeper machine learning models across a number of challenging datasets (see NPL Ramachandran et al., “Searching for Activation Functions”, Abstract). Regarding claim 6, Wu in view of Dbouk, Ng, and Trockman discloses wherein the convolution mixer network comprises at least one convolution mixer network module (Trockman, Fig. 2 shows ConvMixer layer x depth), each convolution mixer network module comprises a space position convolution mixer unit (Trockman, Fig. 2, depthwise convolution mixes spatial locations; pg. 3, 2nd para. “…In particular, we chose depthwise convolution to mix spatial locations…”) and a channel position convolution mixer unit (Trockman, Fig. 2, pointwise convolution mixes channel locations; pg. 3, 2nd para. “…In particular, we chose depthwise convolution to mix spatial locations and pointwise convolution to mix channel locations…”), and a space position mixer processing result output by the space position convolution mixer unit and an input of the space position convolution mixer unit are input to the channel position convolution mixer unit through a residual connection for a channel position mixer processing (Fig. 2, output of space position convolutional mixer unit (i.e. ‘BatchNorm’ connected to Depthwise convolution + GELU) and input of the space position convolutional mixer unit (i.e. ‘BatchNorm’ connected to “Patch Embedding’ + GELU) are input to the channel positional convolutional mixer unit (i.e. ‘Pointwise Convolution’ layer) via a residual connection); wherein the space position convolution mixer unit comprises: a depthwise separable convolution layer (Fig. 2, see Depthwise Convolution block), a second batch normalization layer (Fig. 2, see ‘BatchNorm’ following ‘Depthwise Convolution’ + GELU) and a second activation function layer which are connected sequentially (Fig. 2, see GELU connected sequentially with ‘Depthwise Convolution’ and ‘BatchNorm’ layers); the channel position convolution mixer unit comprises: a pointwise convolution layer (Trockman, Fig. 2, see ‘Pointwise Convolution’ layer), a third batch normalization (Fig. 2, see ‘BatchNorm’ following ‘Pointwise Convolution’ + GELU) layer and a third activation function layer which are connected sequentially (Fig. 2, see GELU connected sequentially with ‘Depthwise Convolution’ and ‘BatchNorm’ layers); and wherein the second activation function layer and the third activation function layer both adopt Swish activation functions (Ng teaches second and third Swish activation functions in ConvMixer blocks, see Fig. 1, Swish in frequency and temporal domain layers). Wu, Dbouk, Ng, and Trockman are considered to be analogous to the claimed invention as Wu, Dbouk, and Ng are in the same field of speech processing, and Trockman is in the same field of machine learning architectures. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Trockman in order to have the convolution mixer network comprises at least one convolution mixer network module, each convolution mixer network module comprises a space position convolution mixer unit and a channel position convolution mixer unit, and a space position mixer processing result output by the space position convolution mixer unit and an input of the space position convolution mixer unit are input to the channel position convolution mixer unit through a residual connection for a channel position mixer processing; wherein the space position convolution mixer unit comprises: a depthwise separable convolution layer, a second batch normalization layer and a second activation function layer which are connected sequentially; the channel position convolution mixer unit comprises: a pointwise convolution layer, a third batch normalization layer and a third activation function layer which are connected sequentially. Doing so would be beneficial, as such architecture is a simple class of model that independently mixes spatial and channel locations using only standard convolutions, with large kernel sizes providing substantial performances boosts, and providing a powerful template for deep learning (see Trockman, pg. 5, section 5). Further, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Ng in order to have the second and third activation function layers both adopt specifically a Swish activation layer. Doing so would be beneficial, given the same rationale as given for claim 5. Regarding claim 19, claim 19 is rejected for analogous reasons to claim 4. Regarding claim 20, claim 20 is rejected for analogous reasons to claim 5. Regarding claim 21, claim 21 is rejected for analogous reasons to claim 6. 5. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Dbouk, Ng, and Trockman, and further in view of Shelhamer et al. (NPL “Fully Convolutional Networks for Semantic Segmentation”, hereinafter Shelhamer). Regarding claim 7, Wu in view of Dbouk, Ng, and Trockman discloses wherein the classification output layer comprises: an average pooling layer and a fully connected layer (Trockman, see Fig. 2, ‘Global Average Pooling’ layer and ‘Fully-Connected’ layer). Wu, Dbouk, Ng, and Trockman are considered to be analogous to the claimed invention as Wu, Dbouk, and Ng are in the same field of speech processing, and Trockman is in the same field of machine learning architectures. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Trockman in order to have the classification output layer comprise an average pooling layer and a fully connected layer. Doing so would be beneficial, as this would enable the output of the ConvMixer layer to be process into class predictions for classification tasks. Wu in view of Dbouk, Ng, and Trockman does not specifically disclose wherein the fully connected layer is realized by adopting a convolution layer. Shelhamer teaches wherein the fully connected layer is realized by adopting a convolution layer (see Fig. 2, fully connected layers realized as convolutional layers instead; pg. 3, section 3.1 “…The fully connected layers of these nets have fixed dimensions and throw away spatial coordinates. However, fully connected layers can also be viewed as convolutions with kernels that cover their entire input regions…”). Wu, Dbouk, Ng, Trockman, and Shelhamer are considered to be analogous to the claimed invention as Wu, Dbouk, and Ng are in the same field of speech processing, and Trockman and Shelhamer are in the same field of machine learning architectures. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Shelhamer in order to have the fully connected layer be realized by adopting a convolutional layer. Doing so would be beneficial, as doing so would enable a fully convolutional network which enables these layers to take input of any size, with the higher computational efficiency (pg. 3, section 3.1, 1st para- 4th para.). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Elders et al. (US 11,693,622 B1): keyword detection to perform actions (Fig. 3) utilizing keyword models (Fig. 2B) Yang et al. (US 2024/0339123 A1): keyword spotting utilizing a convmixer architecture (para. 0062) Ahn et al. (US 2021/0074270 A1): obtaining input feature map from speech, utilizing convolutional keyword spotting model to extract voice keyword (Fig. 4) D’Amato & Smith (US 2021/0035571 A1): control device, playback device, and computing device for keyword spotting (Fig. 6, Fig. 7A) Song et al. (NPL “Knowledge Distillation for In-Memory Keyword Spotting Model”): use of im-memory computing encoder for keyword spotting (pg. 2, section 3.1) Shaefer et al. (NPL “LSTMs for Keyword Spotting with ReRAM-based Compute-In-Memory Architectures”): compute-in-memory architecture implementing keyword spotting (Fig. 2) Gharbieh et al. (NPL “DyConvMixer: Dynamic Convolution Mixer Architecture for Open-Vocabulary Keyword Spotting”): ConvMixer architecture for keyword spotting (Fig. 1) Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CODY DOUGLAS HUTCHESON/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jan 15, 2025
Application Filed
Sep 09, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743450
QUERY FORMATTING SYSTEM, QUERY FORMATTING METHOD, AND INFORMATION STORAGE MEDIUM
3y 6m to grant Granted Sep 22, 2026
Patent 12737563
REGIONAL SIGN LANGUAGE TRANSLATION
3y 5m to grant Granted Sep 15, 2026
Patent 12718828
AUDIO SOURCE CLASSIFICATION FOR HANDSFREE COMMUNICATIONS
3y 5m to grant Granted Aug 25, 2026
Patent 12664970
SPEECH TRANSLATION WITH PERFORMANCE CHARACTERISTICS
3y 2m to grant Granted Jun 23, 2026
Patent 12626715
ROLE SEPARATION METHOD, ELECTRONIC DEVICE, AND COMPUTER STORAGE MEDIUM
3y 4m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
62%
Grant Probability
99%
With Interview (+37.5%)
2y 9m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 32 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month