Prosecution Insights
Last updated: August 06, 2026
Application No. 18/260,196

DISTRIBUTED SPEECH PROCESSING SYSTEM AND METHOD

Non-Final OA §101§112
Filed
Jun 30, 2023
Priority
Dec 31, 2020 — CN 202011628865.4 +1 more
Examiner
OGUNBIYI, OLUWADAMILOL M
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Espressif Systems (Shanghai) Co. Ltd.
OA Round
3 (Non-Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
242 granted / 314 resolved
+15.1% vs TC avg
Strong +19% interview lift
Without
With
+19.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
23 currently pending
Career history
342
Total Applications
across all art units

Statute-Specific Performance

§101
21.1%
-18.9% vs TC avg
§103
49.1%
+9.1% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
13.8%
-26.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 314 resolved cases

Office Action

§101 §112
DETAILED ACTION Claims 1, 2, 4 – 6, 8, 11, 12, 14 – 16 and 18 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment With regard to the Final Office Action from 04 February 2026, the Applicant has filed a response on 24 April 2026. Claims 3, 7, 9, 10, 13, 17, 19 and 20 have been cancelled. With regard to the 35 U.S.C. 112(f) interpretation given to limitations of claim 1, the Applicant has now amended this claim to indicate that ‘the sound acquisition device’ is a ‘microphone array or a multi-channel microphone’ and this is suitable to overcome the claim interpretation. The previous claims were rejected under 35 U.S.C. 101 for being directed to a judicial exception without significantly more. The independent claims have now been amended to provide, particularly by their last limitation, obtaining a final speech recognition result based on a first and a second speech recognition results that get dynamic weighting applied to them by assigning a weight according to the sound quality of each of the two speech recognition results a sound pickup distance of the node devices. This limitation adds significantly more to the mere mental process of performing speech recognition, presenting a particular and unique approach to the idea of obtaining speech recognition results. The Examiner hereby, based on this, withdraws the 35 U.S.C. 101 rejection. Response to Arguments The Applicant has amended claim 1 in an attempt to overcome the 35 U.S.C. 112(f) interpretation, further providing (Remarks: pages 9 – 10) that the amendment is suitable to obviate this interpretation, stating that sufficient structure and algorithmic detail is provided which fully satisfy the 35 U.S.C. 112(f). The Examiner however holds that the amendment only addresses the ‘speech acquisition device’ but is yet to properly address the ‘sound processing module’ and the ‘communication module’ other than maintaining that they are covered under a node device in a network. The Examiner maintains the 112(f) interpretation in this regard. With regard to the 35 U.S.C. 103 rejection given to the claims, the Applicant has amended the claims, and the Applicant’s arguments on the applied prior art of reference are directed to the amendments to the independent claims, stating that the applied prior art do not teach the independent claims as currently amended. Applicant’s arguments with respect to the independent claims have been considered but are moot due to the new grounds of rejection necessitated by the amendment to the claims. The claims will be addressed by their current presentation in the following sections. Claim Objections Claims 4, 5, 6, 8, 14, 15, 16 and 18 are objected to because of the following informalities: These claims are presented as being dependent on claims that have been cancelled. Claims 4, 5, 6 and 8 are presented as being dependent on claim 3 which is currently cancelled, while claims 14, 15, 16 and 18 are presented as being dependent on claim 13, also currently cancelled. For the purpose of this Office Action, the Examiner will consider claims 4, 5, 6 and 8 as being dependent on claim 1, while considering claims 14, 15, 16 and 18 as being dependent on claim 11. The dependency of these claims should be amended. Appropriate correction is required. Claim Interpretations The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: ‘the sound processing module is configured to preprocess the electrical signal …’ in claim 1; and ‘the communication module is configured to send …’ in claim 1. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover a processor and communication modules such as those capable of WiFi/BLE/Zigbee communication protocols (page 8 lines 6 – 17) as the corresponding structure described in the Specification as performing the claimed function, and equivalents thereof. If Applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, Applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recites sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the Applicant regards as his invention. Claims 1, 2, 4 – 6, 8, 11, 12, 14 – 16 and 18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the Applicant), regards as the invention. Regarding claims 1 and 11, the phrase “each node device has peer-to-peer direct communication capability and can act as a router to forward signals” renders the claim indefinite because it is unclear as to if each node device actually is performing the function of a router, or if it simply can perform the function of a router but does not actually perform the function of a router. As it is currently presented, it only indicates the possibility of functioning as a router without the precise recitation of functioning as a router to forward signals. This recitation renders the claims indefinite. See MPEP § 2173.05(d). Dependent claims 2, 4 – 6, 8, 12, 14 – 16 and 18 are also rejected under 35 U.S.C. 112(b) based on their dependence on their respectively rejected base claims, in the absence of addressing the rejection raised for their respective base claims. Claims 1 and 11 recite the limitations ‘the local node device’ and ‘the remote node device’ in their final limitations. There is insufficient antecedent basis for these limitations in the claims. Claims 2, 4 – 6, 8, 12, 14 – 16 and 18 are also hereby rejected based on their dependence on the rejected claims above. Allowable Subject Matter Claims 1 and 11 would be allowable if rewritten or amended to overcome the rejections under 35 U.S.C. 112(b), set forth in this Office action. The following is a statement of reasons for the indication of allowable subject matter: With regard to independent claim 1, the invention states: A distributed speech processing system, comprising: a plurality of node devices in a network, wherein the network is a decentralized self-organizing Mesh network, the plurality of node devices implement ad hoc networking through wired or wireless manner, each node device has peer-to-peer direct communication capability and can act as a router to forward signals: each of the plurality of node devices comprises a processor, a memory, a communication module, and a sound processing module, and at least one of the plurality of node devices comprises a sound acquisition module; the processor is configured to provide a clock to achieve time synchronization of preprocessed data among all node devices: the sound acquisition module is a microphone array or a multi-channel microphone; wherein, the sound acquisition module is configured to acquire an audio signal and convert the audio signal into an electrical signal; the sound processing module is configured to preprocess the electrical signal through signal framing, pre-emphasis, and Fast Fourier Transform (FFT) to obtain a first sound preprocessed result; the communication module is configured to send the first sound preprocessed result to one or more node devices in the network; the communication module is further configured to receive one or more second sound preprocessed results from at least one other node device over the Mesh network; wherein the first sound preprocessed result and the second sound preprocessed result comprise data blocks with a same duration, and each of the data blocks include an incremental sequence number, a sound feature value, a sound quality, and sound time information; the sound processing module is further configured to first filter out data blocks whose sound quality exceeds a predetermined threshold, then splice the first sound preprocessed result and the one or more second sound preprocessed results in the order of incremental sequence numbers, and select the data block with the highest sound quality from the data blocks with the same incremental sequence number to form a complete third sound preprocessed result perform speech recognition based on the first sound preprocessed result and the one or more second sound preprocessed results third sound preprocessed result to obtain a first speech recognition result; wherein each of the first sound preprocessed result and the one or more second sound preprocessed results is an intermediate result of the speech recognition; the communication module is further configured to receive one or more second speech recognition results from at least one other node device over the Mesh network; and the sound processing module is further configured to perform speech recognition based on the first speech recognition result and the one or more second speech recognition results to obtain a final speech recognition result, wherein the speech recognition adopts a dynamic weighting rule, wherein a weight is assigned according to the sound quality of the recognition results and the sound pickup distance of the node devices, and a weight of the recognition result of the local node device is higher than that of the remote node device. Closest Prior Art The reference of Nakadai et al. (US 2016/0055850 A1) provides teaching for first and second speech processing devices connected in a network [0046], the devices having a preprocessing unit as a sound processing module, a database as a memory, a communication unit, containing a sound acquisition module (FIG. 1), a processor [0264], acquiring the input speech signal (FIG. 1 Parts 30, 110), a preprocessing module for performing sound source localisation (FIG. 1 Part 112), a communication unit able to transmit pre-processed speech signal to another device 20A which can perform its own further preprocessing, such as sound source localisation, etc. (FIG. 7 Part 120), the network being a local area network (indicating according to FIG. 7 Part 50, that the devices are all in a network) [0053], communicating results of pre-processed data from one preprocessing step to the next, such as from 112 to 113 to 114 and first and second speech processing units for performing speech recognition on results of first and second sound pre-processed results, with Parts 112, 113 and 114 being preprocessed results that are intermediate results on a path to speech recognition Part 116 (FIG. 7); a first speech processing device 10 and a second speech processing device 20 which are connected through a network (indicating different node devices) [0046], the second speech processing device 20 also performs speech recognition on speech received from the first speech processing device) [0050], a communication unit that receives second speech recognition results over a network 50, whereby data can be transmitted and received between both the first and second speech processing devices, these being two devices qualifying as node devices over a network (FIG 7 Part 220), a second speech processing device 20A which includes preprocessing unit and a second speech recognition unit (able to perform its own speech recognition) [0135], and a teaching that ‘The communication unit 220 transmits transmission data including the second text data input from the second speech recognition unit 216 to the first speech processing device 10’ (showing a transmission of second speech recognition result from another node device) [0078]. The reference of KIM et al. (US 2019/038614 A1) provides teaching for generating a final speech recognition result based on weighting each of the available speech recognition results [0221] and obtaining several speech recognition results, which get combined in order to obtain an aggregate speech recognition result as a final speech recognition result [0222]. KIKUCHI et al. (JP 2020/0160281 A; applying the attached English translation) teaches performing pre-processing at 112-1 of a first device and then transmitting it over to another pre-processing unit 112-2 (indicating a distributed pre-processing over different devices or nodes) (page 4 line 47 – page 5 line 6) and another pre-processing unit 112-2 which obtains a signal that was initially pre-processed by a different unit, before performing its own pre-processing (page 6 lines 20–26). Sereshki et al. (US 2020/010529 A1) provides teaching for the presence of microphone devices all arranged within a network [0105], whereby the different microphone devices in the network capture audio signals which could be misaligned, and these audio signals get aligned through synchronization as provided by a system clock [0107], whereby the devices within the network can transmit data/signals to other devices within the network, indicating that the devices can act with peer-to-peer communication capabilities, forwarding signals amongst themselves [0046]. Jeong et al. (US 2006/0053009 A1) provides teaching for passing a signal through a pre-emphasis filter as well as converting each frame into frequency domain using FFT [0093]. KANDA (US 2019/0139540 A1) provides teaching for a framing unit for framing digitized speech signal from an analogue-to-digital converter by using windows to place the signal in prescribed lengths [0068]. Huang et al. (US 2019/0295542 A1) provides teaching for synchronizing collected audio streams, generating a weighted combination of the different signals (indicating assigning weights to each of the signals), and then performing speech recognition on the weighted combination of the signals [0091]. Georganti (US 2020/0301651 A1) provides teaching for the selection of a better signal of microphone from two microphones (assigning weights to select that with a better quality) [0043], such that a microphone that is closer to the sound source gets selected (weighted higher) than one that is further away [0012], thereby teaching of a higher weight being applied to a closer/local device than a remote/farther-away device. Brown et al. (US 2016/0364681 A1) provides teaching for splicing together, the best quality audio segments together [0089]. Lanham et al. (US 2010/0299131 A1) provides teaching for splicing together an original audio recording with a supplemental audio recording to obtain a modified time-aligned transcript [0056]. Parada et al. (US 2018/0075860 A1) provides teaching for selecting a best microphone pair based on confidence measures [0006], segmenting audio signals based on clustering information [0043], the selection of the microphones being to achieve a robust ASR performance. Degani et al. (US 2009/0313018 A1) provides teaching for preprocessing speech utterances into sequences of strings of equal length blocks [0042]. Nemala et al. (US 2016/0063997 A1) provides teaching for dynamically assigning weights to each audio stream from each audio device as a user walks around a house, so as to ensure optimal audio quality and speech recognition at all times [0031]. The prior art of record taken alone or in combination however fail to teach, inter alia, a distributed processing system having node devices that acquire an audio signal such that first and second sound preprocessed results that comprise data blocks are acquired, each of the data blocks including an incremental sequence number, a sound feature value, a sound quality and sound time information, leading to the selection of data blocks with the highest sound quality to form a complete third sound preprocessed result that gets applied to obtain a first speech recognition result. Claim 1 would hereby be allowable if rewritten or amended to overcome the 35 U.S.C. 112(b) rejection. Claim 11 would also hereby be allowable if rewritten or amended to overcome the 35 U.S.C. 112(b) rejection, based on the reason set forth for claim 1 above. Claims 2 and 12 would be allowable if rewritten to overcome the rejections under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. Claims 4, 5, 6, 8, 14, 15, 16 and 18 would be allowable if rewritten to overcome the rejections under 35 U.S.C. 112(b) and the claim objections set forth in this Office action, and to include all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to Applicant’s disclosure. See the closest prior art provided in the “Allowable Subject Matter” section. Any inquiry concerning this communication or earlier communications from the Examiner should be directed to OLUWADAMILOLA M. OGUNBIYI whose telephone number is (571)272-4708. The Examiner can normally be reached Monday – Thursday (8:00 AM – 5:30 PM Eastern Standard Time). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s Supervisor, PARAS D. SHAH can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /OLUWADAMILOLA M OGUNBIYI/Examiner, Art Unit 2653 /Paras D Shah/Supervisory Patent Examiner, Art Unit 2653 06/24/2026
Read full office action

Prosecution Timeline

Jun 30, 2023
Application Filed
Jul 30, 2025
Non-Final Rejection mailed — §101, §112
Oct 23, 2025
Response Filed
Feb 04, 2026
Final Rejection mailed — §101, §112
Apr 24, 2026
Request for Continued Examination
May 04, 2026
Response after Non-Final Action
Jun 29, 2026
Non-Final Rejection mailed — §101, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694891
METHOD, DEVICE AND COMPUTER PROGRAM FOR EMOTION RECOGNITION FROM A REAL-TIME AUDIO SIGNAL
3y 7m to grant Granted Jul 28, 2026
Patent 12640154
Stylizing Text-to-Speech (TTS) Voice Response for Assistant Systems
3y 5m to grant Granted May 26, 2026
Patent 12608427
Drill Back To Original Audio Clip In Virtual Assistant Initiated Lists And Reminders
1y 11m to grant Granted Apr 21, 2026
Patent 12579979
NAMING DEVICES VIA VOICE COMMANDS
1y 11m to grant Granted Mar 17, 2026
Patent 12537007
METHOD FOR DETECTING AIRCRAFT AIR CONFLICT BASED ON SEMANTIC PARSING OF CONTROL SPEECH
1y 0m to grant Granted Jan 27, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
96%
With Interview (+19.3%)
2y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 314 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month