DETAILED ACTION
Claims 1, 2, 4 – 6, 8, 11, 12, 14 – 16 and 18 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
With regard to the Final Office Action from 04 February 2026, the Applicant has filed a response on 24 April 2026.
Claims 3, 7, 9, 10, 13, 17, 19 and 20 have been cancelled.
With regard to the 35 U.S.C. 112(f) interpretation given to limitations of claim 1, the Applicant has now amended this claim to indicate that ‘the sound acquisition device’ is a ‘microphone array or a multi-channel microphone’ and this is suitable to overcome the claim interpretation.
The previous claims were rejected under 35 U.S.C. 101 for being directed to a judicial exception without significantly more. The independent claims have now been amended to provide, particularly by their last limitation, obtaining a final speech recognition result based on a first and a second speech recognition results that get dynamic weighting applied to them by assigning a weight according to the sound quality of each of the two speech recognition results a sound pickup distance of the node devices. This limitation adds significantly more to the mere mental process of performing speech recognition, presenting a particular and unique approach to the idea of obtaining speech recognition results. The Examiner hereby, based on this, withdraws the 35 U.S.C. 101 rejection.
Response to Arguments
The Applicant has amended claim 1 in an attempt to overcome the 35 U.S.C. 112(f) interpretation, further providing (Remarks: pages 9 – 10) that the amendment is suitable to obviate this interpretation, stating that sufficient structure and algorithmic detail is provided which fully satisfy the 35 U.S.C. 112(f). The Examiner however holds that the amendment only addresses the ‘speech acquisition device’ but is yet to properly address the ‘sound processing module’ and the ‘communication module’ other than maintaining that they are covered under a node device in a network. The Examiner maintains the 112(f) interpretation in this regard.
With regard to the 35 U.S.C. 103 rejection given to the claims, the Applicant has amended the claims, and the Applicant’s arguments on the applied prior art of reference are directed to the amendments to the independent claims, stating that the applied prior art do not teach the independent claims as currently amended. Applicant’s arguments with respect to the independent claims have been considered but are moot due to the new grounds of rejection necessitated by the amendment to the claims. The claims will be addressed by their current presentation in the following sections.
Claim Objections
Claims 4, 5, 6, 8, 14, 15, 16 and 18 are objected to because of the following informalities:
These claims are presented as being dependent on claims that have been cancelled. Claims 4, 5, 6 and 8 are presented as being dependent on claim 3 which is currently cancelled, while claims 14, 15, 16 and 18 are presented as being dependent on claim 13, also currently cancelled. For the purpose of this Office Action, the Examiner will consider claims 4, 5, 6 and 8 as being dependent on claim 1, while considering claims 14, 15, 16 and 18 as being dependent on claim 11. The dependency of these claims should be amended.
Appropriate correction is required.
Claim Interpretations
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
‘the sound processing module is configured to preprocess the electrical signal …’ in claim 1; and
‘the communication module is configured to send …’ in claim 1.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover a processor and communication modules such as those capable of WiFi/BLE/Zigbee communication protocols (page 8 lines 6 – 17) as the corresponding structure described in the Specification as performing the claimed function, and equivalents thereof.
If Applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, Applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recites sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the Applicant regards as his invention.
Claims 1, 2, 4 – 6, 8, 11, 12, 14 – 16 and 18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the Applicant), regards as the invention.
Regarding claims 1 and 11, the phrase “each node device has peer-to-peer direct communication capability and can act as a router to forward signals” renders the claim indefinite because it is unclear as to if each node device actually is performing the function of a router, or if it simply can perform the function of a router but does not actually perform the function of a router. As it is currently presented, it only indicates the possibility of functioning as a router without the precise recitation of functioning as a router to forward signals. This recitation renders the claims indefinite. See MPEP § 2173.05(d).
Dependent claims 2, 4 – 6, 8, 12, 14 – 16 and 18 are also rejected under 35 U.S.C. 112(b) based on their dependence on their respectively rejected base claims, in the absence of addressing the rejection raised for their respective base claims.
Claims 1 and 11 recite the limitations ‘the local node device’ and ‘the remote node device’ in their final limitations. There is insufficient antecedent basis for these limitations in the claims.
Claims 2, 4 – 6, 8, 12, 14 – 16 and 18 are also hereby rejected based on their dependence on the rejected claims above.
Allowable Subject Matter
Claims 1 and 11 would be allowable if rewritten or amended to overcome the rejections under 35 U.S.C. 112(b), set forth in this Office action.
The following is a statement of reasons for the indication of allowable subject matter:
With regard to independent claim 1, the invention states:
A distributed speech processing system, comprising:
a plurality of node devices in a network, wherein the network is a decentralized self-organizing Mesh network, the plurality of node devices implement ad hoc networking through wired or wireless manner, each node device has peer-to-peer direct communication capability and can act as a router to forward signals: each of the plurality of node devices comprises a processor, a memory, a communication module, and a sound processing module, and at least one of the plurality of node devices comprises a sound acquisition module; the processor is configured to provide a clock to achieve time synchronization of preprocessed data among all node devices: the sound acquisition module is a microphone array or a multi-channel microphone; wherein,
the sound acquisition module is configured to acquire an audio signal and convert the audio signal into an electrical signal;
the sound processing module is configured to preprocess the electrical signal through signal framing, pre-emphasis, and Fast Fourier Transform (FFT) to obtain a first sound preprocessed result;
the communication module is configured to send the first sound preprocessed result to one or more node devices in the network;
the communication module is further configured to receive one or more second sound preprocessed results from at least one other node device over the Mesh network; wherein the first sound preprocessed result and the second sound preprocessed result comprise data blocks with a same duration, and each of the data blocks include an incremental sequence number, a sound feature value, a sound quality, and sound time information;
the sound processing module is further configured to first filter out data blocks whose sound quality exceeds a predetermined threshold, then splice the first sound preprocessed result and the one or more second sound preprocessed results in the order of incremental sequence numbers, and select the data block with the highest sound quality from the data blocks with the same incremental sequence number to form a complete third sound preprocessed result perform speech recognition based on the first sound preprocessed result and the one or more second sound preprocessed results third sound preprocessed result to obtain a first speech recognition result; wherein each of the first sound preprocessed result and the one or more second sound preprocessed results is an intermediate result of the speech recognition;
the communication module is further configured to receive one or more second speech recognition results from at least one other node device over the Mesh network; and
the sound processing module is further configured to perform speech recognition based on the first speech recognition result and the one or more second speech recognition results to obtain a final speech recognition result, wherein the speech recognition adopts a dynamic weighting rule, wherein a weight is assigned according to the sound quality of the recognition results and the sound pickup distance of the node devices, and a weight of the recognition result of the local node device is higher than that of the remote node device.
Closest Prior Art
The reference of Nakadai et al. (US 2016/0055850 A1) provides teaching for first and second speech processing devices connected in a network [0046], the devices having a preprocessing unit as a sound processing module, a database as a memory, a communication unit, containing a sound acquisition module (FIG. 1), a processor [0264], acquiring the input speech signal (FIG. 1 Parts 30, 110), a preprocessing module for performing sound source localisation (FIG. 1 Part 112), a communication unit able to transmit pre-processed speech signal to another device 20A which can perform its own further preprocessing, such as sound source localisation, etc. (FIG. 7 Part 120), the network being a local area network (indicating according to FIG. 7 Part 50, that the devices are all in a network) [0053], communicating results of pre-processed data from one preprocessing step to the next, such as from 112 to 113 to 114 and first and second speech processing units for performing speech recognition on results of first and second sound pre-processed results, with Parts 112, 113 and 114 being preprocessed results that are intermediate results on a path to speech recognition Part 116 (FIG. 7); a first speech processing device 10 and a second speech processing device 20 which are connected through a network (indicating different node devices) [0046], the second speech processing device 20 also performs speech recognition on speech received from the first speech processing device) [0050], a communication unit that receives second speech recognition results over a network 50, whereby data can be transmitted and received between both the first and second speech processing devices, these being two devices qualifying as node devices over a network (FIG 7 Part 220), a second speech processing device 20A which includes preprocessing unit and a second speech recognition unit (able to perform its own speech recognition) [0135], and a teaching that ‘The communication unit 220 transmits transmission data including the second text data input from the second speech recognition unit 216 to the first speech processing device 10’ (showing a transmission of second speech recognition result from another node device) [0078].
The reference of KIM et al. (US 2019/038614 A1) provides teaching for generating a final speech recognition result based on weighting each of the available speech recognition results [0221] and obtaining several speech recognition results, which get combined in order to obtain an aggregate speech recognition result as a final speech recognition result [0222].
KIKUCHI et al. (JP 2020/0160281 A; applying the attached English translation) teaches performing pre-processing at 112-1 of a first device and then transmitting it over to another pre-processing unit 112-2 (indicating a distributed pre-processing over different devices or nodes) (page 4 line 47 – page 5 line 6) and another pre-processing unit 112-2 which obtains a signal that was initially pre-processed by a different unit, before performing its own pre-processing (page 6 lines 20–26).
Sereshki et al. (US 2020/010529 A1) provides teaching for the presence of microphone devices all arranged within a network [0105], whereby the different microphone devices in the network capture audio signals which could be misaligned, and these audio signals get aligned through synchronization as provided by a system clock [0107], whereby the devices within the network can transmit data/signals to other devices within the network, indicating that the devices can act with peer-to-peer communication capabilities, forwarding signals amongst themselves [0046].
Jeong et al. (US 2006/0053009 A1) provides teaching for passing a signal through a pre-emphasis filter as well as converting each frame into frequency domain using FFT [0093].
KANDA (US 2019/0139540 A1) provides teaching for a framing unit for framing digitized speech signal from an analogue-to-digital converter by using windows to place the signal in prescribed lengths [0068].
Huang et al. (US 2019/0295542 A1) provides teaching for synchronizing collected audio streams, generating a weighted combination of the different signals (indicating assigning weights to each of the signals), and then performing speech recognition on the weighted combination of the signals [0091].
Georganti (US 2020/0301651 A1) provides teaching for the selection of a better signal of microphone from two microphones (assigning weights to select that with a better quality) [0043], such that a microphone that is closer to the sound source gets selected (weighted higher) than one that is further away [0012], thereby teaching of a higher weight being applied to a closer/local device than a remote/farther-away device.
Brown et al. (US 2016/0364681 A1) provides teaching for splicing together, the best quality audio segments together [0089].
Lanham et al. (US 2010/0299131 A1) provides teaching for splicing together an original audio recording with a supplemental audio recording to obtain a modified time-aligned transcript [0056].
Parada et al. (US 2018/0075860 A1) provides teaching for selecting a best microphone pair based on confidence measures [0006], segmenting audio signals based on clustering information [0043], the selection of the microphones being to achieve a robust ASR performance.
Degani et al. (US 2009/0313018 A1) provides teaching for preprocessing speech utterances into sequences of strings of equal length blocks [0042].
Nemala et al. (US 2016/0063997 A1) provides teaching for dynamically assigning weights to each audio stream from each audio device as a user walks around a house, so as to ensure optimal audio quality and speech recognition at all times [0031].
The prior art of record taken alone or in combination however fail to teach, inter alia, a distributed processing system having node devices that acquire an audio signal such that first and second sound preprocessed results that comprise data blocks are acquired, each of the data blocks including an incremental sequence number, a sound feature value, a sound quality and sound time information, leading to the selection of data blocks with the highest sound quality to form a complete third sound preprocessed result that gets applied to obtain a first speech recognition result.
Claim 1 would hereby be allowable if rewritten or amended to overcome the 35 U.S.C. 112(b) rejection.
Claim 11 would also hereby be allowable if rewritten or amended to overcome the 35 U.S.C. 112(b) rejection, based on the reason set forth for claim 1 above.
Claims 2 and 12 would be allowable if rewritten to overcome the rejections under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
Claims 4, 5, 6, 8, 14, 15, 16 and 18 would be allowable if rewritten to overcome the rejections under 35 U.S.C. 112(b) and the claim objections set forth in this Office action, and to include all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to Applicant’s disclosure.
See the closest prior art provided in the “Allowable Subject Matter” section.
Any inquiry concerning this communication or earlier communications from the Examiner should be directed to OLUWADAMILOLA M. OGUNBIYI whose telephone number is (571)272-4708. The Examiner can normally be reached Monday – Thursday (8:00 AM – 5:30 PM Eastern Standard Time).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s Supervisor, PARAS D. SHAH can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/OLUWADAMILOLA M OGUNBIYI/Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
06/24/2026