Prosecution Insights
Last updated: August 18, 2026
Application No. 19/326,569

DISTINGUISHING USER SPEECH FROM BACKGROUND SPEECH IN SPEECH-DENSE ENVIRONMENTS

Non-Final OA §103
Filed
Sep 11, 2025
Priority
Jul 27, 2016 — continuation of 10/714,121 +5 more
Examiner
YANG, QIAN
Art Unit
2677
Tech Center
2600 — Communications
Assignee
Vocollect Inc.
OA Round
2 (Non-Final)
74%
Grant Probability
Favorable
2-3
OA Rounds
1y 9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
726 granted / 987 resolved
+11.6% vs TC avg
Strong +31% interview lift
Without
With
+31.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
26 currently pending
Career history
1004
Total Applications
across all art units

Statute-Specific Performance

§101
15.8%
-24.2% vs TC avg
§103
52.4%
+12.4% vs TC avg
§102
19.4%
-20.6% vs TC avg
§112
9.3%
-30.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 987 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Applicant's amendment filed on May 20, 2026 has been entered. Claims 1, 7 and 13 – 18 have been amended. No claims have been canceled. Claim 19 has been added. Claims 1 – 19 are still pending in this application, with claims 1, 7 and 13 being independent. Response to Arguments Applicant’s arguments with respect to claim(s) 1 – 19 have been considered but are moot because the new ground of rejection. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1 – 2, 7 – 8, 13 – 14 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Braho et al. (US Patent Application Publication 2014/0278391), hereinafter referred as Braho, in view of Commons (US Patent 8,775,341), and in further view of Chan et al. (“A hybrid noise suppression filter for accuracy enhancement of commercial speech recognizers in varying noisy conditions”, Applied Soft Computing, Volume 14, Part A, January 2014, Pages 132 – 139), hereinafter referred as Chan. Regarding claim 7, Braho discloses a speech recognition device (Figs. 1 – 3) comprising: a microphone (Fig. 1, #120 a-b); at least one processor (Fig. 3, microprocessor 302); and at least one memory (Fig. 3, ROM/RAM 306 and 308) including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the speech recognition device (SRD) to at least: receive, at the microphone, an audio input (Fig. 5, 504, [0110]), wherein the audio input comprises at least one of a speech of at least one user and a background sound in an environment of the at least one user ([0013, 0050, 0081, 0114 - 0118]), and wherein at least a portion of audio input corresponds to a task being executed by the at least one user ([0005 - 0007, 0013]); determine a presence of a background sound in the audio input based on an algorithm, wherein the algorithm is processed based on at least one of a plurality of audio speech samples and a plurality of background sound samples (Fig. 5, step 514 – 518, [0114 - 0118]); in an instance in which the background sound is absent from the audio input, generate at least one of words and phrases related to the task in a workflow (Fig. 5, 518, [0119], “This may advantageously limit the information being sent to be information which has been classified (i.e., determined to likely be) as speech rather than noise”; Fig. 5, 520 - 522, [0120 - 0122], “digitized audio to recognize speech”; [0162], “outputs recognized text”); and in an instance in which the audio input comprises the background sound, filter out the background sound from the audio input (Fig. 8, step 810 – 813, [0147 – 0150]; Fig. 9, step 912 – 916, [0161 – 0612], reject/block background sound). However, Braho fails to explicitly disclose the device wherein determining a presence of a background sound is based on a neural network, wherein the neural network is trained based on at least one of a plurality of audio speech samples and a plurality of background sound samples collected in a training environment having sound associated with an industrial setting. However, in a similar field of endeavor Commons discloses a system of using neural network to detect voice message (col. 45, lines 19 - 37). In addition, Commons discloses the system wherein determining a presence of a background sound is based on a neural network, wherein the neural network is trained based on at least one of a plurality of audio speech samples and a plurality of background sound samples (col. 45, lines 19 - 37) collected in a training environment having sound associated with a setting (col. 45, lines 19 – 37, “Parts of the training set could include noise, either from other speakers, musical instruments, or white noise in the background”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Braho, and determining a presence of a background sound is based on a neural network, wherein the neural network is trained based on at least one of a plurality of audio speech samples and a plurality of background sound samples collected in a training environment having sound associated with a setting. The motivation for doing this is that the device of Braho can be more powerful and advanced for artificial intelligence. However, Commons fails to explicitly disclose wherein the sound associated with an industrial setting. However, in a similar field of endeavor Chan discloses a method for accuracy enhancement of commercial speech recognizers in varying noisy conditions (abstract). In addition, Chan discloses the method wherein the environment of sound associated with an industrial setting (abstract, “under various noisy environments in factories”). There was some teaching, suggestion, or motivation, either in the references themselves or in the knowledge generally available to one of ordinary skill in the art, to modify Commons and Chan or to combine Commons and Chan teachings, and collected in a training environment having sound associated with an industrial setting. There was reasonable expectation of success to achieve collected in a training environment having sound associated with an industrial setting by modifying Commons and Chan or to combining Commons and Chan teachings (KSR scenario G). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Braho and Commons, and collected in a training environment having sound associated with an industrial setting. The motivation for doing this is that the training data can be more close to real environment so that the speech recognition can be more accurate. Regarding claim 8 (depends on claim 7), Braho discloses the device wherein the at least one processor is configured to receive the plurality of audio speech samples and the plurality of background sound samples, wherein the plurality of audio speech samples corresponds to a speech of a plurality of users ([0007], “a plurality of users each wearing respective portable computer systems and headsets interface with the central or server computer system. This approach allows the user(s) to provide spoken or voice input to the voice driven system, including commands and/or information”; also [0110]), and wherein the plurality of background sound samples corresponds to the background sound in the environment of the plurality of users ([0007], “a plurality of users each wearing respective portable computer systems and headsets interface with the central or server computer system”; [0008], “conversations which are not intended as input”). Regarding claims 1 – 2, they are corresponding to claims 7 – 8, respectively, thus, they are interpreted and rejected for a same reason set forth for claims 7 – 8. Regarding claims 13 – 14, they are corresponding to claims 7 – 8, respectively, thus, they are interpreted and rejected for a same reason set forth for claims 7 – 8. Regarding claim 19 (depends on claim 1), Chan discloses the method wherein the plurality of background sound samples comprises industrial sound samples (abstract, “under various noisy environments in factories”). Claim(s) 3, 9 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Braho in view of Commons, in further view of Chan and Faisman et al. (US Patent Application Publication 2008/0319743), hereinafter referred as Faisman. Regarding claim 9 (depends on claim 7), Braho fails to explicitly disclose the device wherein the at least one processor is configured to: generate a first transcript of each speech sample of the plurality of audio speech samples; and generate a second transcript of a set of background sound samples of the plurality of background sound samples, wherein the set of background sound samples include speech of one or more users in the environment. However, in a similar field of endeavor Faisman discloses an ASR-aided transcription system (abstract). In addition, Faisman discloses the system wherein generate transcripts of each speech sample of the plurality of audio speech samples (Fig. 1, [0013 – 0018], generate transcripts for each speech sample of the plurality of audio speech samples). Braho discloses process each speech sample of the plurality of audio speech samples, and process a set of background sound samples of the plurality of background sound samples, wherein the set of background sound samples include speech of one or more users in the environment ([0007 – 0008]). There was some teaching, suggestion, or motivation, either in the references themselves or in the knowledge generally available to one of ordinary skill in the art, to modify Braho and Faisman, or to combine references teachings. There was reasonable expectation of success to modify Braho and Faisman, or to combine references teachings to achieve the claimed limitations. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Braho, and generate a first transcript of each speech sample of the plurality of audio speech samples; and generate a second transcript of a set of background sound samples of the plurality of background sound samples, wherein the set of background sound samples include speech of one or more users in the environment. The motivation for doing this is that the all conversation can be recorded and logged so that it is beneficial for a later checking. Regarding claims 3 and 15, they are corresponding to claim 9, thus, they are interpreted and rejected for a same reason set forth for claim 9. Claim(s) 4, 10 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Braho in view of Commons, in further view of Chan, Faisman and Sak et al. (“LEARNING ACOUSTIC FRAME LABELING FOR SPEECH RECOGNITION WITH RECURRENT NEURAL NETWORKS”, 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)), hereinafter referred as Sak. Regarding claim 10 (depends on claim 9), Braho fails to explicitly disclose the device wherein the at least one processor is configured to train the neural network based at least on the first transcript and the second transcript. However, in a similar field of endeavor Sak discloses a method for acoustic modeling (abstract). In addition, Sak discloses the method wherein train the neural network based at least on transcripts (section 1, 2.2, 2.3, 4.1, transcript). There was some teaching, suggestion, or motivation, either in the references themselves or in the knowledge generally available to one of ordinary skill in the art, to modify Braho and Sak, or to combine references teachings. There was reasonable expectation of success to modify Braho and Sak, or to combine references teachings to achieve the claimed limitations. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Braho, and train the neural network based at least on the first transcript and the second transcript. The motivation for doing this is that training can be more precise and effective with labeled transcripts. Regarding claims 4 and 16, they are corresponding to claim 10, thus, they are interpreted and rejected for a same reason set forth for claim 10. Claim(s) 5 – 6, 11 – 12 and 17 – 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Braho in view of Commons, in further view of Chan, Faisman and Ramalho et al. (US Patent 8,600,750), hereinafter referred as Ramalho. Regarding claim 11 (depends on claim 9), Braho fails to explicitly disclose the device wherein the at least one processor is configured to determine sound characterization for one or more words based on the first transcript and the second transcript. However, in a similar field of endeavor Ramalho discloses an ASR system (abstract). In addition, Ramalho discloses the system wherein determine sound characterization for one or more words based on the first transcript and the second transcript (col. 1, line 47 to col. 2, line 3; col. 4, lines 29 - 43). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Braho, and determine sound characterization for one or more words based on the first transcript and the second transcript. The motivation for doing this is that the application of Braho can be extended to specify a specific user so that the system is more advanced. Regarding claim 12 (depends on claim 11), Braho discloses the device wherein the at least one processor is configured to determine a rejection threshold based on the sound classifier, wherein the rejection threshold is utilized to at least accept or reject the audio input ([0143 – 0150]). However, Braho fails to explicitly disclose wherein the sound classifier is the sound characterization. However, in a similar field of endeavor Ramalho discloses an ASR system (abstract). In addition, Ramalho discloses the system wherein classifies the sound characterization to group different people (col. 2, line 52 to col. 3, line 20). Substituting sound classification with sound characterization were known to the art. One of ordinary skill in the art could have substituted one known element for another, and the results of the substitution would have been predictable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Braho, and determine a rejection threshold based on the sound characterization. The motivation for doing this is that the application of Braho can be extended to specify a specific user so that the system is more advanced. Regarding claims 5 – 6, they are corresponding to claims 11 – 12, respectively, thus, they are interpreted and rejected for a same reason set forth for claims 11 – 12. Regarding claims 17 – 18, they are corresponding to claims 11 – 12, respectively, thus, they are interpreted and rejected for a same reason set forth for claims 11 – 12. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to QIAN YANG whose telephone number is (571)270-7239. The examiner can normally be reached on Monday-Thursday 8am-6pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached on 571-270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /QIAN YANG/ Primary Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Sep 11, 2025
Application Filed
Feb 20, 2026
Non-Final Rejection mailed — §103
May 20, 2026
Response Filed
Jun 04, 2026
Final Rejection mailed — §103
Jul 28, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699716
REDUCTION OF LATENCY IN RETRIEVER-READER ARCHITECTURES
3y 10m to grant Granted Aug 04, 2026
Patent 12694511
SYSTEM FOR PROCESSING A WHOLE SLIDE IMAGE, WSI, OF A BIOPSY
3y 11m to grant Granted Jul 28, 2026
Patent 12682172
Intelligent Interface for Automation Multitasking
2y 7m to grant Granted Jul 14, 2026
Patent 12684080
IMAGE FORMING APPARATUS
2y 3m to grant Granted Jul 14, 2026
Patent 12675874
Graph-based Hemodynamics for Biomarkers of Neurovascular Resilience
4y 3m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+31.2%)
2y 8m (~1y 9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 987 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month