DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
The Supreme Court has long held that “[l]aws of nature, natural phenomena, and abstract ideas are not patentable.” Alice Corp. Pty. Ltd. v. CLS Bank Int’l, 134 S. Ct. 2347, 2354 (2014) (quoting Assoc. for Molecular Pathology v. Myriad Genetics, Inc., 133 S. Ct. 2107, 2116 (2013) (internal quotation marks omitted)). The “abstract ideas” category embodies the longstanding rule that an idea, by itself, is not patentable. Alice Corp., 134S. Ct. at 2355 (quoting Gottschalk v. Benson, 409 U.S. 63, 67 (1972).
In Alice, the Supreme Court sets forth an analytical “framework for distinguishing patents that claim laws of nature, natural phenomena, and abstract ideas [or mental processes ] from those that claim patent-eligible applications of those concepts.” Id. at 2355 (citing Mayo Collaborative Servs. v. Prometheus Labs., Inc., 132 S. Ct. 1289, 1296–97 (2012)). The first step in the analysis is to “determine whether the claims at issue are directed to one of those patent-ineligible concepts.” Id. If the claims are directed to a patent-ineligible concept, the second step in the analysis is to consider the elements of the claims “individually and ‘as an ordered combination’” to determine whether there are additional elements that “‘transform the nature of the claim’ into a patent-eligible application.” Id. (quoting Mayo, 132 S. Ct. at 1298, 1297). In other words, the second step is to “search for an ‘inventive concept’—i.e., an element or combination of elements that is ‘sufficient to ensure that the patent in practice amounts to significantly more than a patent upon the [ineligible concept] itself’”. Id. (brackets in original) (quoting Mayo, 132 S. Ct. at 1294). The prohibition against patenting an abstract idea “‘cannot be circumvented by attempting to limit the use of the formula to a particular technological environment’ or adding ‘insignificant post-solution activity.’” Bilski v. Kappos, 561 U.S. 593, 610–11 (2010) (citation omitted).
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any statutory category. See MPEP 2106.03. Independent Claim 1 recites the method of receiving a sequence of acoustic frames characterizing portions of an utterance, employing a speech recognition model to detect a disfluency in the utterance and determine an end of speech when the user has finished speaking, and thus is a process (a series of steps or acts). A process is a statutory category of invention. Additionally, Independent Claim 11 recites a system comprising data processing hardware and memory hardware configured to execute operations similar to Claim 1. A system is a Statutory category of invention.
Step 2A, Prong One: This part of the eligibility analysis evaluates whether the claim recites a judicial exception. As explained in MPEP 2106.04, subsection II, a claim “recites” a judicial exception when the judicial exception is “set forth” or “described” in the claim. In applying the framework set out in Alice, examiner found Applicant’s claims 1 and 11 are directed to a patent-ineligible abstract concept of detecting disfluency in an utterance and an endpoint corresponding to an end of speech of the received utterance. The steps of Applicant’s claims 1 and 11 are an abstract concept that would fall under the judicial exception of mental processes. Specifically, the claims recite the step of “as a user speaks an utterance… receiving a first sequence of acoustic frames characterizing a first portion of the utterance spoken by the user.” This limitation is directed to mental process because a human is able to receive data either audibly by listening to audible speech, or visually on a sheet of paper. Furthermore, the step of “processing… the first sequence of acoustic frames to generate first speech recognition results for the first portion of the utterance.” The claim does not place any limits on how the first speech recognition result is generated. The “processing” may simply involve a human interpreting received utterances. Therefore, this step is directed to a mental process. Furthermore, the step of “after receiving the first sequence of acoustic frames, receiving a second sequence of acoustic frames characterizing a second portion of the utterance spoken by the user” recites steps that are directed to a mental process. The claim purports to use a particular speech processing model on multiple segments of a received utterance by the use of “acoustic frames characterizing a second portion”, however, the portions as claimed may correspond to a human receiving multiple parts of an utterance and interpreting them as they are received. Therefore, the above steps are also directed to mental processes. Furthermore, the claim recites the step of “processing… the second sequence of acoustic frames to detect a presence of a disfluency in the second portion of the utterance.” The claim does not specify what kind of processing is involved in detecting a disfluency in the utterance. Under broadest reasonable interpretation, a human is capable of detecting a disfluency using his or her own judgement. Furthermore, the claim recites the step of “based on detecting the presence of the disfluency in the second portion of the utterance, determining that the user has not finished speaking the utterance.” Again, the claim does not specify what kind of processing is involved in detecting the disfluency in the utterance, thus, under broadest reasonable interpretation, a human is capable of detecting the disfluency and verifying whether the disfluency corresponds to an end of speech using his or her own judgement. Furthermore, the claim recites the step of “after receiving the second sequence of acoustic frames, receiving a third sequence of acoustic frames characterizing a third portion of the utterance spoken by the user.” As noted above, the limitation corresponds to a processing of an utterance in parts, which is something a human is capable of performing in the mind. Furthermore, the claim recites the step of “processing… the third sequence of acoustic frames to: generate second speech recognition results for the third portion of the utterance; and detect an end of speech event at an end of the third portion of the utterance.” As noted above, no limit is placed on how the end of speech is determined, thus, a human is capable of determining an end of speech using his or her own judgement. Finally, the step of “based on detecting the end of speech event, generating a response to the utterance” recites steps that are directed to a mental process. The limitation provides no further limit on how the response is generated, therefore, as noted above, a human is capable of interpreting and detecting an utterance endpoint and provide a response to the received utterance once the utterance has determined to be finished. Under broadest reasonable interpretation, the limitations correspond to normal human-to-human conversation. Additionally, claim 11 contains subject matter different from claim 1 in the last limitation reciting “based on detecting the end of speech event, triggering a microphone closing event by the user device.” Although a particular technological environment is recited such as a microphone closing event, the closing of a microphone may be represented by an environment where a human operator may “mute” another person’s microphone once they have determined to have finished speaking, such as in television or radio broadcast environments, or public seminars involving multiple people speaking using multiple microphones. Furthermore, dependent claims 2-10 and 12-20 simply recite further description of the detected disfluency being a filler word or detecting a pause in the utterance, displaying a transcription of the user utterance, providing synthesized output corresponding to the response and generating an acknowledgement of the detected disfluency. The recited elements correspond to mental processes.
Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a whole integrates the recited judicial exception into a practical application of the exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additional elements individually and in combination to determine whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d).
Independent claims 1 and 11 recite “…a digital assistant application…” and “…a speech recognition model” as additional elements beyond the judicial exception. The examiner has found, however, that the use of a “digital assistant application” and “speech recognition model” step provides no further detail and is recited at such a high-level of generality that this limitation is merely a post-solution step. Therefore, this step is an insignificant extra-solution activity and does not integrate the judicial exception into a practical application. See MPEP 2106.05(g). Independent Claim 11 further recites “data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations” and “a microphone” as additional elements beyond the judicial exception. However, these additional elements do not amount to significantly more than the abstract idea because the additional elements constitute a generic computer environment. Alice, 134 S. Ct. at 2357. The Claims need meaningful limitations that go beyond generally linking the use of an abstract idea to a particular technological environment. Therefore, the steps are all abstract and the Claim as a whole is abstract. “[S]imply appending generic computer functionality to lend speed or efficiency to the performance of an otherwise abstract concept does not meaningfully limit claim scope for purposes of patent eligibility.” CLS Bank, 2013 U.S. App. LEXIS 9493, at *29 (citing Bancorp, 687 F.3d at 1278, and Dealertrack, Inc. v. Huber, 674 F.3d 1315, 1333-34 (Fed. Cir. 2012) (finding that the claimed computer-aided clearinghouse process is a patent-ineligible abstract idea)); SiRF Tech., Inc. v. Int'l Trade Comm'n, 601 F.3d 1319, 1333 (Fed. Cir. 2010) (“In order for the addition of a machine to impose a meaningful limit on the scope of a claim, it must play a significant part in permitting the claimed method to be performed, rather than function solely as an obvious mechanism for permitting a solution to be achieved more quickly, i.e., through the utilization of a computer for performing calculations.”). Additionally, dependent claims 2-10 and 12-20 do not integrate the judicial exception into a practical application.
Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole amounts to significantly more than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. See MPEP 2106.05.
At step 2A, prong two, the additional elements of “…a digital assistant application…” “…a speech recognition model” “data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations” and “a microphone” were found to be extra-solution activity and a generic computer environment. At Step 2B, the re-evaluation of the insignificant extra-solution activity consideration takes into account whether or not the extra-solution activity is well understood, routine, and conventional in the field. See MPEP 2106.05(g). Here, the extra-solution activity and generic computer environment are recited at such a high level of generality that it does not provide an inventive concept. Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, and therefore do not provide an inventive concept. Additionally, dependent claims 2-10 and 12-20 do not add an inventive concept.
In conclusion, Examiner notes that none of recited steps in Applicant's claims 1-20 refer to a specific machine by reciting structural limitations of any apparatus or to any specific operations that would cause a machine to be the mechanism to perform these steps. Although the claims may be processed by a computing system having a processor, the computing system is merely a general purpose computing system. Therefore, all of the claims 1-20 are abstract.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 5-7 and 9 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Bratt (US PG Pub 20220115001).
As per claim 1, Bratt discloses: A computer-implemented method executing on data processing hardware that causes the data processing hardware to perform operations comprising: as a user speaks an utterance directed toward a digital assistant application and captured by a microphone of a user device (Bratt; Fig. 1, item 102; p. 0036 - …The Automatic Audio Processing module 102 receives speech input from the user via one or more microphones. The links and/or interfaces exchange information with the Automatic Audio Processing Module 102 to detect and comprehend the user's audio input from the one or more microphones; see also p. 0164-0165 - The CI manager module 106 cooperates with the Automated Speech Recognition module 102 and the Spoken Language Understanding module 104 to enable functions to distinguish system/non-system directed speech…; see also p. 0177 - The input of a speech signal to the VDA is captured as a speech waveform associated with utterances spoken by user. The speech data processing subsystem produces speech data corresponding to audio input captured from a human in the speech waveforms…): receiving a first sequence of acoustic frames characterizing a first portion of the utterance spoken by the user (Bratt; p. 0177 - The acoustic front end includes a plurality of analytics engines each comprising a plurality of algorithms are configured different types of user state analytics including the timing, the pitch and the loudness of phones and phrases to convey prosody over frames of the speech waveform); processing, using a speech recognition model, the first sequence of acoustic frames to generate first speech recognition results for the first portion of the utterance (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…); after receiving the first sequence of acoustic frames, receiving a second sequence of acoustic frames characterizing a second portion of the utterance spoken by the user (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…; see also p. 0183 - …The CI manage module 106 cooperating with the spoken language understanding module 104 looks for grammatical completeness of sentence syntax in the flow of speech coming from the user. If a user initially responds “Yeah that looks good but . . . ” (first sequence), the CI manager module 106 is configured to understand that this in an incomplete human sentence. Then the user might subsequently state after the long pause “I am not sure on Tuesday, maybe Wednesday! (second sequence)” Thus, if the CI manager module 106 pairs this initial flow of speech with a subsequent flow of speech from the user, then possibly a grammatically complete sentence can be sent to the spoken language understanding module 104 to get a correct interpretation of the speech from the user and without taking the conversational floor from the user before they completely convey the concept in the flow of speech in which they were attempting to convey with those two broken up phrases…; see also p. 0197 - …speaking with pauses between two or more (three) user utterances (processing first sequence followed by a second sequence, followed by a third sequence) so that the user can respond initially incompletely with a first utterance followed by a pause and then a second utterance to complete a thought the user is trying to convey with that speech activity…); processing, using the speech recognition model, the second sequence of acoustic frames to detect a presence of a disfluency in the second portion of the utterance (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0038-0039 - The CI manager module 106 has a disfluency detector for a micro-interaction on an analysis of timing information on the flow of speech coming from the user… The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…; see also p. 0182-0183 – …The CI manager module 106 has an input from a disfluency detector to trigger a micro-interaction on speech repair to detect disfluency information of various breaks of i) words and sentences that are cut off mid-utterance, and/or ii) non-lexical vocables uttered while the user is speaking and holding the conversational floor…); based on detecting the presence of the disfluency in the second portion of the utterance, determining that the user has not finished speaking the utterance (Bratt; p. 0183 - …The spoken language understanding module 104 may indicate when a current flow of speech does not contain a completed thought…; see also p. 0076-0089 - Micro-Interaction: User Has Not Completed Their Utterance/Thought); after receiving the second sequence of acoustic frames, receiving a third sequence of acoustic frames characterizing a third portion of the utterance spoken by the user (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample (sequence) of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…; see also p. 0183 - …The CI manage module 106 cooperating with the spoken language understanding module 104 looks for grammatical completeness of sentence syntax in the flow of speech coming from the user. If a user initially responds “Yeah that looks good but . . . ” (first sequence), the CI manager module 106 is configured to understand that this in an incomplete human sentence. Then the user might subsequently state after the long pause “I am not sure on Tuesday, maybe Wednesday! (second sequence)” Thus, if the CI manager module 106 pairs this initial flow of speech with a subsequent flow of speech from the user, then possibly a grammatically complete sentence can be sent to the spoken language understanding module 104 to get a correct interpretation of the speech from the user and without taking the conversational floor from the user before they completely convey the concept in the flow of speech in which they were attempting to convey with those two broken up phrases…; see also p. 0197 - …speaking with pauses between two or more user utterances (processing first sequence followed by a second sequence, followed by a third sequence) so that the user can respond initially incompletely with a first utterance followed by a pause and then a second utterance to complete a thought the user is trying to convey with that speech activity…); and processing, using the speech recognition model, the third sequence of acoustic frames to: generate second speech recognition results for the third portion of the utterance (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample (sequence) of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…); and detect an end of speech event at an end of the third portion of the utterance (Bratt; p. 0040 - The CI manager module 106 uses the input from the prosodic detector to determine i) whether the user has indeed yielded the conversational floor (end of speech) or ii) whether the user is inserting pauses into a flow of their speech to convey additional information. Note, the additional information can include 1) speaking with pauses to help convey and understand a long list of information, 2) speaking with pauses between two or more user utterances such that the user responds initially incompletely with a first utterance followed by a pause and then a later utterance to complete a thought the user is trying to convey with that speech activity, as well as 3) any combination of these two; see also p. 0177 – Each of these analytics engines can have executable software using algorithms specifically for performing that particular function. For example, the pause analytics engine can utilize a conventional “speech/no-speech” algorithm that detects when a pause in the speech occurs. The output is a binary value that indicates whether the present speech signal sample is a portion of speech or not a portion of speech. This output and determination information can be used to identify an endpoint; see also p. 0182-0183); and based on detecting the end of speech event, generating a response to the utterance (Bratt; p. 0056 - The dialogue management module 108 uses rules, codified through the dialogue specification language (or again alternatively implemented with a decision tree and/or trained artificial intelligence model), to detect for when a topic shift initiated by the user occurs, as well as, when the conversational assistant should try a topic shift, and then generates an adapted user-state aware response(s) based on the conversational context; see also p. 0048 - The links and/or interfaces exchange information with i) the TTS module 112 to generate audio output from a text format or waveform format, as well as ii) work with a natural language generation module 110 to generate audio responses and queries from the CI manager module 106. The TTS module 112 uses one or more speakers to generate the audio output for the user to hear).
As per claim 2, Bratt discloses: The computer-implemented method of claim 1, wherein detecting the presence of the disfluency in the second portion of the utterance comprises detecting a presence of a filler word spoken by the user in the second portion of the utterance (Bratt; p. 0067 - The CI manager module 106 uses a rule-based engine on conversational intelligence to understand and generate beyond-the-words conversational cues to establish trust while smoothly navigating complex conversations, such as i) non-lexical vocal cues, such as an “Uhmm” utterance (filler words), and ii) pitch, such as “60;Right!!” or “Right??”…; see also p. 0042-0043).
As per claim 3, Bratt discloses:
The computer-implemented method of claim 2, wherein processing the second sequence of acoustic frames to detect the presence of the disfluency further comprises processing, using the speech recognition model, the second sequence of acoustic frames to generate third speech recognition results for the second portion of the utterance (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0038-0039 - The CI manager module 106 has a disfluency detector for a micro-interaction on an analysis of timing information on the flow of speech coming from the user… The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…; see also p. 0182-0183 – …The CI manager module 106 has an input from a disfluency detector to trigger a micro-interaction on speech repair to detect disfluency information of various breaks of i) words and sentences that are cut off mid-utterance, and/or ii) non-lexical vocables uttered while the user is speaking and holding the conversational floor…), the third speech recognition results comprising a transcript of the filler word spoken by the user (Bratt; p. 0067 - The CI manager module 106 uses a rule-based engine on conversational intelligence to understand and generate beyond-the-words conversational cues to establish trust while smoothly navigating complex conversations, such as i) non-lexical vocal cues, such as an “Uhmm” utterance (filler words), and ii) pitch, such as “Right!!” or “Right??”…; see also p. 0042-0043).
As per claim 5, Bratt discloses:
The computer-implemented method of claim 1, wherein detecting the presence of the disfluency in the second portion of the utterance comprises detecting a presence of a pause in speech in the second portion of the utterance (Bratt; p. 0031 - The CI manager module 106 uses the rule-based engine to analyze and make determinations on factors of conversational cues. The rule-based engine has rules to analyze and make determinations on two or more conversational cues of i) non-lexical items, ii) pitch of spoken words, iii) prosody of spoken words, iv) grammatical completeness of sentence syntax in the user's flow of speech, and v) pause duration, vi) degree of semantic constraints of a user's utterance. Note, the pitch of spoken words can be a part of prosody. Also, a degree of semantic constraints of a user's utterance can be when a user is looking for a restaurant and then pauses a little, the system will just offer a ton of restaurant options; see also p. 0038 - The CI manager module 106 has a disfluency detector for a micro-interaction on an analysis of timing information on the flow of speech coming from the user. The timing information can be used for prosodic analysis. The timing information can also be used for a timer for determining time durations, such as a 0.75 second pause after receiving final word in a completed thought from the user, which indicates the user is yielding the conversation floor).
As per claim 6, Bratt discloses:
The computer-implemented method of claim 1, wherein the operations further comprise, based on detecting the end of speech event, processing the first speech recognition results and the second speech recognition results to execute a query specified by the utterance spoken by the user (Bratt; p. 0029 - the CI manager module 106 using the rule-based engine choice between a response of i) a full sentence and ii) a back channel may depend on the level of confidence of the conversational engagement 100 on understanding of the meaning behind what the user recently conveyed. Note, the full sentence response can be when the system determines that the user has given enough information (e.g. reservation for hotel in Rome near Trevi Fountain), the CI manager module 106 directs a look up for hotels meeting the criteria and simply responds with the information the user is looking for. The response of information the user is looking for implicitly conveys the confirmation of the conversational grounding for the topic and issue of the current dialog).
As per claim 7, Bratt discloses:
The computer-implemented method of claim 6, wherein the operations further comprise providing, for audible output from the user device, a synthesized speech representation of the response to the query (Bratt; p. 0189 - The VDA is able to interpret human speech including Human conversational cues that go beyond just the spoken words and respond via synthesized voices).
As per claim 9, Bratt discloses:
The computer-implemented method of claim 1, wherein the operations further comprise, based on detecting the presence of the disfluency in the second portion of the utterance, generating an acknowledgement response, the acknowledgement response indicating to the user that the digital assistant application is waiting for the user to finish speaking the utterance (Bratt; p. 0028 - the CI manager module 106 uses the rule-based engine to analyze and make an example determination on whether to issue a backchannel utterance, such as “Uh-mm” or “Okay”, to quickly indicate that the modules of the VDA understood both the words and the conveyed meaning behind the initial thought of “Find me a hotel in Rome by Trevi Fountain” by generating this short backchannel utterance while the user still holds the conversational floor and without the VDA attempting to take the floor. The flow of the speech and its conversational cues from the user indicate that the user intends to continue with conveying additional information after this initial thought so a short back channel acknowledgement is appropriate).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 4, 8 and 11-19 are rejected under 35 U.S.C. 103 as being unpatentable over Bratt in view of Simko (US PG Pub 20180350395).
As per claim 4, Bratt discloses: The computer-implemented method of claim 3, wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device (Bratt; p. 0212 - a display screen 491 to display at least some of the information stored in the one or more memories 430-432 and other components), a transcription of the utterance that comprises the first speech recognition results generated for the first portion of the utterance, the third speech recognition results generated for the second portion of the utterance, and the second speech recognition results generated for the third portion of the utterance. Bratt, however, fails to explicitly disclose providing, for display in a graphical user interface displayed on a screen of the user device a transcription of the utterance that comprises the first speech recognition results generated for the first portion of the utterance, the third speech recognition results generated for the second portion of the utterance, and the second speech recognition results generated for the third portion of the utterance. Simko does teach providing, for display in a graphical user interface displayed on a screen of the user device a transcription of the utterance that comprises the first speech recognition results generated for the first portion of the utterance, the third speech recognition results generated for the second portion of the utterance, and the second speech recognition results generated for the third portion of the utterance (Simko; p. 0043 - At stage I, the server 110 outputs the transcription 128 of the utterance 104. In some implementations, the server 110 may transmit the transcription 128 to the computing device 106. In this instance, the computing device 106 may display the transcription on the 128 on the display of the computing device 106). Therefore, it would have been obvious to one of ordinary skill In the art to modify the system of Bratt to include providing, for display in a graphical user interface displayed on a screen of the user device a transcription of the utterance that comprises the first speech recognition results generated for the first portion of the utterance, the third speech recognition results generated for the second portion of the utterance, and the second speech recognition results generated for the third portion of the utterance, as taught by Simko, because a user may use the voice input capabilities of a computing device and speak at a pace that is comfortable for the user. This may increase the utility of the computing device for users, in particular for users with speech disorders or impediments. An utterance may be endpointed at the intended end of the utterance, leading to more accurate or desirable natural language processing outputs, and to faster processing by the natural language processing system. This can reduce the use of computational resources and can conserve power. Moreover, closing of the microphone at a more suitable point can further reduce the use of computational resources and conserve power, since the microphone does not need to be activated, and the use of computational resources in interpreting and performing tasks based on additional audio detected by the microphone can be avoided (Simko; p. 0016).
As per claim 8, Bratt discloses:
The computer-implemented method of claim 6, upon which claim 8 depends. Bratt, however, fails to specifically disclose wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device: the first speech recognition results for the first portion of the utterance; the second speech recognition results for the second portion of the utterance; and the response to the query. Simko does teach wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device: the first speech recognition results for the first portion of the utterance; the second speech recognition results for the second portion of the utterance (Simko; p. 0043 - At stage I, the server 110 outputs the transcription 128 of the utterance 104. In some implementations, the server 110 may transmit the transcription 128 to the computing device 106. In this instance, the computing device 106 may display the transcription on the 128 on the display of the computing device 106; see also p. 0058); and the response to the query (Simko; p. 0058 - In this instance, the computing device 202 may display the transcription on the 228 on the display of the computing device 202. In some implementations, the server 210 may perform an action based on the transcription 228 such as initiate a phone call, send a message, open an application, initiate a search query, or any other similar action).
Therefore, it would have been obvious to one of ordinary skill In the art to modify the system of Bratt to include wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device: the first speech recognition results for the first portion of the utterance; the second speech recognition results for the second portion of the utterance; and the response to the query, as taught by Simko, because a user may use the voice input capabilities of a computing device and speak at a pace that is comfortable for the user. This may increase the utility of the computing device for users, in particular for users with speech disorders or impediments. An utterance may be endpointed at the intended end of the utterance, leading to more accurate or desirable natural language processing outputs, and to faster processing by the natural language processing system. This can reduce the use of computational resources and can conserve power. Moreover, closing of the microphone at a more suitable point can further reduce the use of computational resources and conserve power, since the microphone does not need to be activated, and the use of computational resources in interpreting and performing tasks based on additional audio detected by the microphone can be avoided (Simko; p. 0016).
As per claim 11, Bratt discloses: A system comprising:
data processing hardware (Bratt; Fig. 4, item 420; p. 0212 - The computing device may include one or more processors or processing units 420 to execute instructions); and memory hardware in communication with the data processing hardware (Bratt; Fig. 4, items 430-432; p. 0212 - one or more memories 430-432 to store information) and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations (Bratt; p. 0212 - …Note, portions of this design implemented in software 444, 445, 446 are stored in the one or more memories 430-432 and are executed by the one or more processors 420…) comprising: as a user speaks an utterance directed toward a digital assistant application and captured by a microphone of a user device (Bratt; Fig. 1, item 102; p. 0036 - …The Automatic Audio Processing module 102 receives speech input from the user via one or more microphones. The links and/or interfaces exchange information with the Automatic Audio Processing Module 102 to detect and comprehend the user's audio input from the one or more microphones; see also p. 0164-0165 - The CI manager module 106 cooperates with the Automated Speech Recognition module 102 and the Spoken Language Understanding module 104 to enable functions to distinguish system/non-system directed speech…; see also p. 0177 - The input of a speech signal to the VDA is captured as a speech waveform associated with utterances spoken by user. The speech data processing subsystem produces speech data corresponding to audio input captured from a human in the speech waveforms…): receiving a first sequence of acoustic frames characterizing a first portion of the utterance spoken by the user (Bratt; p. 0177 - The acoustic front end includes a plurality of analytics engines each comprising a plurality of algorithms are configured different types of user state analytics including the timing, the pitch and the loudness of phones and phrases to convey prosody over frames of the speech waveform); processing, using a speech recognition model, the first sequence of acoustic frames to generate first speech recognition results for the first portion of the utterance (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…); after receiving the first sequence of acoustic frames, receiving a second sequence of acoustic frames characterizing a second portion of the utterance spoken by the user (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection… see also p. 0183 - …The CI manage module 106 cooperating with the spoken language understanding module 104 looks for grammatical completeness of sentence syntax in the flow of speech coming from the user. If a user initially responds “Yeah that looks good but . . . ” (first sequence), the CI manager module 106 is configured to understand that this in an incomplete human sentence. Then the user might subsequently state after the long pause “I am not sure on Tuesday, maybe Wednesday! (second sequence)” Thus, if the CI manager module 106 pairs this initial flow of speech with a subsequent flow of speech from the user, then possibly a grammatically complete sentence can be sent to the spoken language understanding module 104 to get a correct interpretation of the speech from the user and without taking the conversational floor from the user before they completely convey the concept in the flow of speech in which they were attempting to convey with those two broken up phrases…; see also p. 0197 - …speaking with pauses between two or more user utterances (processing first sequence followed by a second sequence, followed by a third sequence) so that the user can respond initially incompletely with a first utterance followed by a pause and then a second utterance to complete a thought the user is trying to convey with that speech activity…); processing, using the speech recognition model, the second sequence of acoustic frames to detect a presence of a disfluency in the second portion of the utterance (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0038-0039 - The CI manager module 106 has a disfluency detector for a micro-interaction on an analysis of timing information on the flow of speech coming from the user… The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…; see also p. 0182-0183 – …The CI manager module 106 has an input from a disfluency detector to trigger a micro-interaction on speech repair to detect disfluency information of various breaks of i) words and sentences that are cut off mid-utterance, and/or ii) non-lexical vocables uttered while the user is speaking and holding the conversational floor…); based on detecting the presence of the disfluency in the second portion of the utterance, determining that the user has not finished speaking the utterance (Bratt; p. 0183 - …The spoken language understanding module 104 may indicate when a current flow of speech does not contain a completed thought…; see also p. 0076-0089 - Micro-Interaction: User Has Not Completed Their Utterance/Thought); after receiving the second sequence of acoustic frames, receiving a third sequence of acoustic frames characterizing a third portion of the utterance spoken by the user (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection… see also p. 0183 - …The CI manage module 106 cooperating with the spoken language understanding module 104 looks for grammatical completeness of sentence syntax in the flow of speech coming from the user. If a user initially responds “Yeah that looks good but . . . ” (first sequence), the CI manager module 106 is configured to understand that this in an incomplete human sentence. Then the user might subsequently state after the long pause “I am not sure on Tuesday, maybe Wednesday! (second sequence)” Thus, if the CI manager module 106 pairs this initial flow of speech with a subsequent flow of speech from the user, then possibly a grammatically complete sentence can be sent to the spoken language understanding module 104 to get a correct interpretation of the speech from the user and without taking the conversational floor from the user before they completely convey the concept in the flow of speech in which they were attempting to convey with those two broken up phrases…; see also p. 0197 - …speaking with pauses between two or more user utterances (processing first sequence followed by a second sequence, followed by a third sequence) so that the user can respond initially incompletely with a first utterance followed by a pause and then a second utterance to complete a thought the user is trying to convey with that speech activity…); and processing, using the speech recognition model, the third sequence of acoustic frames to: generate second speech recognition results for the third portion of the utterance (Bratt; p. 0177 - The VDA computes and compares the data from the frames of the speech waveform to a database and subsequent classification module (speech recognition). Note, each sample of the speech signal is processed to generate the endpoint signal, then the next sample is processed (sequential processing of a sequence of frames). The new sample will be used to update the endpoint signal. The acoustic front end can include a pause analysis analytics engine, a duration pattern analytics engine, a loudness analytics engine, and a pitch processing analytics engine; see also p. 0039 - The automatic audio processing module 102 includes components and performs the functions of automated speech recognition including speech activity detection…); and detect an end of speech event at an end of the third portion of the utterance (Bratt; p. 0040 - The CI manager module 106 uses the input from the prosodic detector to determine i) whether the user has indeed yielded the conversational floor (end of speech) or ii) whether the user is inserting pauses into a flow of their speech to convey additional information. Note, the additional information can include 1) speaking with pauses to help convey and understand a long list of information, 2) speaking with pauses between two or more user utterances such that the user responds initially incompletely with a first utterance followed by a pause and then a later utterance to complete a thought the user is trying to convey with that speech activity, as well as 3) any combination of these two; see also p. 0177 – Each of these analytics engines can have executable software using algorithms specifically for performing that particular function. For example, the pause analytics engine can utilize a conventional “speech/no-speech” algorithm that detects when a pause in the speech occurs. The output is a binary value that indicates whether the present speech signal sample is a portion of speech or not a portion of speech. This output and determination information can be used to identify an endpoint; see also p. 0182-0183); and based on detecting the end of speech event, triggering a microphone closing event by the user device. Bratt, however, fails to disclose based on detecting the end of speech event, triggering a microphone closing event by the user device. Simko does teach based on detecting the end of speech event, triggering a microphone closing event by the user device (Simko; p. 0077 - In instances where the speech decoder and the end of query model determine at approximately the same time whether the user has likely finished speaking, then the system may use both determinations to generate an instruction to close the microphone). Therefore, it would have been obvious to one of ordinary skill In the art to modify the system of Bratt to include based on detecting the end of speech event, triggering a microphone closing event by the user device, as taught by Simko, because a user may use the voice input capabilities of a computing device and speak at a pace that is comfortable for the user. This may increase the utility of the computing device for users, in particular for users with speech disorders or impediments. An utterance may be endpointed at the intended end of the utterance, leading to more accurate or desirable natural language processing outputs, and to faster processing by the natural language processing system. This can reduce the use of computational resources and can conserve power. Moreover, closing of the microphone at a more suitable point can further reduce the use of computational resources and conserve power, since the microphone does not need to be activated, and the use of computational resources in interpreting and performing tasks based on additional audio detected by the microphone can be avoided (Simko; p. 0016).
As per claim 12, the claim is directed to the system of claim 11, and contains similar subject matter as claim 2. Therefore, the claim is rejected similarly. As per claim 13, the claim is directed to the system of claim 12, and contains similar subject matter as claim 3. Therefore, the claim is rejected similarly.
As per claim 14, the claim is directed to the system of claim 13, and contains similar subject matter as claim 4. Therefore, the claim is rejected similarly.
As per claim 15, the claim is directed to the system of claim 11, and contains similar subject matter as claim 5. Therefore, the claim is rejected similarly.
As per claim 16, the claim is directed to the system of claim 11, and contains similar subject matter as claim 6. Therefore, the claim is rejected similarly.
As per claim 17, the claim is directed to the system of claim 16, and contains similar subject matter as claim 7. Therefore, the claim is rejected similarly.
As per claim 18, the claim is directed to the system of claim 16, and contains similar subject matter as claim 8. Therefore, the claim is rejected similarly.
As per claim 19, the claim is directed to the system of claim 11, and contains similar subject matter as claim 9. Therefore, the claim is rejected similarly.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Bratt in view of Liu et al. (US Patent 12,002,451; hereinafter “Liu”).
As per claim 10, Bratt discloses: The computer-implemented method of claim 1, upon which claim 10 depends. Bratt, however, fails to disclose wherein the speech recognition model comprises a stack of self-attention blocks. Liu does teach wherein the speech recognition model comprises a stack of self-attention blocks (Liu; Col. 8, lines 13-35 - The audio encoder 160, in some embodiments, may be comprised of stacked self-attention transformer layers). Therefore, it would have been obvious to one of ordinary skill In the art to modify the method of Bratt to include wherein the speech recognition model comprises a stack of self-attention blocks, as taught by Liu, in order to determine a relevance of the context data to a spoken input from a user (Liu; Col. 3, lines 5-8).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Bratt in view of Simko and further in view of Liu.
As per claim 20, the claim is directed to the system of claim 11, and contains similar subject matter as claim 10. Therefore, the claim is rejected similarly.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The prior art made of record and not relied upon includes: Wu (US PG Pub 20220351718) discloses a computing system that is configured to generate a transformer-transducer-based deep neural network. The transformer-transducer-based deep neural network comprises a transformer encoder network and a transducer predictor network. The transformer encoder network has a plurality of layers, each of which includes a multi-head attention network sublayer and a feed-forward network sublayer. The computing system trains an end-to-end (E2E) automatic speech recognition (ASR) model, using the transformer-transducer-based deep neural network. The E2E ASR model has one or more adjustable hyperparameters that are configured to dynamically adjust an efficiency or a performance of E2E ASR model when the E2E ASR model is deployed onto a device or executed by the device (Wu; Abstract). Chen, Qian et al. “Controllable Time-Delay Transformer for Real-Time Punctuation Prediction and Disfluency Detection.” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2020): 8069-8073. This paper discloses a Controllable Time-delay Transformer (CT-Transformer) model that jointly completes the punctuation prediction and disfluency detection tasks in real time (Abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Rodrigo A Chavez whose telephone number is (571)270-0139. The examiner can normally be reached Monday - Friday 9-6 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached on 5712727602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RODRIGO A CHAVEZ/Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658