DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Terminal Disclaimer
The Terminal Disclaimer filed 7/10/26 linking the instant application to its parent Patent US 12,327,553 B2 is acknowledged and approved (7/10/26)
Response to Arguments
Applicant's arguments filed 7/10/26 have been fully considered but they are not persuasive.
Regarding claims 1, 9 and 17, Applicant argues that the cited portions of Bratt fail to disclose to “iteratively process audio-based characteristics" as input to generate each of "a first predicted measure associated with synthesized speech audio data that captures fulfillment output to be audibly rendered via one or more speakers of the client device, a second predicted measure associated with synthesized speech audio data that captures natural conversation output to be audibly rendered via one or more speakers of the client device, and a third predicted measure associated with refraining from causing any synthesized speech audio data to be audibly rendered via one or more speakers of the client device" or "the next interaction state" that is implemented to be determined based on the different "predicted measures". (Arguments, pg. 13 – pg. 15, third para.). Examiner respectfully disagrees.
Bratt discloses a voice-based digital assistant (VDA) system including a CI manager module 106 that uses its rule-based engine to analyze and make determinations based on the flow of the speech to and from a user via analyzing conversational cues including the pitch and prosody of spoken words of a user as well as pausing of the user (para. [0027]; para. [0031]; para. [0039]) i.e., audio-based characteristics, corresponding to limitation “determining, based on processing the stream of audio data, audio-based characteristics associated with the one or more spoken utterances”.
Bratt also discloses the CI manager module 106 using hybrid approach of a rule-based engine in the CI manager module 106 as well as a trained machine-learning model portion to analyze the user’s speech/spoken words and make decisions and determinations on what steps to take next based on the obtained conversational cues of pitch and prosody determined via analyzing the user’s speech flow (para. [0027]; para. [0188]), where the analysis involves performing prosodic analysis at the end of the user’s utterance and during the user’s utterance (para. [0071]; para. [0197]) to determine next/predicted decisions based on the obtained pitch and prosody audio characteristics/cues (i.e., continual/iterative processing), corresponding to limitation “iteratively processing, using a machine learning (ML) model, at least the audio-based characteristics to generate a plurality of predicted measures associated with a next interaction state to be implemented”.
Bratt further discloses the CI manager module as making decisions/determinations on what steps to take next based on the analyzed pitch and prosody, where the decisions/determinations include to indicate that the VDA system has a desire to grab the conversational floor to query the user or respond to the user's request (para. [0032]), where the response to the user’s request is generated via a loudspeaker Text-To-Speech (TTS) response (para. [0048]), corresponding to limitation “the plurality of predicted measures including at least: a first predicted measure associated with causing synthesized speech audio data that captures fulfillment output to be audibly rendered via one or more speakers of the client device”
Bratt also discloses the CI manager module as making decisions/determinations on what steps to take next based on the analyzed pitch and prosody, where the decisions/determinations include to 1) to prompt additional information from the user including asking the user to confirm a request the user provided e.g. the system stating “So then you want to make a reservation for a hotel room in Rome near walking distance within Trevi Fountain?” (para. [0029]; para. [0032]), where the prompt in response to the user’s request is generated via a loudspeaker Text-To-Speech (TTS) response (para. [0048]), corresponding to limitation “a second predicted measure associated with causing synthesized speech audio data that captures natural conversation output to be audibly rendered via the one or more speakers of the client device”.
Bratt further discloses the CI manager module as making decisions/determinations on what steps to take next based on the analyzed pitch and prosody, where If user is holding the floor prosodically and the user’s utterance is incomplete, then the VDA system sets a wait time to long fixed setting while waiting to take over a conversational floor from a user before prompting a user to continue speaking (para. [0062]; para. [0083]-[0086]), corresponding to limitation “a third predicted measure associated with refraining from causing any synthesized speech audio data to be audibly rendered via the one or more speakers of the client device”.
Also given Bratt discloses the first, second and third predicted measures to be associated with a next step performed by the VDA as provided above, Bratt discloses limitations “determining, based on the plurality of predicted measures, the next interaction state to be implemented” and “causing the next interaction state to be implemented”.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claims 1, 9 and 17 contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. In particular, claims 1, 9 and 17 recite steps of “determining, based on processing the stream of audio data, audio-based characteristics associated with the one or more spoken utterances” and “iteratively processing, using a machine learning (ML) model, at least the audio- based characteristics to generate a plurality of predicted measures associated with a next interaction state to be implemented, the plurality of predicted measures including at least:…”. There is no disclosure of “iteratively processing, using a machine learning (ML) model, at least the audio- based characteristics to generate a plurality of predicted measures” in Applicant’s original specification/drawings.
Paragraphs [0005]-[0006] and [0060]-[0064] of the original specification which Applicant provides as supporting the amendments to the claims, as well as figure 3B, describe a step 360/260A of processing, using a classification ML model, the current state of the stream of NLU output, the stream of fulfillment data, and/or the audio-based characteristics to generate corresponding predicted measures, as well as a subsequent step 384 of determine, based on the corresponding predicted measures, whether a next interaction state is: (i) causing fulfillment output to be implemented, (ii) causing natural conversation output to be audibly rendered for presentation to a user, or (iii) refraining from causing any interaction to be implemented, where at an iteration of step 384, the system determines, based on the corresponding predicted measures, the next interaction state is (i) causing fulfillment output to be implemented, then the system may proceed to perform the state.
The portions do not describe iteratively processing, using a machine learning (ML) model, at least the audio-based characteristics to generate a plurality of predicted measures associated with a next interaction state to be implemented. The dependent claims are rejected based on their dependency.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
1. Claims 1-3, 6-11, 14-18 and 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Bratt et al US 2022/0115001 A1 (“Bratt” - IDS)
Per claim 1, Bratt discloses a method implemented by one or more processors, the method comprising:
processing a stream of audio data, the stream of audio data being generated by one or more microphones of a client device of a user, and the stream of audio data capturing one or more spoken utterances of the user that are directed to an automated assistant implemented at least in part at the client device (A voice-based digital assistant (VDA) …, Abstract; para. [0020]; For example, the user may verbally state, "Find me a hotel in Rome by Trevi Fountain." para. [0027]; para. [0036]; para. [0039]; para. [0041]; suppose the user says, "what hours is Wells Fargo open?", para. [0053]; A user may enter commands and information into the computing device 402 through input devices such as a keyboard, touchscreen a microphone 463 The microphone 463 can cooperate with speech recognition software, para. [0215]);
determining, based on processing the stream of audio data, audio-based characteristics associated with the one or more spoken utterances (The CI manager module 106 uses the rule-based engine to analyze and make determinations on the flow of the speech to and from the user via analyzing, for example, non-lexical sounds, pitch and/or prosody of the spoken words, pausing, and grammatical completeness …, para. [0027]; The rule-based engine has rules to analyze and make determinations on two or more conversational cues of i) non-lexical items, ii) pitch of spoken words, iii) prosody of spoken words … Note, the pitch of spoken words can be a part of prosody…., para. [0031]; The CI manager module 106 has a prosodic detector …, para. [0039], prosody and/or pitch as example audio-based characteristics);
iteratively processing, using a machine learning (ML) model, at least the audio- based characteristics to generate a plurality of predicted measures associated with a next interaction state to be implemented (para. [0020]; Based on the prosody and pitch of those words, and optionally a pause after the last words “Trevi Fountain,” the CI manager module 106 uses the rule-based engine to analyze and make a determination. For example …, para. [0027]; The CI manager module 106 uses the rule-based engine to analyze and make determinations on factors of conversational cues. The rule-based engine has rules to analyze and make determinations on two or more conversational cues of i) non-lexical items, ii) pitch of spoken words, iii) prosody of spoken words …, para. [0031]; para. [0039]; para. [0071]; para. [0093]; the CI manager module 106 can be configured to use a hybrid approach of 1) a rule-based engine in the CI manager module 106 as well as a trained, machine-learning model portion to analyze and make decisions …, para. [0188]; The prosodic detector initially checks to detect whether any speech activity is occurring from the Automatic Audio Processing module and then to apply the prosodic analysis at ‘an end of’ and/or ‘during’ a user's utterance …, para [0197], analyzing prosody of received utterance at the end of the utterance and during the utterance as iteratively processing prosody/audio-based characteristics),
the plurality of predicted measures including at least: a first predicted measure associated with causing synthesized speech audio data that captures fulfillment output to be audibly rendered via one or more speakers of the client device (para. [0027]- [0031]; The CI manager module, after making these determinations and analysis, can then decide whether to generate an utterance during the time frame when the user still holds the conversational floor in order to at least one of 1) to prompt additional information from the user, 2) to signal the user to hold the conversational floor and continue to speak, or 3) to indicate that the VDA has a desire to grab the conversational floor … Thus, the CI manager module 106 may take the conversational floor to query the user or respond to the user's request …, para. [0032]; para. [0048], responding to user request via loudspeaker Text-To-Speech (TTS) response based on analyzed prosody/pitch as utilizing predicted fulfillment output),
a second predicted measure associated with causing synthesized speech audio data that captures natural conversation output to be audibly rendered via the one or more speakers of the client device (Alternatively, the CI manager module 106 uses the rule-based engine to analyze and make an example determination on when the user issues a single utterance forming a complete thought, then the CI manager module 106 will know to take over the conversational floor in the on-going dialogue between the user and the VDA. For example, the VDA may then reference the dialogue manager module 108 and repeat back the current topic of the dialogue to the user with a complete utterance. For example, the VDA may state, “So then you want to make a reservation for a hotel room in Rome near walking distance within Trevi Fountain?” …, para. [0029]; para. [0048], loudspeaker produced Text-To-Speech (TTS ) system response requesting confirmation based on analyzed prosody/pitch as synthesized speech audio data that captures natural conversation output), and
a third predicted measure associated with refraining from causing any synthesized speech audio data to be audibly rendered via the one or more speakers of the client device (para. [0028]; para. [0030]-[0032]; para. [0053]; when the user has the conversational floor, the VDA can take an example action of waiting longer for input from the user before prompting the user to continue, para. [0062]; para. [0083]; If user is holding the floor prosodically, then: … set wait time to long fixed setting and then take over conversational floor …, para. [0084]-[0086]; para. [0183], waiting to take over the conversation floor while the user is yet to complete their thought based on prosodic analysis as implying refraining from providing synthesized speech);
determining, based on the plurality of predicted measures, the next interaction state to be implemented (para. [0027]-[0032]; para. [0071]; para. [0083]-[0086]); and
causing the next interaction state to be implemented (para. [0027]-[0032]; para. [0071]; para. [0083]-[0086]).
Per claim 2, Bratt discloses the method of claim 1, wherein determining the next interaction state to be implemented based on the plurality of predicted measures comprises determining to cause the synthesized speech audio data that captures the fulfillment output to be audibly rendered via the one or more speakers of the client device (para. [0027]- [0031]; The CI manager module, after making these determinations and analysis, can then decide whether to generate an utterance during the time frame when the user still holds the conversational floor in order to at least one of 1) to prompt additional information from the user, 2) to signal the user to hold the conversational floor and continue to speak, or 3) to indicate that the VDA has a desire to grab the conversational floor … Thus, the CI manager module 106 may take the conversational floor to query the user or respond to the user's request …, para. [0032]; para. [0048], responding to user request via loudspeaker Text-To-Speech (TTS) response based on analyzed prosody/pitch as utilizing predicted fulfillment output).
Per claim 3, Bratt discloses the method of claim 2, wherein the fulfillment output is generated by the automated assistant and based on processing the stream of audio data (para. [0027]-[0032]; para. [0053]; para. [0063]).
Per claim 6, Bratt discloses the method of claim 1, wherein determining the next interaction state to be implemented based on the plurality of predicted measures comprises: determining to cause the synthesized speech audio data that captures the natural conversation output to be audibly rendered via the one or more speakers of the client device (para. [0027]; Alternatively, the CI manager module 106 uses the rule-based engine to analyze and make an example determination on when the user issues a single utterance forming a complete thought, then the CI manager module 106 will know to take over the conversational floor in the on-going dialogue between the user and the VDA. For example, the VDA may then reference the dialogue manager module 108 and repeat back the current topic of the dialogue to the user with a complete utterance. For example, the VDA may state, “So then you want to make a reservation for a hotel room in Rome near walking distance within Trevi Fountain?” …, para. [0029]; para. [0031]; para. [0048], loudspeaker produced Text-To-Speech (TTS ) system response requesting confirmation based on analyzed prosody/pitch as synthesized speech audio data that captures natural conversation output).
Per claim 7, Bratt discloses the method of claim 1, wherein the audio-based characteristics associated with one or more of the spoken utterances comprise the one or more of: an intonation of the one or more of the spoken utterances, a cadence of the one or more of the spoken utterances, or a duration of time that has elapsed between speaking the one or more of the spoken utterances (Abstract; The CI manager module 106 uses the rule-based engine to analyze and make determinations on the flow of the speech to and from the user via analyzing, for example, non-lexical sounds, pitch and/or prosody of the spoken words, pausing, and grammatical completeness …, para. [0027]-[0028]; para. [0133]; para. [0179]; para. [0189], pitch/prosody as intonation).
Per claim 8, Bratt discloses the method of claim 1, wherein determining the next interaction state to be implemented based on the plurality of predicted measures comprises: determining to refrain from causing any synthesized speech audio data to be audibly rendered via the one or more speakers of the client device (para. [0028]; para. [0030]-[0032]; para. [0053]; when the user has the conversational floor, the VDA can take an example action of waiting longer for input from the user before prompting the user to continue, para. [0062]; para. [0083]; If user is holding the floor prosodically, then: … set wait time to long fixed setting and then take over conversational floor …, para. [0084]-[0086]; para. [0183], waiting to take over the conversation floor while the user is yet to complete their thought based on prosodic analysis as implying refraining from providing synthesized speech); and
in response to refraining from causing any synthesized speech audio data to be audibly rendered via the one or more of the speakers of the client device: continuing to process the stream of audio data (para. [0083]-[0086]; para. [0183], waiting for a predetermined wait time to present response after receiving user utterance while the user is holding the floor as implying continuing to process the user utterance/audio stream).
Per claim 9, Bratt discloses a system comprising:
at least one processor (para. [0212]); and
memory storing instructions that, when executed, cause the at least one processor to be operable to: process a stream of audio data, the stream of audio data being generated by one or more microphones of a client device of a user, and the stream of audio data capturing one or more spoken utterances of the user that are directed to an automated assistant implemented at least in part at the client device (A voice-based digital assistant (VDA) …, Abstract; para. [0020]; For example, the user may verbally state, "Find me a hotel in Rome by Trevi Fountain." para. [0027]; para. [0036]; para. [0039]; para. [0041]; suppose the user says, "what hours is Wells Fargo open?" , para. [0053]; para. [0212]; A user may enter commands and information into the computing device 402 through input devices such as a keyboard, touchscreen a microphone 463 The microphone 463 can cooperate with speech recognition software, para. [0215]);
determine, based on processing the stream of audio data, audio-based characteristics associated with the one or more spoken utterances (The CI manager module 106 uses the rule-based engine to analyze and make determinations on the flow of the speech to and from the user via analyzing, for example, non-lexical sounds, pitch and/or prosody of the spoken words, pausing, and grammatical completeness …, para. [0027]; The rule-based engine has rules to analyze and make determinations on two or more conversational cues of i) non-lexical items, ii) pitch of spoken words, iii) prosody of spoken words … Note, the pitch of spoken words can be a part of prosody…., para. [0031]; The CI manager module 106 has a prosodic detector …, para. [0039], prosody and/or pitch as example audio-based characteristics);
iteratively process, using a machine learning (ML) model, at least the audio- based characteristics to generate a plurality of predicted measures associated with a next interaction state to be implemented (para. [0027]; The CI manager module 106 uses the rule-based engine to analyze and make determinations on factors of conversational cues. The rule-based engine has rules to analyze and make determinations on two or more conversational cues of i) non-lexical items, ii) pitch of spoken words, iii) prosody of spoken words …, para. [0031]; para. [0039]; para. [0071]; para. [0093]; the CI manager module 106 can be configured to use a hybrid approach of 1) a rule-based engine in the CI manager module 106 as well as a trained, machine-learning model portion to analyze and make decisions …, para. [0188]; The prosodic detector initially checks to detect whether any speech activity is occurring from the Automatic Audio Processing module and then to apply the prosodic analysis at ‘an end of’ and/or ‘during’ a user's utterance …, para [0197], analyzing prosody of received utterance at the end of the utterance and during the utterance as iteratively processing prosody/audio-based characteristics),
the plurality of predicted measures including at least: a first predicted measure associated with causing synthesized speech audio data that captures fulfillment output to be audibly rendered via one or more speakers of the client device (para. [0027]- [0031]; The CI manager module, after making these determinations and analysis, can then decide whether to generate an utterance during the time frame when the user still holds the conversational floor in order to at least one of 1) to prompt additional information from the user, 2) to signal the user to hold the conversational floor and continue to speak, or 3) to indicate that the VDA has a desire to grab the conversational floor … Thus, the CI manager module 106 may take the conversational floor to query the user or respond to the user's request …, para. [0032]; para. [0048], responding to user request via loudspeaker Text-To-Speech (TTS) response based on analyzed prosody/pitch as utilizing predicted fulfillment output),
a second predicted measure associated with causing synthesized speech audio data that captures natural conversation output to be audibly rendered via the one or more speakers of the client device (Alternatively, the CI manager module 106 uses the rule-based engine to analyze and make an example determination on when the user issues a single utterance forming a complete thought, then the CI manager module 106 will know to take over the conversational floor in the on-going dialogue between the user and the VDA. For example, the VDA may then reference the dialogue manager module 108 and repeat back the current topic of the dialogue to the user with a complete utterance. For example, the VDA may state, “So then you want to make a reservation for a hotel room in Rome near walking distance within Trevi Fountain?” …, para. [0029]; para. [0048], loudspeaker produced Text-To-Speech (TTS ) system response requesting confirmation based on analyzed prosody/pitch as synthesized speech audio data that captures natural conversation output), and
a third predicted measure associated with refraining from causing any synthesized speech audio data to be audibly rendered via the one or more speakers of the client device (para. [0028]; para. [0030]-[0032]; para. [0053]; when the user has the conversational floor, the VDA can take an example action of waiting longer for input from the user before prompting the user to continue, para. [0062]; para. [0083]; If user is holding the floor prosodically, then: … set wait time to long fixed setting and then take over conversational floor …, para. [0084]-[0086]; para. [0183], waiting to take over the conversation floor while the user is yet to complete their thought based on prosodic analysis as implying refraining from providing synthesized speech);
determine, based on the plurality of predicted measures, the next interaction state to be implemented (para. [0027]-[0032]; para. [0071]; para. [0083]-[0086]); and
cause the next interaction state to be implemented (para. [0027]-[0032]; para. [0071]; para. [0083]-[0086]).
Per claim 10, Bratt discloses the system of claim 9, wherein the instructions to determine the next interaction state to be implemented based on the plurality of predicted measures comprise instructions to determine to cause the synthesized speech audio data that captures the fulfillment output to be audibly rendered via the one or more speakers of the client device (para. [0027]- [0031]; The CI manager module, after making these determinations and analysis, can then decide whether to generate an utterance during the time frame when the user still holds the conversational floor in order to at least one of 1) to prompt additional information from the user, 2) to signal the user to hold the conversational floor and continue to speak, or 3) to indicate that the VDA has a desire to grab the conversational floor … Thus, the CI manager module 106 may take the conversational floor to query the user or respond to the user's request …, para. [0032]; para. [0048], responding to user request via loudspeaker Text-To-Speech (TTS) response based on analyzed prosody/pitch as utilizing predicted fulfillment output).
Per claim 11, Bratt discloses the system of claim 10, wherein the fulfillment output is generated by the automated assistant and based on processing the stream of audio data (para. [0029]-[0032]; para. [0053]; para. [0063]).
Per claim 14, Bratt discloses the system of claim 9, wherein the instructions to determine the next interaction state to be implemented based on the plurality of predicted measures comprise instructions to: determine to cause the synthesized speech audio data that captures the natural conversation output to be audibly rendered via the one or more speakers of the client device (para. [0027]; Alternatively, the CI manager module 106 uses the rule-based engine to analyze and make an example determination on when the user issues a single utterance forming a complete thought, then the CI manager module 106 will know to take over the conversational floor in the on-going dialogue between the user and the VDA. For example, the VDA may then reference the dialogue manager module 108 and repeat back the current topic of the dialogue to the user with a complete utterance. For example, the VDA may state, “So then you want to make a reservation for a hotel room in Rome near walking distance within Trevi Fountain?” …, para. [0029]; para. [0031]; para. [0048], loudspeaker produced Text-To-Speech (TTS ) system response requesting confirmation based on analyzed prosody/pitch as synthesized speech audio data that captures natural conversation output).
Per claim 15, Bratt discloses the system of claim 14, wherein the audio-based characteristics associated with the one or more of the spoken utterances comprise one or more of: an intonation of the one or more of the spoken utterances, a cadence of the one or more of the spoken utterances, or a duration of time that has elapsed between speaking the one or more of the spoken utterances (Abstract; The CI manager module 106 uses the rule-based engine to analyze and make determinations on the flow of the speech to and from the user via analyzing, for example, non-lexical sounds, pitch and/or prosody of the spoken words, pausing, and grammatical completeness …, para. [0027]-[0028]; para. [0133]; para. [0179]; para. [0189], pitch/prosody as intonation).
Per claim 16, Bratt discloses the system of claim 9, wherein the instructions to determine the next interaction state to be implemented based on the plurality of predicted measures comprise instructions to: determine to refrain from causing any synthesized speech audio data to be audibly rendered via the one or more speakers of the client device (para. [0028]; para. [0030]-[0032]; para. [0053]; when the user has the conversational floor, the VDA can take an example action of waiting longer for input from the user before prompting the user to continue, para. [0062]; para. [0083]; If user is holding the floor prosodically, then: … set wait time to long fixed setting and then take over conversational floor …, para. [0084]-[0086]; para. [0183], waiting to take over the conversation floor while the user is yet to complete their thought based on prosodic analysis as implying refraining from providing synthesized speech); and
in response to refraining from causing any synthesized speech audio data to be audibly rendered via the one or more of the speakers of the client device: continuing to process the stream of audio data (para. [0083]-[0086]; para. [0183], waiting for a predetermined wait time to present response after receiving user utterance while the user is holding the floor as implying continuing to process the user utterance/audio stream).
Per claim 17, Bratt discloses a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations to:
process a stream of audio data, the stream of audio data being generated by one or more microphones of a client device of a user, and the stream of audio data capturing one or more spoken utterances of the user that are directed to an automated assistant implemented at least in part at the client device (A voice-based digital assistant (VDA) …, Abstract; para. [0020]; For example, the user may verbally state, "Find me a hotel in Rome by Trevi Fountain." para. [0027]; para. [0036]; para. [0039]; para. [0041]; suppose the user says, "what hours is Wells Fargo open?" , para. [0053]; A user may enter commands and information into the computing device 402 through input devices such as a keyboard, touchscreen a microphone 463 The microphone 463 can cooperate with speech recognition software, para. [0215]);
determine, based on processing the stream of audio data, audio-based characteristics associated with the one or more spoken utterances (The CI manager module 106 uses the rule-based engine to analyze and make determinations on the flow of the speech to and from the user via analyzing, for example, non-lexical sounds, pitch and/or prosody of the spoken words, pausing, and grammatical completeness …, para. [0027]; The rule-based engine has rules to analyze and make determinations on two or more conversational cues of i) non-lexical items, ii) pitch of spoken words, iii) prosody of spoken words … Note, the pitch of spoken words can be a part of prosody…., para. [0031]; The CI manager module 106 has a prosodic detector …, para. [0039], prosody and/or pitch as example audio-based characteristics);
iteratively process, using a machine learning (ML) model, at least the audio- based characteristics to generate a plurality of predicted measures associated with a next interaction state to be implemented (para. [0027]; The CI manager module 106 uses the rule-based engine to analyze and make determinations on factors of conversational cues. The rule-based engine has rules to analyze and make determinations on two or more conversational cues of i) non-lexical items, ii) pitch of spoken words, iii) prosody of spoken words …, para. [0031]; para. [0039]; para. [0071]; para. [0093]; the CI manager module 106 can be configured to use a hybrid approach of 1) a rule-based engine in the CI manager module 106 as well as a trained, machine-learning model portion to analyze and make decisions …, para. [0188]; The prosodic detector initially checks to detect whether any speech activity is occurring from the Automatic Audio Processing module and then to apply the prosodic analysis at ‘an end of’ and/or ‘during’ a user's utterance …, para [0197], analyzing prosody of received utterance at the end of the utterance and during the utterance as iteratively processing prosody/audio-based characteristics),
the plurality of predicted measures including at least: a first predicted measure associated with causing synthesized speech audio data that captures fulfillment output to be audibly rendered via one or more speakers of the client device (para. [0027]- [0031]; The CI manager module, after making these determinations and analysis, can then decide whether to generate an utterance during the time frame when the user still holds the conversational floor in order to at least one of 1) to prompt additional information from the user, 2) to signal the user to hold the conversational floor and continue to speak, or 3) to indicate that the VDA has a desire to grab the conversational floor … Thus, the CI manager module 106 may take the conversational floor to query the user or respond to the user's request …, para. [0032]; para. [0048], responding to user request via loudspeaker Text-To-Speech (TTS) response based on analyzed prosody/pitch as utilizing predicted fulfillment output),
a second predicted measure associated with causing synthesized speech audio data that captures natural conversation output to be audibly rendered via the one or more speakers of the client device (Alternatively, the CI manager module 106 uses the rule-based engine to analyze and make an example determination on when the user issues a single utterance forming a complete thought, then the CI manager module 106 will know to take over the conversational floor in the on-going dialogue between the user and the VDA. For example, the VDA may then reference the dialogue manager module 108 and repeat back the current topic of the dialogue to the user with a complete utterance. For example, the VDA may state, “So then you want to make a reservation for a hotel room in Rome near walking distance within Trevi Fountain?” …, para. [0029]; para. [0048], loudspeaker produced Text-To-Speech (TTS ) system response requesting confirmation based on analyzed prosody/pitch as synthesized speech audio data that captures natural conversation output), and
a third predicted measure associated with refraining from causing any synthesized speech audio data to be audibly rendered via the one or more speakers of the client device (para. [0028]; para. [0030]-[0032]; para. [0053]; when the user has the conversational floor, the VDA can take an example action of waiting longer for input from the user before prompting the user to continue, para. [0062]; para. [0083]; If user is holding the floor prosodically, then: … set wait time to long fixed setting and then take over conversational floor …, para. [0084]-[0086]; para. [0183], waiting to take over the conversation floor while the user is yet to complete their thought based on prosodic analysis as implying refraining from providing synthesized speech);
determine, based on the plurality of predicted measures, the next interaction state to be implemented (para. [0027]-[0032]; para. [0071]; para. [0083]-[0086]); and
cause the next interaction state to be implemented (para. [0027]-[0032]; para. [0071]; para. [0083]-[0086]).
Per claim 18, Bratt discloses the non-transitory computer-readable storage medium of claim 17, wherein the operations to determine the next interaction state to be implemented based on the plurality of predicted measures comprise operations to determine to cause the synthesized speech audio data that captures the fulfillment output to be audibly rendered via the one or more speakers of the client device (para. [0027]- [0031]; The CI manager module, after making these determinations and analysis, can then decide whether to generate an utterance during the time frame when the user still holds the conversational floor in order to at least one of 1) to prompt additional information from the user, 2) to signal the user to hold the conversational floor and continue to speak, or 3) to indicate that the VDA has a desire to grab the conversational floor … Thus, the CI manager module 106 may take the conversational floor to query the user or respond to the user's request …, para. [0032]; para. [0048], responding to user request via loudspeaker Text-To-Speech (TTS) response based on analyzed prosody/pitch as utilizing predicted fulfillment output).
Per claim 20, Bratt discloses the non-transitory computer-readable storage medium of claim 17, wherein the audio-based characteristics associated with the one or more of the spoken utterances comprise one or more of: an intonation of the one or more of the spoken utterances, a cadence of the one or more of the spoken utterances, or a duration of time that has elapsed between speaking the one or more of the spoken utterances (Abstract; The CI manager module 106 uses the rule-based engine to analyze and make determinations on the flow of the speech to and from the user via analyzing, for example, non-lexical sounds, pitch and/or prosody of the spoken words, pausing, and grammatical completeness …, para. [0027]-[0028]; para. [0133]; para. [0179]; para. [0189], pitch/prosody as intonation).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
2. Claims 4 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Bratt in view of Mois et al US 10,102,844 B1 (“Mois” - IDS)
Per claim 4, Bratt discloses the method of claim 2,
Bratt does not explicitly disclose wherein the fulfillment output is generated by a third-party agent and based on processing the stream of audio data
However, this feature is taught by Mois (Category servers/skills module 262 may further correspond to one or more first party applications and/or third party applications capable of performing various tasks or actions, as well as providing response information for responses to user commands …, col. 11, ln 58- col. 12, ln 13)
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Mois with the method of Bratt in arriving at the missing features of Bratt because such combination would have resulted in ensuring responses tailored for a particular individual (Mois, col. 15, ln 28-43).
Per claim 12, Bratt discloses the system of claim 10,
System claim 12 and method claim 4 are related as the system and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 12 is similarly rejected under the same rationale as applied above with respect to claim 4.
3. Claims 5, 13 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bratt in view of Kanaan et al US 2016/0203002 A1 (“Kanaan”)
Per claim 5, Bratt discloses the method of claim 2, wherein the fulfillment output is selected from a set of multiple candidate fulfillment outputs (para. [0034]; para. [0053]), and
Bratt does not explicitly disclose wherein the method further comprises: as the user continues to provide the one or more spoken utterances: initiating, for each of the multiple candidate given fulfillment outputs, corresponding partial fulfillment, wherein the corresponding partial fulfillment comprises at least: for a first candidate given fulfillment output, from among the multiple candidate given fulfillment outputs, establishing a corresponding connection with one of: a given software application that is accessible by the client device, a given first-party agent, a given third-party agent, or a given additional client device that is in addition to the client device; and for a second candidate given fulfillment output, from among the multiple candidate given fulfillment outputs, establishing a corresponding connection with another one of: the given software application that is accessible by the client device, the given first-party agent, the given third-party agent, or the given additional client device that is in addition to the client device
However, this feature is taught by Kanaan:
wherein the method further comprises: as the user continues to provide the one or more spoken utterances: initiating, for each of the multiple candidate given fulfillment outputs, corresponding partial fulfillment, wherein the corresponding partial fulfillment comprises at least: for a first candidate given fulfillment output, from among the multiple candidate given fulfillment outputs, establishing a corresponding connection with one of: a given software application that is accessible by the client device, a given first-party agent, a given third-party agent, or a given additional client device that is in addition to the client device (As another example, some partial command strings can be limited to a small set of applications defined in the command data structure 140, and the set of applications can be warmed up in parallel when there is a match on the partial command string. Specifically, the command data structure 140 may have only two applications with commands having the word “take,” such as a camera application with a command “take a picture,” and a memo application with a command “take a memo.” The control logic 224 can begin warming up both the camera application and the memo application when the word “take” is recognized …, para. [0042]; the command data structure 140 can map user voice commands to functions supported by available third-party voice-enabled applications, para. [0045]-[0046]; para. [0048]); and
for a second candidate given fulfillment output, from among the multiple candidate given fulfillment outputs, establishing a corresponding connection with another one of: the given software application that is accessible by the client device, the given first-party agent, the given third-party agent, or the given additional client device that is in addition to the client device (As another example, some partial command strings can be limited to a small set of applications defined in the command data structure 140, and the set of applications can be warmed up in parallel when there is a match on the partial command string. Specifically, the command data structure 140 may have only two applications with commands having the word “take,” such as a camera application with a command “take a picture,” and a memo application with a command “take a memo.” The control logic 224 can begin warming up both the camera application and the memo application when the word “take” is recognized …, para. [0042]; the command data structure 140 can map user voice commands to functions supported by available third-party voice-enabled applications, para. [0045]-[0046]; para. [0048], functions performed by each application (memo application or camera application) as example fulfilment output)
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Kanaan with the method of Bratt in in arriving at the missing features of Bratt because such combination would have resulted in reducing the time it takes to respond to the user (Kanaan, para. [0027])
Per claim 13, Bratt discloses the system of claim 10,
System claim 13 and method claim 5 are related as the system and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 13 is similarly rejected under the same rationale as applied above with respect to claim 5.
Per claim 19, Bratt discloses the non-transitory computer-readable storage medium of claim 18,
Medium claim 19 and method claim 5 are related as the medium and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claim 19 is similarly rejected under the same rationale as applied above with respect to claim 5.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO 892 form.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLUJIMI A ADESANYA whose telephone number is (571)270-3307. The examiner can normally be reached Monday-Friday 8:30-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/OLUJIMI A ADESANYA/Primary Examiner, Art Unit 2658