Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the application and claims filed 02/26/2024. Claims 21-40 are pending and have been examined. Claims 21-40 are rejected.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on Receipt Date 02/26/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim 27, 30, 33, 39 are objected to because of the following informalities:
Claim 27: “so that not decipherable” should be “so that it is not decipherable.”
Claim 30: “fixed-sized representation” should be “fixed-size representation” for consistent terminology.
Claim 33: “processes … to fixed-size representation” should be “processes … into a fixed-size representation.”
Claim 39: “further comprises generate GUI output” should be “further comprises generating GUI output.”
Appropriate correction is required.
Claim 25 is objected to because of the following informalities:
The claim recites “searching for a product” as one of the two or more things. The remaining things that recite “… the product” may have antecedent basis if the selected “two or more” excludes the “searching for a product” option. Suggested fix: “a product” in each instance instead of “the”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 21-40 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
As per claims 21, 35, and 40, these claims contain the limitations "receive audio input from a group of users over a time period..." and "...a fixed-size representation of the group of users, wherein the fixed-size representation indicates a member composition of the group based on user classes detected in the group." However, the specification at no time describes audio input from a group of users, or a group representation whose member composition is detected from audio input. The voice-controlled device throughout the specification interacts with a single user 310, whose audio commands are provided to the trained encoder to generate representations of that one user (paragraphs 0065-0067). Paragraph 0066 comes closest, but describes a multiuser device generating a separate representation for each user based on login, not one representation of the group, and with no class detection. Member composition and user classes appear only in the multi-member account embodiment of FIG. 6 (paragraphs 0080-0084), where composition is predicted from account event sequences (paragraph 0082); no audio input, voice commands, or voice-controlled device is discussed anywhere in that embodiment. Nowhere in the specification does it describe audio input from a group of users or user classes detected from voice commands. This causes the claims to lack adequate written description and therefore rejected under U.S.C. 112(a). In the next response, please indicate where support for these limitations are found within the specification.
As per claims 22-34 and 36-39, these claims are rejected as being dependent on a claim rejected under U.S.C. 112(a) for lack of adequate written description.
As per claim 33, this claim calls for "the encoder model comprises a recurrent neural network (RNN) that processes a sequence of words in a voice command to fixed-size representation." However, the RNN in the specification never reads/processes individual words. Voice input is turned into text commands, and the RNN receives each whole command as one event (paragraphs 0025, 0066). The only disclosed process that turns words into a fixed-size representation is the word model of FIG. 4 (paragraphs 0070-0074). That word model is not a recurrent neural network or the encoder model. Paragraph 0071's representation from a voice command refers to that same process. This causes the claims to lack adequate written description and therefore rejected under U.S.C. 112(a).
As per claim 34, this claim calls for further training of the encoder model after deployment "wherein the further training adapts the encoder model to the group of users." Paragraph 0069 describes partial on-device training that adapts the encoder to attributes of the single user 310; adapting the encoder model to a group of users is nowhere discussed. This causes the claims to lack adequate written description and therefore rejected under U.S.C. 112(a).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 21-40 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 21
Step 1: The claim recites a system; therefore, it is directed to the statutory category of machine.
Step 2A Prong 1: The claim recites the following abstract ideas:
"generate, based on the audio input (...), a fixed-size representation of the group of users, wherein the fixed-size representation indicates a member composition of the group based on user classes detected in the group;" (This limitation is a mental process. A person mentally or with a pen and paper can listen to audio input from a group of users, detect the classes of the users (e.g., whether a speaker is an adult or a child), and write down a representation of the group (e.g., a list of a fixed length) that indicates the member composition of the group based on the user classes the person detected.)
"wherein the (...) uses the fixed-size representation to generate a personalized output for the group of users;" (This limitation is a mental process. A person mentally or with a pen and paper can use a written representation of a group of users to determine a personalized output, such as a personalized suggestion, for the group of users.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"A system, comprising: one or more computers that implement a voice-controlled device, configured to:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off the shelf computers implementing a voice-controlled device as tools to perform the recited abstract ideas.)
"receive audio input from a group of users over a time period, wherein the audio input includes voice commands associated with different items over the time period;" (Data Gather - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)).)
"generate, based on the audio input and via an encoder model trained using one or more machine learning techniques, a fixed-size representation of the group of users..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "encoder model trained using one or more machine learning techniques" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of "generate, based on the audio input (...), a fixed-size representation of the group of users". The claim denotes a generic encoder model and generic machine learning techniques with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"upload the fixed-size representation to a remote service (...);" (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): This limitation amounts to merely transmitting data to a remote service, recited at a high level of generality.)
"(...) a remote service that implements a machine learning model, wherein the machine learning model uses the fixed-size representation to generate a personalized output for the group of users;" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "machine learning model" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of using the fixed-size representation to generate a personalized output for the group of users. The claim denotes a generic machine learning model with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"receive the personalized output from the remote service and generate audio output indicating the personalized output." (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): Receiving the personalized output from the remote service amounts to mere data gathering, and generating audio output indicating the personalized output is the nominal act of outputting or delivering the end result after the core process is complete; therefore, this is interpreted as insignificant pre- and post-solution activity.)
Step 2B:
"A system, comprising: one or more computers that implement a voice-controlled device, configured to:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off the shelf computers implementing a voice-controlled device as tools to perform the recited abstract ideas.)
"receive audio input from a group of users over a time period, wherein the audio input includes voice commands associated with different items over the time period;" (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
"generate, based on the audio input and via an encoder model trained using one or more machine learning techniques, a fixed-size representation of the group of users..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "encoder model trained using one or more machine learning techniques" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of "generate, based on the audio input (...), a fixed-size representation of the group of users". The claim denotes a generic encoder model and generic machine learning techniques with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"upload the fixed-size representation to a remote service (...);" (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
"(...) a remote service that implements a machine learning model, wherein the machine learning model uses the fixed-size representation to generate a personalized output for the group of users;" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "machine learning model" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of using the fixed-size representation to generate a personalized output for the group of users. The claim denotes a generic machine learning model with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"receive the personalized output from the remote service and generate audio output indicating the personalized output." (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 22
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 22 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein the voice-controlled device comprises a smartphone or a television." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off the shelf smartphone or television. The smartphone or television is performing a generic computer function which consists of the abstract ideas recited in claim 21, Step 2A Prong 1. Therefore, it amounts to no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.)
Step 2B:
"wherein the voice-controlled device comprises a smartphone or a television." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a generic, off the shelf smartphone or television to perform the abstract ideas amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 23
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 23 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein the voice-controlled device comprises a vehicle-based computer." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off the shelf vehicle-based computer. The vehicle-based computer is performing a generic computer function which consists of the abstract ideas recited in claim 21, Step 2A Prong 1. Therefore, it amounts to no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.)
Step 2B:
"wherein the voice-controlled device comprises a vehicle-based computer." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a generic, off the shelf vehicle-based computer to perform the abstract ideas amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 24
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 24 depends on. Claim 24 further recites:
"wherein: the group of users are members of a family account; and the user classes indicate different ages and genders of the members." (This limitation is a mental process. A person mentally or with a pen and paper can identify that a group of users are members of a family and can note the different ages and genders of the members.)
Step 2A Prong 2: The claim does not recite additional elements; therefore, the judicial exception is not integrated into a practical application.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 25
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 25 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein the voice commands indicate interactions with different products including two or more of: searching for a product, viewing the product, purchasing the product, returning the product, and providing feedback on the product." (Data Gather - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner's Note (EN): The limitation merely specifies the content of the received data, i.e., that the received voice commands indicate interactions with different products.)
Step 2B:
"wherein the voice commands indicate interactions with different products including two or more of: searching for a product, viewing the product, purchasing the product, returning the product, and providing feedback on the product." (This falls under well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). See MPEP 2106.05(d)(II).)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 26
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 25 above, which claim 26 depends on. Claim 26 further recites:
"wherein the personalized output indicates a product recommendation to the group of users." (This limitation is a mental process. A person mentally or with a pen and paper can determine a product recommendation for a group of users.)
Step 2A Prong 2: The claim does not recite additional elements; therefore, the judicial exception is not integrated into a practical application.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 27
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 27 depends on. Claim 27 further recites:
"the fixed-size representation is generated so that not decipherable by a third-party observer on the public network to determine a private or confidential information about the group of users." (This limitation is a mental process. A person mentally or with a pen and paper can write down a representation of a group of users in a form, such as a personal shorthand or code, that is not decipherable by another person observing it to determine private or confidential information about the group of users.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the fixed-size representation is uploaded to the remote service over a public network; and" (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): This limitation amounts to merely transmitting data over a generic public network, recited at a high level of generality.)
Step 2B:
"wherein: the fixed-size representation is uploaded to the remote service over a public network; and" (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 28
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 28 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the voice-controlled device encrypts the fixed-size representation before uploading the fixed-size representation to the remote service." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic encryption performed by a generic voice-controlled device. The encrypting is recited at a high level of generality (i.e., as a generic computer function of encrypting data) such that it amounts to no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.)
Step 2B:
"wherein: the voice-controlled device encrypts the fixed-size representation before uploading the fixed-size representation to the remote service." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic encryption performed by a generic voice-controlled device, recited at a high level of generality as a generic computer function of encrypting data. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 29
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 29 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the voice-controlled device uploads the fixed-size representation to the remote service using an encrypted communication protocol." (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): This limitation amounts to merely transmitting data to a remote service using a generic, off the shelf encrypted communication protocol recited at a high level of generality.)
Step 2B:
"wherein: the voice-controlled device uploads the fixed-size representation to the remote service using an encrypted communication protocol." (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 30
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 30 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the encoder model is trained as part of a multitask neural network that uses a plurality of decoders to predict a plurality of attributes of different user groups; and" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training of a generic multitask neural network using a plurality of generic decoders, with no additional details or limitations beyond generic, off the shelf neural network models.)
"the training of the multitask neural network trains the encoder to embed signals of the attributes in the fixed-sized representation." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): This limitation amounts to performing generic training of a generic neural network model, which is merely an instruction to apply the abstract idea using generic machine learning components.)
Step 2B:
"wherein: the encoder model is trained as part of a multitask neural network that uses a plurality of decoders to predict a plurality of attributes of different user groups; and" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training of a generic multitask neural network using a plurality of generic decoders, with no additional details or limitations beyond generic, off the shelf neural network models.)
"the training of the multitask neural network trains the encoder to embed signals of the attributes in the fixed-sized representation." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): This limitation amounts to performing generic training of a generic neural network model, which is merely an instruction to apply the abstract idea using generic machine learning components.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 31
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 31 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the encoder model is trained using labeled training data that indicates ground truth group compositions of groups associated with the training data." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training of a generic encoder model using labeled training data, recited at a high level of generality with no additional details or limitations beyond generic, off the shelf machine learning training.)
Step 2B:
"wherein: the encoder model is trained using labeled training data that indicates ground truth group compositions of groups associated with the training data." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training of a generic encoder model using labeled training data, recited at a high level of generality with no additional details or limitations beyond generic, off the shelf machine learning training.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 32
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 32 depends on. Claim 32 further recites:
"(...) generate a representation of the group that indicates a respective probability or likelihood of individual user classes in the group." (This limitation is a mental process. A person mentally or with a pen and paper can estimate a probability or likelihood that each individual user class (e.g., an adult or a child) is present in a group and write down a representation of the group indicating the respective probability or likelihood of the individual user classes.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the encoder model is trained to (...)" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic training of a generic, off the shelf encoder model as a tool to perform the recited abstract idea.)
Step 2B:
"wherein: the encoder model is trained to (...)" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic training of a generic, off the shelf encoder model as a tool to perform the recited abstract idea.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 33
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 33 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the encoder model comprises a recurrent neural network (RNN) that processes a sequence of words in a voice command to fixed-size representation." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off the shelf recurrent neural network. This limitation merely amounts to applying a generic neural network model as a recurrent neural network to process data, which is merely an instruction to apply the abstract idea using a generic computer component.)
Step 2B:
"wherein: the encoder model comprises a recurrent neural network (RNN) that processes a sequence of words in a voice command to fixed-size representation." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off the shelf recurrent neural network. This limitation merely amounts to applying a generic neural network model as a recurrent neural network to process data, which is merely an instruction to apply the abstract idea using a generic computer component.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 34
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 21 above, which claim 34 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the voice-controlled device is configured to perform further training of the encoder model after the encoder model is deployed to the voice-controlled device, wherein the further training adapts the encoder model to the group of users." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic further training of a generic encoder model on a generic voice-controlled device, with no additional details or limitations beyond generic, off the shelf machine learning training. This limitation amounts to performing a generic retraining step.)
Step 2B:
"wherein: the voice-controlled device is configured to perform further training of the encoder model after the encoder model is deployed to the voice-controlled device, wherein the further training adapts the encoder model to the group of users." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic further training of a generic encoder model on a generic voice-controlled device, with no additional details or limitations beyond generic, off the shelf machine learning training. This limitation amounts to performing a generic retraining step.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 35
Step 1: The claim recites a method; therefore, it is directed to the statutory category of processes.
Step 2A Prong 1: The claim recites the following abstract ideas:
"generating, based on the audio input (...), a fixed-size representation of the group of users, wherein the fixed-size representation indicates a member composition of the group based on user classes detected in the group;" (This limitation is a mental process. A person mentally or with a pen and paper can listen to audio input from a group of users, detect the classes of the users (e.g., whether a speaker is an adult or a child), and write down a representation of the group (e.g., a list of a fixed length) that indicates the member composition of the group based on the user classes the person detected.)
"wherein the (...) uses the fixed-size representation to generate a personalized output for the group of users;" (This limitation is a mental process. A person mentally or with a pen and paper can use a written representation of a group of users to determine a personalized output, such as a personalized suggestion, for the group of users.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"A method, comprising: performing, by a voice-controlled device implemented by one or more computers:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites a generic, off the shelf voice-controlled device implemented by generic computers as a tool to perform the recited abstract ideas.)
"receiving audio input from a group of users over a time period, wherein the audio input includes voice commands associated with different items over the time period;" (Data Gather - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)).)
"generating, based on the audio input and via an encoder model trained using one or more machine learning techniques, a fixed-size representation of the group of users..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "encoder model trained using one or more machine learning techniques" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of "generating, based on the audio input (...), a fixed-size representation of the group of users". The claim denotes a generic encoder model and generic machine learning techniques with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"uploading the fixed-size representation to a remote service (...);" (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): This limitation amounts to merely transmitting data to a remote service, recited at a high level of generality.)
"(...) a remote service that implements a machine learning model, wherein the machine learning model uses the fixed-size representation to generate a personalized output for the group of users;" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "machine learning model" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of using the fixed-size representation to generate a personalized output for the group of users. The claim denotes a generic machine learning model with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"receiving the personalized output from the remote service and generating audio output indicating the personalized output." (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): Receiving the personalized output from the remote service amounts to mere data gathering, and generating audio output indicating the personalized output is the nominal act of outputting or delivering the end result after the core process is complete; therefore, this is interpreted as insignificant pre- and post-solution activity.)
Step 2B:
"A method, comprising: performing, by a voice-controlled device implemented by one or more computers:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites a generic, off the shelf voice-controlled device implemented by generic computers as a tool to perform the recited abstract ideas.)
"receiving audio input from a group of users over a time period, wherein the audio input includes voice commands associated with different items over the time period;" (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
"generating, based on the audio input and via an encoder model trained using one or more machine learning techniques, a fixed-size representation of the group of users..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "encoder model trained using one or more machine learning techniques" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of "generating, based on the audio input (...), a fixed-size representation of the group of users". The claim denotes a generic encoder model and generic machine learning techniques with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"uploading the fixed-size representation to a remote service (...);" (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
"(...) a remote service that implements a machine learning model, wherein the machine learning model uses the fixed-size representation to generate a personalized output for the group of users;" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "machine learning model" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of using the fixed-size representation to generate a personalized output for the group of users. The claim denotes a generic machine learning model with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"receiving the personalized output from the remote service and generating audio output indicating the personalized output." (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 36
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 35 above, which claim 36 depends on. Claim 36 further recites:
"the personalized output indicates a product recommendation to the group of users." (This limitation is a mental process. A person mentally or with a pen and paper can determine a product recommendation for a group of users.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the voice commands indicate interactions with different products by the group of users; and" (Data Gather - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner's Note (EN): The limitation merely specifies the content of the received data, i.e., that the received voice commands indicate interactions with different products.)
Step 2B:
"wherein: the voice commands indicate interactions with different products by the group of users; and" (This falls under well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). See MPEP 2106.05(d)(II).)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 37
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 35 above, which claim 37 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"encrypting, by the voice-controlled device, the fixed-size representation before or during the uploading of the fixed-size representation to the remote service." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic encryption performed by a generic voice-controlled device. The encrypting is recited at a high level of generality (i.e., as a generic computer function of encrypting data) such that it amounts to no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.)
Step 2B:
"encrypting, by the voice-controlled device, the fixed-size representation before or during the uploading of the fixed-size representation to the remote service." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic encryption performed by a generic voice-controlled device, recited at a high level of generality as a generic computer function of encrypting data. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 38
Claim 38 is a method claim that recites identical limitations to system claim 24. Therefore, claim 38 is rejected using the same rationale as claim 24.
Claim 39
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 35 above, which claim 39 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"wherein: the voice-controlled device implements a graphical user interface (GUI); and" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off the shelf graphical user interface performing a generic computer function of providing output. Therefore, it amounts to no more than mere instructions to apply the exception using a generic computer component.)
"the method further comprises generate GUI output via the GUI that indicates the personalized output." (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): Generating GUI output that indicates the personalized output is the nominal act of outputting or delivering the end result after the core process is complete; therefore, this is interpreted as an insignificant post-solution activity.)
Step 2B:
"wherein: the voice-controlled device implements a graphical user interface (GUI); and" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off the shelf graphical user interface performing a generic computer function of providing output. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept.)
"the method further comprises generate GUI output via the GUI that indicates the personalized output." (MPEP 2106.05(d)(II) indicates that merely receiving or transmitting data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 40
Step 1: The claim recites one or more non-transitory computer-accessible storage media; therefore, it is directed to the statutory category of manufacture.
Step 2A Prong 1: The claim recites the following abstract ideas:
"generate, based on the audio input (...), a fixed-size representation of the group of users, wherein the fixed-size representation indicates a member composition of the group based on user classes detected in the group;" (This limitation is a mental process. A person mentally or with a pen and paper can listen to audio input from a group of users, detect the classes of the users (e.g., whether a speaker is an adult or a child), and write down a representation of the group (e.g., a list of a fixed length) that indicates the member composition of the group based on the user classes the person detected.)
"wherein the (...) uses the fixed-size representation to generate a personalized output for the group of users;" (This limitation is a mental process. A person mentally or with a pen and paper can use a written representation of a group of users to determine a personalized output, such as a personalized suggestion, for the group of users.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
"One or more non-transitory computer-accessible storage media storing program instructions that when executed on one or more processors of a voice-controlled device cause the voice-controlled device to:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off the shelf storage media, program instructions, and processors of a generic voice-controlled device as tools to perform the recited abstract ideas.)
"receive audio input from a group of users over a time period, wherein the audio input includes voice commands associated with different items over the time period;" (Data Gather - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)).)
"generate, based on the audio input and via an encoder model trained using one or more machine learning techniques, a fixed-size representation of the group of users..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "encoder model trained using one or more machine learning techniques" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of "generate, based on the audio input (...), a fixed-size representation of the group of users". The claim denotes a generic encoder model and generic machine learning techniques with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"upload the fixed-size representation to a remote service (...);" (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): This limitation amounts to merely transmitting data to a remote service, recited at a high level of generality.)
"(...) a remote service that implements a machine learning model, wherein the machine learning model uses the fixed-size representation to generate a personalized output for the group of users;" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "machine learning model" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of using the fixed-size representation to generate a personalized output for the group of users. The claim denotes a generic machine learning model with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"receive the personalized output from the remote service and generate audio output indicating the personalized output." (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- Examiner's Note (EN): Receiving the personalized output from the remote service amounts to mere data gathering, and generating audio output indicating the personalized output is the nominal act of outputting or delivering the end result after the core process is complete; therefore, this is interpreted as insignificant pre- and post-solution activity.)
Step 2B:
"One or more non-transitory computer-accessible storage media storing program instructions that when executed on one or more processors of a voice-controlled device cause the voice-controlled device to:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off the shelf storage media, program instructions, and processors of a generic voice-controlled device as tools to perform the recited abstract ideas.)
"receive audio input from a group of users over a time period, wherein the audio input includes voice commands associated with different items over the time period;" (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
"generate, based on the audio input and via an encoder model trained using one or more machine learning techniques, a fixed-size representation of the group of users..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "encoder model trained using one or more machine learning techniques" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of "generate, based on the audio input (...), a fixed-size representation of the group of users". The claim denotes a generic encoder model and generic machine learning techniques with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"upload the fixed-size representation to a remote service (...);" (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
"(...) a remote service that implements a machine learning model, wherein the machine learning model uses the fixed-size representation to generate a personalized output for the group of users;" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "machine learning model" in the claim is interpreted as applying a generic, off the shelf machine learning model on the abstract idea of using the fixed-size representation to generate a personalized output for the group of users. The claim denotes a generic machine learning model with no additional details or limitations beyond a generic, off the shelf machine learning model.)
"receive the personalized output from the remote service and generate audio output indicating the personalized output." (MPEP 2106.05(d)(II) indicates that receiving or transmitting data over a network is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Examiner’s Note: Some rejections will include an Examiner’s Note (labeled ‘EN’) to provide additional context or rationale explaining the basis for the rejection.
Claims 21-27, 31-33, 35, 36, and 38-40 are rejected under 35 U.S.C. 103 as being unpatentable over Mixter (US 2017/0329573 A1) in view of Sullivan et al. (US 2017/0064358 A1), and further in view of Vinyasa et al. (US 2015/0356401 A1).
Regarding claim 21, Mixter teaches:
A system, comprising: one or more computers that implement a voice-controlled device, configured to: (Para 53, "Examples of the voice assistant client device 104 include, but are not limited to, a mobile phone, a tablet computer, a laptop computer, a desktop computer, a wireless speaker... a voice command device... a television, a soundbar, a casting device... an in-vehicle system, and a wearable personal device. The voice assistant client device 104... or casting device 106..., typically, includes one or more processing units (CPUs) 202, one or more network interfaces 204, memory 206, and one or more communication buses 208... The voice assistant client device 104 or casting device 106 includes one or more input devices 210 that facilitate user input, including an audio input device 108 or 132 (e.g., a voice-command input unit or microphone)..." - EN: this denotes a computer-implemented voice-command device having processors, memory, network interfaces, and audio-input hardware.)
receive audio input from (Para 117, "The device receives (502) a verbal input at the device. The client device 104/casting device 106 captures a verbal input (e.g., voice input) uttered by a user." Para 118, "The processing may include hotword detection, conversion to textual data, and identification of words and phrases corresponding to commands, requests, and/or parameters provided by the user." Para 26, "Voice commands may be used to initiate playback and control of music, radio, podcasts, news, and other audio media through voice. For example, a user can utter voice commands (e.g., "play jazz music," "play 107.5 FM," "skip to next song," "play 'Serial'") to play or control various types of audio media." Para 27, "The user can also specify generic categories or specific content in the command, and the appropriate media is played in accordance with the specified category or content in the command." Para 25, "an objective of a voice assistant is to provide users a personalized voice interface available across a variety of devices and enabling a wide variety of use cases, providing consistent experience throughout a user's day." Para 63, "Usage history 234 for storing information associated with the operation and usage of the voice assistant module 136 (e.g., logs), such as commands and requests received, responses to the commands and requests, operations performed in response to commands and requests, and so on" - EN: this denotes receiving voice audio, recognizing commands directed to different media/content items, and maintaining those commands as a usage history over time.)
generate, based on the audio input and (...), a (Para 118, "For example, the processing may include encoding the verbal input audio for transmission to server 114, or preparing the captured raw audio of the verbal input for transmission to server 114." Para 48, "The voice assistant module/library 136 detects (e.g., receives) verbal input picked up (e.g., captured) by the audio input device 108/132, processes the verbal input (e.g., to detect hotwords), and transmits the processed verbal input or an encoding of the processed verbal input to the server 114." - EN: this denotes generating an encoded, machine-readable representation based on the received audio input.)
upload the (Para 119, "The client device 104/casting device 106 transmits the determined request to the remote system (e.g., server 114)... In some implementations, the client device 104/casting device 106 transmits the verbal input (e.g., as an encoded audio, as raw audio data) to the server 114..." Para 48, "The casting device 106 and the voice assistant client device 104 are communicatively coupled to a server system 114 through one or more communicative networks 112 (e.g., local area networks, wide area networks, the Internet)... The server 114 receives the processed verbal input or an encoding thereof, and processes the received verbal input to determine the appropriate response to the verbal input... As part the processing, the server 114 may communicate with one or more content or information sources 138 to obtain content or information, or references to such, for the response. In some implementations, the content or information sources 138 include, for example, search engines, databases, information associated with the user's account (e.g., calendar, task list, email), websites, and media streaming services." Para 25, "an objective of a voice assistant is to provide users a personalized voice interface available across a variety of devices and enabling a wide variety of use cases" - EN: this denotes uploading the generated encoding or representation from the voice-controlled device to a remote network service, and the remote server using the received encoding and user-account information to determine a personalized response.)
receive the personalized output from the remote service and generate audio output indicating the personalized output. (Para 120, "The remote system (e.g., the server 114) determines and generates a response to the request, and transmits the response to the client device 104/casting device 106." Para 121, "if the response is a command to the device to output certain information by audio, the client device 104/casting device 106 retrieves the information, converts the information to speech audio output, and outputs the speech audio through the speaker." Para 48, "The server 114 sends the response to the casting device 106 or voice assistant client device 104, where the content or information is output (e.g., output through audio output device 110/134) and/or a function is performed." - EN: this denotes receiving the server-generated response and rendering that response as speech or other audio through the device speaker.)
Mixter does not explicitly teach:
"a group of users";
"via an encoder model trained using one or more machine learning techniques";
"a fixed-size representation of the group of users";
"wherein the fixed-size representation indicates a member composition of the group based on user classes detected in the group";
"that implements a machine learning model, wherein the machine learning model uses the fixed-size representation"; and
"for the group of users," i.e. the personalized output is directed to the group collectively.
However, Sullivan teaches:
"a group of users" (Para 22, "As used herein, "consumption data" refers to information pertaining to media exposure events presented via a media presentation device... of a household (e.g., a panelist household) and associated with a person and/or a group of persons of the household (e.g., panelist(s), member(s) of the panelist household)." Para 53, "For example, the tuning event database 120 stores a household (e.g., the household 102), a channel, and a time associated with each tuning event of the tuning data 108." Para 94, "For example, tuning data and a corresponding demographic distribution is obtained for tuning events of the household 102 associated with NBC between 6:00 P.M. and 6:15 P.M., NBC between 6:15 P.M. and 6:30 P.M., NBCSports between 7:00 P.M. and 7:15 P.M... and NBC between 10:30 P.M. and 10:45 P.M." - EN: this denotes collecting a sequence of item-related interaction events attributable to a household or other group of persons over a defined period of time. When Sullivan's household-event teaching is applied to Mixter's stored voice-command history, the combined system receives voice-command events associated with a group of users over time.)
"a fixed-size representation of the group of users" (Para 70, "The score calculator 206 of the illustrated example constructs a score vector based on the calculated scores. Each element of the score vector represents the calculated score of a respective demographic constraint. In an example score vector, a first element represents a score associated with males (e.g., 0.59), a second element represents a score associated with females (e.g., 1.43), a third element represents a score associated with young adults (e.g., 0.56), a fourth element represents a score associated with middle-aged adults (e.g., 1.34), and a fifth element represents a score associated with seniors (e.g., 1.06)." Para 71, "For example, the score calculator 206 calculates a score vector for the household 102 and calculates another score vector for another non-panelist household." - EN: this denotes a household representation having a predetermined set of vector elements, with respective elements corresponding to defined user or demographic classes. Under the broadest reasonable interpretation (BRI), Sullivan's score vector having predetermined male, female, young-adult, middle-aged, and senior fields constitutes a fixed-size representation of the household/group.)
"wherein the fixed-size representation indicates a member composition of the group based on user classes detected in the group" (Para 28, "a demographic score (e.g., a ratio) is calculated for the demographic constraints of interest... For example, a higher score for a particular demographic marginal corresponds to a higher likelihood that the non-panelist household includes a member of that particular demographic marginal." Para 25, "to predict or estimate a household characteristic (e.g., a demographic composition such as a number of household members and demographics of the household members... etc.) of a non-panelist household..." Para 87, "If the value obtained from the decision tree satisfies the threshold (e.g., is greater than or equal to the threshold value), the household estimator 210 identifies that the non-panelist household includes the corresponding household feature. For example, the household estimator 210 identifies that the household 102 includes a female..." - EN: each element of the score vector (Para 70) is a score for one user class, and a higher score means the group more likely includes a member of that class (Para 28). The score vector itself therefore indicates the member composition of the group, i.e., which user classes are present and thus the makeup of the household (Para 25). Identifying the classes present in the group from the vector (Para 87) constitutes detecting user classes in the group.)
"that implements a machine learning model, wherein the machine learning model uses the fixed-size representation" (Para 51, "the STB 110 communicates the collected tuning data 108 to the AME 104 via the network 106 (e.g., the Internet, a local area network, a wide area network, a cellular network, etc.)" Para 52, "The AME 104 of the illustrated example utilizes the collected tuning data 108 to estimate household characteristics of the household 102... As illustrated in FIG. 1, the AME 104 includes a tuning event database 120, a panelist database 122, a distribution calculator 124, and a characteristic estimator 126." Para 72, "The decision tree ensembles are subsequently applied to data associated with the non-panelist households (e.g., the score vector of the household 102) by the household estimator 210 of the illustrated example to identify household characteristics of the non-panelist households (e.g., number of household members, demographics of the respective household members, etc.)." Para 83, "Subsequently, the household estimator 210 applies the decision tree ensembles constructed by the decision tree trainer 208 to the non-panelist feature matrix." Para 84, "By applying the decision trees of the respective decision tree ensembles to the non-panelist feature matrix, the household estimator 210 obtains values associated with likelihoods that the non-panelist households (e.g., the household 102) include members satisfying the corresponding household features of interest." Para 86, "alternative examples of the household estimator 210 utilize other forms of machine learning (e.g., neural networks, support vector machines, clustering, Bayesian networks, etc.) to estimate the demographics of the household 102." - EN: this denotes a network-accessible service implementing a trained machine-learning classifier that consumes a class-based household vector or feature matrix and produces household-member classifications. In the combination, Mixter's server 114 is the remote service that implements Sullivan's trained model and applies it to the uploaded representation.)
"for the group of users" (Para 26, "As used herein, a "household characteristic" refers to a characteristic of a household and/or a characteristic of a member of the household. Example household characteristics include a number of household members, demographics of the household members, a number of television sets within the household, locations of the respective televisions within the household, etc.)." - EN: Sullivan's estimated characteristics pertain to the household group and its members collectively. In the combination, Mixter's personalized server response is generated for the group of users of the shared voice-controlled device based on the member composition detected by Sullivan's model, and the personalized output is therefore for the group of users rather than for a single user.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the voice-controlled device of Mixter, which captures voice commands directed to different media items from its users, logs those commands as a usage history, and obtains personalized responses from a remote server, with the fixed-size, per-class group representation and trained machine-learning household-composition estimator of Sullivan, applied to Mixter's stored voice-command history. The motivation for doing so would be to automatically determine, from usage data the device already collects, the number of users sharing the device and the user classes (e.g., ages and genders) of those users, without requiring the users to enroll, self-report demographics, or operate special metering equipment, so that the remote server's responses are personalized to the actual group using the device, furthering Mixter's stated objective of "provid[ing] users a personalized voice interface" (Mixter Para 25). Sullivan discloses this benefit: "methods and apparatus disclosed herein enable AMEs (or any other entity) to associate the tuning data of the non-panelist households with demographics data of its household members" (Sullivan Para 25), whereas collecting such demographics directly requires enlisting panelists, which "can be a difficult and costly process" in which members "must diligently perform specific tasks" (Sullivan Para 23). In the combination, the representation is generated at Mixter’s device, which already processes and encodes the verbal input before transmission to the server, and the device uploads the generated representation in place of the verbal input itself; uploading the per-class representation rather than the users’ raw voice data is itself beneficial, because users are more willing to share usage data that carries no personalized information: "Many households are willing to provide tuning data via a STB, because personalized information is not collected by the STB and repeated actions are not required of the household members" (Sullivan Para 24).
Further, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the personalized output generation at the remote server of Mixter with the trained machine learning model of Sullivan, such that Mixter's server generates its personalized response using the machine learning model's processing of the uploaded fixed-size representation. The motivation for doing so would be that a machine learning model, trained on groups whose member composition is already known, can then automatically and accurately determine the composition of new groups from their uploaded representations alone, allowing the server to personalize its output to the users actually sharing the device without requiring the users to identify themselves. Sullivan discloses this benefit: "the decision tree trainer 208 tests the constructed decision tree ensembles on the testing group panelist households to determine whether the decision tree ensembles are able to be accurately applied to households on which they were not trained" (Sullivan Para 80); see also "[b]ased on the scores of the decision tree ensembles, the AME is able to estimate household characteristics of the non-panelist household (e.g., a number of members of the non-panelist household, demographics of each of the members...)" (Sullivan Para 30).
Mixter in view of Sullivan does not explicitly teach:
"via an encoder model trained using one or more machine learning techniques"
However, Vinyals teaches:
"via an encoder model trained using one or more machine learning techniques" (Para 17, "The encoder LSTM neural network 110 is a recurrent neural network that receives an input sequence and generates an alternative representation from the input sequence." Para 18, "The encoder LSTM neural network 110 has been configured, e.g., through training, to process each input in a given input sequence to generate the alternative representation of the input sequence in accordance with a set of parameters." Para 51, "In order to configure the encoder LSTM neural network and the decoder LSTM neural network, the system can train the networks using conventional machine learning training techniques, e.g., using Stochastic Gradient Descent." - EN: this denotes an encoder neural network, trained using machine learning training techniques such as stochastic gradient descent, that converts a received input sequence into a representation of the sequence, which reads on "an encoder model trained using one or more machine learning techniques." In the combination, Vinyals's trained encoder is the encoder model at Mixter's voice-controlled device that generates the group representation of Mixter in view of Sullivan from the received voice-command sequence, which Mixter converts to textual data comprising words and phrases.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the voice-controlled device of Mixter in view of Sullivan, which generates at the device a fixed-size, class-based representation of the group of users from the voice commands the device converts to words and phrases (Mixter Para 118; Sullivan Para 70), with the trained encoder neural network of Vinyals (Vinyals Paras 17-18, 51). The motivation for doing so would be to generate the group representation with a single trained model that accepts voice-command sequences of any length while always producing a representation of the same fixed size, ensuring that the uploaded representation can be consumed by the remote service's machine learning model regardless of how many commands were received over the time period. Vinyals discloses this benefit: "the alternative representation of the input sequence is a fixed-length representation, i.e., the number of elements in the alternative representation is fixed and is not dependent on the number of inputs in the input sequence" (Vinyals Para 28).
Regarding claim 22,
Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Mixter further teaches:
wherein the voice-controlled device comprises a smartphone or a television (Para 53, "Examples of the voice assistant client device 104 include, but are not limited to, a mobile phone, a tablet computer, a laptop computer, a desktop computer, a wireless speaker..., a voice command device..., a television, a soundbar, a casting device..." Para 41, "The voice assistant client device 104 (e.g., a smartphone, a laptop or desktop computer, a tablet computer, a voice command device...)")
Regarding claim 23,
Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Mixter further teaches:
wherein the voice-controlled device comprises a vehicle-based computer (Para 53, "Examples of the voice assistant client device 104 include... an in-vehicle system, and a wearable personal device." - EN: the in-vehicle system is a vehicle-based computer.)
Regarding claim 24, Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Mixter further teaches:
the (this limitation is split between the references: Mixter teaches the voice-controlled device is associated with a user account (Para 33, "The configurations or setups may include specifying a device location, association with a user account..."; Para 48, "information associated with the user's account (e.g., calendar, task list, email)").)
Mixter does not explicitly teach:
"the group of users are members of a family account,"; and
"wherein the user classes indicate different ages and genders of the members"
However, Sullivan further teaches:
"the group of users are members of a family account," (Sullivan teaches that the group of users is a household and its members (Para 22, "a group of persons of the household"; Para 26, "a number of household members, demographics of the household members"). In the combination, the members of Sullivan's household share Mixter's device and its associated user account, and a common account shared by the members of one household is a family account under the broadest reasonable interpretation.)
wherein the user classes indicate different ages and genders of the members (Para 70, "a first element represents a score associated with males (e.g., 0.59), a second element represents a score associated with females (e.g., 1.43), a third element represents a score associated with young adults (e.g., 0.56), a fourth element represents a score associated with middle-aged adults (e.g., 1.34), and a fifth element represents a score associated with seniors (e.g., 1.06)." - EN: the user classes of claim 21 are Sullivan's demographic constraints, which are genders (male, female) and ages (young adult, middle-aged, senior).).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the voice-controlled device of Mixter, which is shared by its users and associated with a user account (Mixter Paras 33, 48), with the household group of members and the age and gender user classes of Sullivan (Sullivan Paras 22, 26, 70). The motivation for doing so would be the same as that set forth for claim 21: to automatically determine, from usage data the device already collects, the makeup of the family sharing the device and its account, i.e., the ages and genders of its members, without requiring the members to enroll or self-report their demographics, so that the remote server's responses are personalized to the actual family using the device. Sullivan discloses this benefit: "methods and apparatus disclosed herein enable AMEs (or any other entity) to associate the tuning data of the non-panelist households with demographics data of its household members" (Sullivan Para 25), such demographics data being, e.g., "gender and age" (Sullivan Para 2).
Regarding claim 25, Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Mixter further teaches:
wherein the voice commands indicate interactions with different products including two or more of: searching for a product, viewing the product, purchasing the product, returning the product, and providing feedback on the product (Para 26, "such commands may be used to play audio media from a variety of sources, such as online streaming of terrestrial radio stations, music subscription services..." Para 33, "linking to and prioritizing media services (e.g., video streaming services, music streaming services)" Para 27, "The user can also specify generic categories or specific content in the command, and the appropriate media is played in accordance with the specified category or content in the command." Para 28, "questions and answers powered by a search engine (e.g., search queries)" Para 121, "if the response is a command to the device to play media content, the client device 104/casting device 106 retrieves the media content and plays the media content." - EN: the media items of Mixter's subscription and streaming services are products under the broadest reasonable interpretation. A command specifying content or a category for the system to locate (Para 27), or a voice search query (Para 28), is searching for a product; a command causing a video media item to be played on the device's display (Paras 121, 33; Para 53, display 144) is viewing the product. This satisfies "two or more of" the recited interactions.)
Regarding claim 26,
Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 25.
Mixter further teaches:
wherein the personalized output indicates a product recommendation to the (Para 27, "The user can also specify generic categories or specific content in the command, and the appropriate media is played in accordance with the specified category or content in the command." Para 48, "The server 114 receives the processed verbal input or an encoding thereof, and processes the received verbal input to determine the appropriate response to the verbal input. The appropriate response may be content, information, or instructions or commands or metadata to the casting device 106 or voice assistant client device 104 to perform a function or operation." - EN: the personalized output is the server's response (see mapping for claim 21), and the products are the media items (see mapping for claim 25). For a generic command like "play jazz music" (Para 26), the server, not the user, picks the specific media item (product); the response identifying that pick is a product recommendation.)
Mixter does not explicitly teach:
"group of users"
However, Sullivan further teaches:
"group of users" (As stated previously in claim 24, Sullivan discloses a group of members of a household: Para 22, "a group of persons of the household"; Para 26, "a number of household members, demographics of the household members")
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the product recommendation of Mixter, in which the server's response picks the specific media item to be played for the user's generic command (Mixter Paras 26, 27, 48), with the group of users of Sullivan (Sullivan Paras 22, 26). The motivation for doing so would be the same as that set forth for claim 21: media played on a shared voice-controlled device is consumed by the household as a group, so directing the recommendation to the group of users, rather than to a single user, personalizes the recommendation to the audience that will actually consume the recommended item. Sullivan discloses this benefit: media exposure events are "associated with a person and/or a group of persons of the household" (Sullivan Para 22), and the value of collected usage data lies in "measuring the composition and size of audiences consuming media" (Sullivan Para 25).
Regarding claim 27,
Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Mixter further teaches:
wherein the (Para 48, "communicatively coupled to a server system 114 through one or more communicative networks 112 (e.g., local area networks, wide area networks, the Internet)" - EN: the Internet is a public network.).
Mixter does not explicitly teach:
“fixed-size”
“the fixed-size representation is generated so that not decipherable by a third-party observer on the public network to determine a private or confidential information about the group of users.”
However, Sullivan teaches:
the fixed-size representation is generated so that not decipherable by a third-party observer on the public network to determine a private or confidential information about the group of users. (Para 70, "The score calculator 206 of the illustrated example constructs a score vector based on the calculated scores. Each element of the score vector represents the calculated score of a respective demographic constraint. In an example score vector, a first element represents a score associated with males (e.g., 0.59), a second element represents a score associated with females (e.g., 1.43), a third element represents a score associated with young adults (e.g., 0.56)...")
Additional Examiner's Note: Sullivan's score vector as generated is a set of numeric scores, one per demographic constraint. The scores are ratio values computed from the household's collected events and carry no names, no utterances, and no record of the underlying events themselves. An observer of the vector on the network, lacking Sullivan's constraint definitions and trained classifier, obtains only unlabeled numbers, and the members' identities and the raw record of what the members said and did cannot be determined from those numbers. The vector as generated is therefore a representation from which a third-party observer on the public network cannot determine the private or confidential information about the group, and in the combination the vector is what the device uploads in place of the verbal input itself. See paragraphs 56, 70, and 84 of Sullivan.
the device's upload of Sullivan's score vector as the fixed-size representation, in place of the verbal input itself, is the combination already made and motivated in the rejection of claim 21, and the recited non-decipherability is a property of the score vector as generated in that combination, as explained in the Examiner's Note above. Moreover, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art that uploading such a representation, rather than the users' raw voice data, is itself beneficial, because users are more willing to share usage data that carries no personalized information. Sullivan discloses this benefit: "Many households are willing to provide tuning data via a STB, because personalized information is not collected by the STB and repeated actions are not required of the household members" (Sullivan Para 24).
Regarding claim 31,
Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Sullivan further teaches:
(Para 76, "the example decision tree trainer 208 constructs truth vectors for the respective household features of interest of the training group and the testing group based on known household characteristics (e.g., demographic characteristics) of the panelist households... the decision tree trainer 208 constructs a first element to indicate a known number of male members of a first panelist household, a second element to indicate a known number of male members of a second panelist household..." - EN: the panelist households are the training groups, and the truth vectors are labels giving each training household's actual (ground truth) member composition.).
Mixter in view of Sullivan does not explicitly disclose:“The encoder model is trained”
However, Vinyals teaches that
the encoder model is trained (Para 51, "the system can train the networks using conventional machine learning training techniques, e.g., using Stochastic Gradient Descent."). This limitation is split between the references: Vinyals supplies training the encoder, and Sullivan supplies the labeled training data; in the combination, the encoder is trained using Sullivan's panelist-household data labeled with the households' ground-truth compositions. Motivation for the combination is provided below.
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the trained encoder model of Vinyals, as combined for claim 21, with the labeled training data of Sullivan, whose truth vectors give each training household's known, ground-truth member composition (Sullivan Para 76), such that the encoder is trained using that labeled data. The motivation for doing so would be the same as that set forth for claim 21: a model trained on groups whose true compositions are known can be tested for accuracy before it is applied to new groups whose compositions are unknown. Sullivan discloses this benefit: "the decision tree trainer 208 tests the constructed decision tree ensembles on the testing group panelist households to determine whether the decision tree ensembles are able to be accurately applied to households on which they were not trained" (Sullivan Para 80).
Regarding claim 32,
Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Sullivan further teaches:
a representation of the group that indicates a respective probability or likelihood of individual user classes in the group (Para 28, "a higher score for a particular demographic marginal corresponds to a higher likelihood that the non-panelist household includes a member of that particular demographic marginal." Para 70 (per-class score vector) - EN: each element of the score vector is a likelihood that the corresponding user class is present in the group.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the representation generated and uploaded by the voice-controlled device of Mixter with the per-class likelihood score vector of Sullivan. The motivation for doing so would be the same as that set forth for claim 21: a representation whose elements indicate the likelihood of each individual user class tells the remote service which user classes are present in the group, which is what allows the service to determine the group's composition and personalize its output accordingly. Sullivan discloses this benefit: "a higher score for a particular demographic marginal corresponds to a higher likelihood that the non-panelist household includes a member of that particular demographic marginal" (Sullivan Para 28); "[b]ased on the scores of the decision tree ensembles, the AME is able to estimate household characteristics of the non-panelist household" (Sullivan Para 30).
Mixter in view of Sullivan does not explicitly disclose:“the encoder model is trained to generate a representation”
However, Vinyals teaches that
“the encoder model is trained to generate a representation” (Para 18, "The encoder LSTM neural network 110 has been configured, e.g., through training, to process each input in a given input sequence to generate the alternative representation of the input sequence in accordance with a set of parameters."). – EN: In the combination, the encoder is trained to generate Sullivan's per-class likelihood vector. Motivation for the combination is provided below.
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the per-class likelihood representation of Mixter in view of Sullivan, above, with the encoder model of Vinyals, which is trained to generate a representation from an input sequence. The motivation for doing so would be the same as that set forth for claim 21: a single trained encoder generates the representation directly from the device's voice-command sequence, whatever the sequence's length, while always producing the fixed-size output that the remote service's machine learning model expects. Vinyals discloses this benefit: "the alternative representation of the input sequence is a fixed-length representation, i.e., the number of elements in the alternative representation is fixed and is not dependent on the number of inputs in the input sequence" (Vinyals Para 28).
Regarding claim 33, Mixter in view of Sullivan, further in view of Vinyals, teaches the system of claim 21.
Vinyals further teaches:
wherein the encoder model comprises a recurrent neural network (RNN) that processes a sequence of words in a voice command to fixed-size representation (Para 17, "The encoder LSTM neural network 110 is a recurrent neural network that receives an input sequence and generates an alternative representation from the input sequence." Para 13, "if the input sequence 102 is a sequence of words in an original language, e.g., a sentence or phrase..." Para 28, "the alternative representation of the input sequence is a fixed-length representation" - EN: the encoder is an LSTM recurrent neural network that processes a sequence of words into a fixed-length, i.e., fixed-size, representation. Vinyals does not describe the words as being words in a voice command; as set forth for claim 21, in the combination the encoder's input sequence is the sequence of words of the voice commands received by Mixter's device, which Mixter converts to textual data and identifies as words and phrases (Mixter Para 118). The encoder of the combination therefore processes a sequence of words in a voice command to a fixed-size representation.).
The encoder model as combined for claim 21 is Vinyals's encoder LSTM neural network, which is itself a recurrent neural network that processes a sequence of words into a fixed-length representation, so the motivation set forth for claim 21 applies equally here. In any event, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to implement the encoder model of Mixter in view of Sullivan, further in view of Vinyals, as the recurrent neural network of Vinyals. The motivation for doing so would be that a recurrent neural network generates the representation from its hidden state, so a single network can process word sequences of any length while always producing a representation of the same fixed size that the remote service's machine learning model expects. Vinyals discloses this benefit: "[b]ecause the system generates the alternative representation from the hidden state of the encoder LSTM neural network, the alternative representation of the input sequence is a fixed-length representation" (Vinyals Para 28).
Regarding claim 35, the claim recites a method performed by a voice-controlled device implemented by one or more computers, with steps (receiving, generating, uploading, and receiving and generating audio output) corresponding to the functions recited in claim 21. Mixter teaches performing these operations as a method at the device (Para 116, "FIG. 5 illustrates a flow diagram of a method 500 for processing verbal inputs on a device, in accordance with some implementations. The method 500 is performed at an electronic device (e.g., voice assistant client device 104, casting device 106) with an audio input system (e.g., audio input device 108/132), one or more processors (e.g., processing unit(s) 202), and memory (e.g., memory 206) storing one or more programs for execution by the one or more processors."). Claim 35 is therefore rejected over Mixter in view of Sullivan, further in view of Vinyals, for the same reasons and with the same motivations set forth for claim 21.
Regarding claim 36, Mixter in view of Sullivan, further in view of Vinyals, teaches the method of claim 35.
Mixter further teaches:
wherein: the voice commands indicate interactions with different products by (Para 26, "Voice commands may be used to initiate playback and control of music, radio, podcasts, news, and other audio media through voice. For example, a user can utter voice commands (e.g., "play jazz music," "play 107.5 FM," "skip to next song," "play ‘Serial’") to play or control various types of audio media." Para 27, "The user can also specify generic categories or specific content in the command, and the appropriate media is played in accordance with the specified category or content in the command." Para 28, "questions and answers powered by a search engine (e.g., search queries)" Para 121, "if the response is a command to the device to play media content, the client device 104/casting device 106 retrieves the media content and plays the media content." - EN: as set forth for claim 25, the media items of Mixter's subscription and streaming services are the claimed products under the broadest reasonable interpretation. Voice commands to search for, play, and control the different media items (Paras 26, 27, 28, 121) are voice commands indicating interactions with those different products by the device's users (see the mappings for claims 21 and 25).)
the personalized output indicates a product recommendation to the (Para 27, "The user can also specify generic categories or specific content in the command, and the appropriate media is played in accordance with the specified category or content in the command." Para 48, "The server 114 receives the processed verbal input or an encoding thereof, and processes the received verbal input to determine the appropriate response to the verbal input. The appropriate response may be content, information, or instructions or commands or metadata to the casting device 106 or voice assistant client device 104 to perform a function or operation." - EN: as set forth for claim 26, the personalized output is the server’s response (see the mapping for claim 21) and the products are the media items (see the mapping for claim 25); for a generic command like "play jazz music" (Para 26), the server, not the user, picks the specific media item (product) to be played, and the response identifying that pick is a product recommendation.)
Mixter does not explicitly teach:
"by the group of users"; and
"to the group of users"
However, Sullivan further teaches:
"by the group of users" and "to the group of users" (Para 22, "As used herein, "consumption data" refers to information pertaining to media exposure events presented via a media presentation device... of a household (e.g., a panelist household) and associated with a person and/or a group of persons of the household (e.g., panelist(s), member(s) of the panelist household)." Para 26, "Example household characteristics include a number of household members, demographics of the household members..." - EN: Sullivan attributes the media exposure/interaction events of the household to the group of persons of the household. In the combination (see claims 21 and 35), the voice commands received by Mixter's shared device are the interaction events of the members of Sullivan's household, i.e., the interactions with the different products are by the group of users; and, as set forth for claim 26, the personalized output of the combined system is directed to the users of the shared device as a group, i.e., the product recommendation is to the group of users.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the voice-command product interactions and server product recommendation of Mixter with the group of users of Sullivan. The motivation for doing so would be the same as that set forth for claims 21 and 26: the voice commands received by a shared voice-controlled device originate from the household members that share the device, and media played on the shared device is consumed by the household as a group, so attributing the product interactions to the group of users and directing the product recommendation to the group of users personalizes the recommendation to the audience that will actually consume the recommended item. Sullivan discloses this benefit: media exposure events are "associated with a person and/or a group of persons of the household" (Sullivan Para 22), and the value of collected usage data lies in "measuring the composition and size of audiences consuming media" (Sullivan Para 25).
Regarding claim 38, the claim recites, in method form, the limitations of claim 24; therefore it is rejected for the same reasons set forth for claim 24 above.
Regarding claim 39,
Mixter in view of Sullivan, further in view of Vinyals, teaches the method of claim 35.
Mixter further teaches:
wherein: the voice-controlled device implements a graphical user interface (GUI); and the method further comprises generate GUI output via the GUI that indicates the personalized output (Para 53, "The voice assistant client device 104 or casting device 106 also includes one or more output devices 212, including an audio output device 110 or 134..., and optionally one or more visual displays (e.g., display 144)... that enable presentation of user interfaces and display content and information" and "other input devices such as a keyboard, a mouse, a touch screen display, a touch-sensitive input pad..." Para 121, "if the response is a command to the device to play media content, the client device 104/casting device 106 retrieves the media content and plays the media content." - EN: a device that presents user interfaces on a visual display, including a touch screen display, implements a GUI. Presenting the responsive content on that display (e.g., media content output via display 144) is GUI output indicating the personalized output.)
Regarding claim 40, the claim recites one or more non-transitory computer-accessible storage media storing program instructions that when executed on one or more processors of a voice-controlled device cause the voice-controlled device to perform operations corresponding to the functions recited in claim 21. Mixter teaches this form (Para 54, "Memory 206, or alternatively the non-volatile memory within memory 206, includes a non-transitory computer readable storage medium. In some implementations, memory 206, or the non-transitory computer readable storage medium of memory 206, stores the following programs, modules, and data structures..." Para 116, "a non-transitory computer readable storage medium stores one or more programs, the one or more programs including instructions which, when executed by an electronic device with an audio input system (e.g., audio input device 108/132)... the electronic device to perform the method 500."). Claim 40 is therefore rejected over Mixter in view of Sullivan, further in view of Vinyals, for the same reasons and with the same motivations set forth for claim 21.
Claims 28, 29, and 37 are rejected under 35 U.S.C. 103 as being unpatentable over Mixter (US 2017/0329573 A1) in view of Sullivan et al. (US 2017/0064358 A1) and Vinyals et al. (US 2015/0356401 A1), and further in view of Dadu et al. (US 2014/0093083 A1), hereinafter "Dadu".
Regarding claim 28, Mixter in view of Sullivan, further in view of Vinyals, further teaches:
"wherein: the voice-controlled device (as previously mapped in claim 21, the combination teaches the voice-controlled device uploading the fixed-size representation to the remote service - Mixter Para 119, "The client device 104/casting device 106 transmits the determined request to the remote system (e.g., server 114)... In some implementations, the client device 104/casting device 106 transmits the verbal input (e.g., as an encoded audio, as raw audio data) to the server 114..."; Mixter Para 48; the representation being fixed-size per Sullivan Para 70 and Vinyals Para 28. See the rejection of claim 21 above.)
Mixter in view of Sullivan, further in view of Vinyals, does not explicitly teach:
"encrypts [the fixed-size representation before]" the uploading, i.e., encrypting the data before sending to the server.
However, Dadu teaches:
encrypts [the fixed-size representation before] (Para 35, "In block 320, the security engine 118 of the client computing device 102 encrypts the audio response data using the shared cryptographic key. Thereafter, in block 322, the client computing device 102 transfers the encrypted audio response data to the server 106 using the application 202." - EN: Dadu's client device captures the user's voice response, encrypts it, and only then transmits it to the server, i.e., the device encrypts its voice-derived data before uploading. In the combination, the voice-controlled device likewise encrypts the fixed-size representation, the voice-derived data it uploads per claim 21, before uploading it to the remote service.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the upload of the fixed-size representation from the voice-controlled device to the remote service of Mixter in view of Sullivan, further in view of Vinyals, with the client-side encryption before transmission to server of Dadu. The motivation for doing so would be to protect the users' voice-derived data from interception or tampering while it travels over the network to the remote service, securing the voice interaction from end to end. Dadu discloses this benefit: "the client computing device 102 may securely transfer audio to the server 106 via a network 104. To facilitate such secure transfer, the client computing device 102 and the server 106 exchange cryptographic keys and establish a secure connection" (Para 14).
Regarding claim 29, Mixter in view of Sullivan, further in view of Vinyals, further teaches:
"wherein: the voice-controlled device uploads the fixed-size representation to the remote service (as previously mapped in claim 21, the combination teaches the voice-controlled device uploading the fixed-size representation to the remote service – Mixter Para 119, "The client device 104/casting device 106 transmits the determined request to the remote system (e.g., server 114)... In some implementations, the client device 104/casting device 106 transmits the verbal input (e.g., as an encoded audio, as raw audio data) to the server 114..."; Mixter Para 48; the representation being fixed-size per Sullivan Para 70 and Vinyals Para 28. See the rejection of claim 21 above.)
Mixter in view of Sullivan, further in view of Vinyals, does not explicitly teach: "using an encrypted communication protocol."
However, Dadu teaches:
[uploading data from a device to a server] “using an encrypted communication protocol" (Para 14, "the client computing device 102 and the server 106 exchange cryptographic keys and establish a secure connection." Para 32, "the client computing device 102 and the server 106 may use the cryptographic symmetric key as a shared key for encryption and decryption. In other embodiments, the shared cryptographic symmetric key may be used to derive additional symmetric keys for encrypting and signing subsequent audio data packets transmitted between the client computing device 102 and the server 106." - EN: Dadu's device and server exchange keys, establish a secure connection, and encrypt the data packets sent between them under a shared key, thus communicating by an encrypted communication protocol. In the combination, the device uploads the fixed-size representation to the remote service using that protocol. See motivation to combine below.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the upload of the fixed-size representation from the voice-controlled device to the remote service of Mixter in view of Sullivan, further in view of Vinyals, with the client-side encryption and secure client-server connection of Dadu (Paras 14, 32, 35). The motivation for doing so would be to protect the users' voice-derived data from interception or tampering while it travels over the network to the remote service, securing the voice interaction from end to end. Dadu discloses this benefit: "the client computing device 102 may securely transfer audio to the server 106 via a network 104. To facilitate such secure transfer, the client computing device 102 and the server 106 exchange cryptographic keys and establish a secure connection" (Para 14).
Claim 37 recites substantially the same limitations as system claim 28, with the encryption performed "before or during" the uploading; the showing at claim 28 of encrypting before the upload satisfies at least one of the recited alternatives. Therefore, claim 37 is rejected under the same rationale as claim 28.
Claim 30 is rejected under 35 U.S.C. 103 as being unpatentable over Mixter (US 2017/0329573 A1) in view of Sullivan et al. (US 2017/0064358 A1) and Vinyals et al. (US 2015/0356401 A1), and further in view of Wang et al., "Multi-task Representation Learning for Demographic Prediction," hereinafter "Wang".
Regarding claim 30, Mixter in view of Sullivan, further in view of Vinyals, teaches all the limitations of claim 21, including the encoder model that generates the fixed-size representation (Vinyals, as combined for claim 21) and the groups of users and their demographic attributes (Mixter in view of Sullivan, as combined for claim 21).
Wang further teaches:
wherein: the encoder model is trained as part of a multitask neural network that uses a plurality of decoders to predict a plurality of attributes of different user groups; and the training of the multitask neural network trains the encoder to embed signals of the attributes in the fixed-sized representation (Abstract, "MTRL uses a supervised way to learn the shared semantic representation across multiple tasks, thus it can obtain a more general and robust representation by considering the constraints among tasks. Experiments are conducted on a real-world retail dataset where three attributes (gender, marital status, and education level) are predicted." Fig. 1 (p. 6), "The lower two layers are shared across all the tasks, while top layers are task-specific. The input is represented as a bag of items, then a non-linear projection W is used to generate a shared representation. Finally, for each task, additional non-linear projection V generates task-specific representations." Section 3 (pp. 6-7), "H^t = [h_j,k] is the matrix that maps task specific representation to the output layer for task t"; Eq. (1) (p. 7), "The objective function of MTRL is then defined as the cross-entropy over the outputs of all the users and all tasks"; Section 4.1 (p. 8), "we set the dimensionality of shared representation layer and task-specific representation layer as 200 and 100 respectively." - EN: Wang's MTRL is a multitask neural network in which one shared layer generates a fixed-dimensionality shared representation from a user's item history, and a separate task-specific branch (V^t, H^t, and a softmax output layer) maps that representation to a prediction of one attribute. Each branch decodes the shared representation into an attribute prediction and is therefore a decoder, and the three branches predicting gender, marital status, and education level are a plurality of decoders predicting a plurality of attributes. Because a single objective spans all tasks (Eq. (1)) and one model is updated by gradients from every task (Algorithm 1, p. 7), the training trains the shared layer to embed the signals of all the attributes in the shared representation. Wang does not teach the claimed encoder model or the claimed user groups; those are taught by the claim 21 combination (encoder model, Vinyals; groups and their demographic attributes, Mixter in view of Sullivan). In the combination, Vinyals's encoder takes the place of Wang's shared layer and is trained as part of Wang's multitask arrangement over the command histories of different groups, with decoders predicting the groups' attributes, so the training embeds the attribute signals in the fixed-size representation of claim 21.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the trained encoder generating the fixed-size group representation of Mixter in view of Sullivan, further in view of Vinyals (as set forth for claim 21), with the multitask training arrangement of Wang, in which a shared representation layer is trained jointly with a plurality of per-attribute prediction branches. The motivation for doing so would be to train the encoder using the labeled data of all the attribute-prediction tasks at once, so that attributes with limited labeled data are still learned well and the resulting representation is more general and robust than one trained for a single task; Wang discloses this benefit: "By using a multi-task approach to learn the tasks, our model leverages the large amounts of cross-task data, which is helpful to task with limited data," and it "can obtain a more general and robust representation by considering the constraints among tasks" (Abstract).
Claim 34 is rejected under 35 U.S.C. 103 as being unpatentable over Mixter (US 2017/0329573 A1) in view of Sullivan et al. (US 2017/0064358 A1) and Vinyals et al. (US 2015/0356401 A1), and further in view of Kim et al. (US 2015/0255068 A1), hereinafter "Kim".
Regarding claim 34,
Mixter in view of Sullivan, further in view of Vinyals, teaches all the limitations of claim 21, including the encoder model deployed to the voice-controlled device (Vinyals, as combined for claim 21) and the group of users of the device (Mixter in view of Sullivan, as combined for claim 21).
Kim further teaches:
wherein: the voice-controlled device is configured to perform further training of the encoder model after the encoder model is deployed to the voice-controlled device, wherein the further training adapts the encoder model to the group of users (Para 16, "the device/system 102 can be configured to automatically update user voice models over time." Para 36, "According to one embodiment, voice model updating is performed locally on the associated device/system." Para 27, "The fingerprint generator 106 can automatically create new voice models and/or incrementally update/refine existing voice models with new voice data." Para 35, "It will be appreciated that training data can be used to create a new voiceprint or update an existing voiceprint, whether generated while user(s) are speaking or based on previously collected voice data." - EN: Kim's device updates its voice models locally, using training data from its users' ongoing speech. Updating a model with training data is further training; because the updating refines models already existing on the device (Para 27), the training occurs after the model is deployed to the device; and because the training data is the device's own users' new voice data, the training adapts the model to those users. Kim does not teach the claimed encoder model or the claimed group of users, those are taught by the claim 21 combination (encoder model, Vinyals; group of users, Mixter in view of Sullivan) - and in the combination, Kim's local updating is applied to that deployed encoder model using the voice commands received from that group. See the motivation to combine below.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the encoder model deployed at the voice-controlled device of Mixter in view of Sullivan, further in view of Vinyals (as set forth for claim 21), with the local, automatic updating of deployed models of Kim, such that the device performs further training of the deployed encoder model using the voice commands it continues to receive from the group of users. The motivation for doing so would be to keep the deployed encoder model current and accurate for the actual group of users sharing the device, whose members and usage change over time, without requiring the users to perform any separate enrollment or retraining procedure. Kim discloses this benefit: "the device/system 102 can be configured to automatically update user voice models over time" (Para 16), which avoids "an untimely or potentially disrupting enrollment phase" that "can interrupt the natural flow of conversation" (Para 29).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAYMUR RAHMAN ALI whose telephone number is (571)272-0007. The examiner can normally be reached Mon-Fri. 9:30-6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAYMUR RAHMAN ALI/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123