Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
2. The information disclosure statement (IDS) submitted on October 5, 2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Double Patenting
3. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto- processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
4. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12183321. Although the claims at issue are not identical, they are not patentably distinct from each other because simply removing inherent and/or unnecessary limitations/step would be within the level of one of ordinary skill in the art. In re Karlson, 136 USPQ 184 (CCPA 1963). Also note Ex parte Rainu, 168 USPQ 375
(Bd. App. 1969). Omission of a reference element or step whose function is not needed would be obvious to one of ordinary skill in the art.
Regarding claim 1,
Instant Application
U.S. Patent No. 12,183,321
A method performed by one or more processors of a client device, the method comprising:
A method performed by one or more processors of a client device, the method comprising:
receiving, via one or more microphones of the client device, audio data that captures a spoken utterance of a user of the client device;
receiving, via one or more microphones of the client device, audio data that captures a spoken utterance of a user of the client device;
processing, using a speech recognition model that is stored locally at the client device, the audio data to generate a plurality of predicted textual segments for a portion of the spoken utterance;
processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance;
determining to cause at least two predicted textual segments, of the plurality of predicted textual segments, to be visually rendered at a display of the client device;
causing the at least two predicted textual segments to be visually rendered at the display of the client device:
causing at least part of the predicted textual segment to be visually rendered at a display of the client device;
subsequent to the causing the at least two predicted textual segments to be visually rendered at the display of the client device:
subsequent to the predicted textual segment being visually rendered:
receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance;
receiving, via one or more of the microphones of the client device, additional audio data that captures an additional spoken utterance of the user;
and processing, using the speech recognition model that is stored locally at the client device, the additional audio data to determine an alternate textual segment, the alternate textual segment being a correction to the predicted textual segment;
determining whether the alternate textual segment that is the correction to the predicted textual segment is directed to performance of the speech recognition model that is stored locally at the client device, wherein determining whether the alternate textual segment that is the correction to the predicted textual segment is directed to performance of the speech recognition model that is stored locally at the client device comprises:
determining a measure of similarity between the predicted textual segment and the alternate textual segment;
determining that the correction is directed to performance of the speech recognition model that is stored locally at the client device based on the measure of similarity satisfying a threshold;
and in response to receiving the user selection of the given predicted textual segment: and in response to receiving the user selection of the given predicted textual segment:
and in response to determining that the alternate textual segment that is the correction to the predicted textual segment is directed to performance of the speech recognition model that is stored locally at the client device:
utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance;
generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient;
generating a gradient based on comparing at least the predicted textual segment to the alternate textual segment,
and one or both of:
updating, based on the gradient, one or more weights of the speech recognition model;
and updating one or more weights of the speech recognition model based on the generated gradient.
or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments,
wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model.
Claim Rejections - 35 USC § 103
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically taught as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. Claims 1-6, 8-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Thomson (U.S. Patent No. 10388272) in view of Jost (U.S. Publication No. 20190318731) in view of Case (U.S. Publication No. 20200380369).
Regarding claim 1, Thomson discloses a method performed by one or more processors of a client device, (Col 6 – Rows 50-54 – each of the first device 104 and the second device 106 may include memory and at least one processor, which are configured to perform operations as described in this disclosure) the method comprising:
receiving, via one or more microphones of the client device, audio data that captures a spoken utterance of a user of the client device (Col 7 – Rows 19-22 – The first audio may include a first voice of the first user 110…the first audio from a microphone of the first device 104);
causing at least part of the predicted textual segment to be visually rendered at a display of the client device (Col 5 – Rows 29-30 – single transcription that is provided to a device for display to a user);
subsequent to the predicted textual segment being visually rendered at the display of the client device (Col 5 – Rows 29-30 – single transcription that is provided to a device for display to a user…[Col 144, Rows 24-27] – on a two-way communication session or conference communication session, the audio of the speaking party may be identified for the CA using visual and/or audio indicators):
However, Thomson does not teach processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance;
determining to cause at least two predicted textual segments to be visually rendered at the display of the client device;
receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance.
However, Jost does teach processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming):
determining to cause at least two predicted textual segments to be visually rendered at the display of the client device ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming [0143] - Also attached to the system bus are typically I/O device interfaces for connecting various input and output devices, e.g., keyboard, mouse, displays, printers, speakers, etc., to the computer);
receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson to include the teachings of Jost in order to implement processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance; determining to cause at least two predicted textual segments to be visually rendered at the display of the client device; receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance. Doing so allows for the improvement of ASR accuracy while making optimal use of the external information source (i.e., reduce the human effort that correction requires) (Jost [0024]).
However, Thomson in view of Jost does not teach in response to receiving the user selection of the given predicted textual segment:
utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance;
generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient;
and one or both of:
updating, based on the gradient, one or more weights of the speech recognition model;
or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments,
wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model.
Case does teach in response to receiving the user selection of the given predicted textual segment:
utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance ([0075] - a gradient is computed based on an error that is computed using ground truth data);
generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient ([0075] - a gradient is computed based on an error that is computed using ground truth data);
and one or both of:
updating, based on the gradient, one or more weights of the speech recognition model ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s));
or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s)),
wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s)).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement in response to receiving the user selection of the given predicted textual segment: utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance; generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient; and one or both of: updating, based on the gradient, one or more weights of the speech recognition model; or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments, wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Regarding claim 2, Thomson in view of Jost in view of Case teaches all limitations of claim 1, above.
Thomson does not disclose the method, wherein determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises:
determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments;
and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments;
and determining that both the first confidence measure and the second confidence measure satisfy a confidence measure threshold.
Jost does teach the method, wherein determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises:
determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments ([0079] - linguistic probability of the word sequence w_1 . . . w_m, and…);
and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments ([0080] - linguistic probability of the word w_n given that the preceding c words in the sequence are w_n−1, . . . ,w_n−c…);
and determining that both the first confidence measure and the second confidence measure satisfy a confidence measure threshold ([0054] - the process can stop ASR when any remaining hypotheses are below a threshold or a time limit is reached. The process can present the n-best structure to the external source for confirmation, selection or correction)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson to include the teachings of Jost in order to implement wherein determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises: determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments; and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments; and determining that both the first confidence measure and the second confidence measure satisfy a confidence measure threshold. Doing so allows for the improvement of ASR accuracy while making optimal use of the external information source (i.e., reduce the human effort that correction requires) (Jost [0024]).
Regarding claim 3, Thomson in view of Jost in view of Case teaches all limitations of claim 1, above.
Thomson does not disclose the method, wherein determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises:
determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments;
and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments;
and determining that the first confidence measure and the second confidence measure both fail to satisfy a confidence measure threshold.
Jost does teach the method, wherein determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises:
determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments ([0079] - linguistic probability of the word sequence w_1 . . . w_m, and…);
and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments ([0080] - linguistic probability of the word w_n given that the preceding c words in the sequence are w_n−1, . . . ,w_n−c…);
and determining that the first confidence measure and the second confidence measure both fail to satisfy a confidence measure threshold ([0054] - the process can stop ASR when any remaining hypotheses are below a threshold or a time limit is reached. The process can present the n-best structure to the external source for confirmation, selection or correction).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson to include the teachings of Jost in order to implement determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises: determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments; and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments; and determining that the first confidence measure and the second confidence measure both fail to satisfy a confidence measure threshold. Doing so allows for the improvement of ASR accuracy while making optimal use of the external information source (i.e., reduce the human effort that correction requires) (Jost [0024]).
Regarding claim 4, Thomson in view of Jost in view of Case teaches all limitations of claim 1, above.
Thomson does not disclose the method, wherein determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises:
determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments;
and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments;
and determining that the first confidence measure and the second confidence measure are within a threshold confidence range.
Jost does teach the method, wherein determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises:
determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments ([0079] - linguistic probability of the word sequence w_1 . . . w_m, and…);
and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments ([0080] - linguistic probability of the word w_n given that the preceding c words in the sequence are w_n−1, . . . ,w_n−c…);
and determining that the first confidence measure and the second confidence measure are within a threshold confidence range ([0054] - the process can stop ASR when any remaining hypotheses are below a threshold or a time limit is reached. The process can present the n-best structure to the external source for confirmation, selection or correction).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson to include the teachings of Jost in order to implement determining to cause the at least two predicted textual segments to be visually rendered at the display of the client device comprises: determining a first confidence measure that is associated with a first predicted textual segment of the at least two predicted textual segments; and determining a second confidence measure that is associated with a second predicted textual segment of the at least two predicted textual segments; and determining that the first confidence measure and the second confidence measure are within a threshold confidence range. Doing so allows for the improvement of ASR accuracy while making optimal use of the external information source (i.e., reduce the human effort that correction requires) (Jost [0024]).
Regarding claim 5, Thomson in view of Jost in view of Case teaches all limitations of claim 1, above.
However, Thomson does not disclose the method, further comprising:
storing, locally at the client device, the plurality of predicted textual segments and the ground truth textual segment for the portion of the spoken utterance;
determining one or more conditions are satisfied;
and wherein generating the gradient based on the ground truth textual segment for the portion of the spoken utterance is in response to determining the one or more conditions are satisfied.
Case does teach the method, further comprising:
storing, locally at the client device, the plurality of predicted textual segments and the ground truth textual segment for the portion of the spoken utterance ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s) [0075] - a gradient is computed based on an error that is computed using ground truth data);
determining one or more conditions are satisfied ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s) [0075] - a gradient is computed based on an error that is computed using ground truth data);
and wherein generating the gradient based on the ground truth textual segment for the portion of the spoken utterance is in response to determining the one or more conditions are satisfied ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s) [0075] - a gradient is computed based on an error that is computed using ground truth data).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement the method, further comprising: storing, locally at the client device, the plurality of predicted textual segments and the ground truth textual segment for the portion of the spoken utterance; determining one or more conditions are satisfied; and wherein generating the gradient based on the ground truth textual segment for the portion of the spoken utterance is in response to determining the one or more conditions are satisfied. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Regarding claim 6, Thomson in view of Jost in view of Case teaches all limitations of claim 5, above.
However, Thomson does not disclose the method, wherein generating the gradient based on the ground truth textual segment for the portion of the spoken utterance comprises: ‘
comparing the ground truth textual segment for the portion of the spoken utterance to one or more of the plurality of predicted textual segments;
and generating, based on comparing the ground truth textual segment for the portion of the spoken utterance to one or more of the plurality of predicted textual segments, the gradient.
Case does teach the method, wherein generating the gradient based on the ground truth textual segment for the portion of the spoken utterance comprises: ‘
comparing the ground truth textual segment for the portion of the spoken utterance to one or more of the plurality of predicted textual segments ([0041] - using partial/sparse weight updates such that weights. In at least one embodiment, an embedding based neural network has a sparse input domain such that only a small portion of weights is used by each step (e.g., batch or minibatch) of training. In at least one embodiment, a neural network is trained on a set of weights which are adjusted as part of training [0075] - a gradient is computed based on an error that is computed using ground truth data);
and generating, based on comparing the ground truth textual segment for the portion of the spoken utterance to one or more of the plurality of predicted textual segments, the gradient ([0041] - using partial/sparse weight updates such that weights. In at least one embodiment, an embedding based neural network has a sparse input domain such that only a small portion of weights is used by each step (e.g., batch or minibatch) of training. In at least one embodiment, a neural network is trained on a set of weights which are adjusted as part of training [0075] - a gradient is computed based on an error that is computed using ground truth data).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement the method, wherein generating the gradient based on the ground truth textual segment for the portion of the spoken utterance comprises: comparing the ground truth textual segment for the portion of the spoken utterance to one or more of the plurality of predicted textual segments; and generating, based on comparing the ground truth textual segment for the portion of the spoken utterance to one or more of the plurality of predicted textual segments, the gradient. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Regarding claim 8, Thomson in view of Jost in view of Case teaches all limitations of claim 1, above.
However, Thomson does not disclose the method, wherein one or more weights of the speech recognition model are updated based on the gradient.
Case does teach the method, wherein one or more weights of the speech recognition model are updated based on the gradient ([0041] - using partial/sparse weight updates such that weights. In at least one embodiment, an embedding based neural network has a sparse input domain such that only a small portion of weights is used by each step (e.g., batch or minibatch) of training. In at least one embodiment, a neural network is trained on a set of weights which are adjusted as part of training [0075] - a gradient is computed based on an error that is computed using ground truth data).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement the method, wherein one or more weights of the speech recognition model are updated based on the gradient. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Regarding claim 9, Thomson in view of Jost in view of Case teaches all limitations of claim 8, above.
However, Thomson does not disclose the method, wherein the gradient is transmitted over the network and to the remote system for utilization in updating the one or more global weights of the global speech recognition model.
Case does teach the method, wherein the gradient is transmitted over the network and to the remote system for utilization in updating the one or more global weights of the global speech recognition model ([0041] - using partial/sparse weight updates such that weights. In at least one embodiment, an embedding based neural network has a sparse input domain such that only a small portion of weights is used by each step (e.g., batch or minibatch) of training. In at least one embodiment, a neural network is trained on a set of weights which are adjusted as part of training [0075] - a gradient is computed based on an error that is computed using ground truth data).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement the method, wherein the gradient is transmitted over the network and to the remote system for utilization in updating the one or more global weights of the global speech recognition model. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Regarding claim 10, Thomson in view of Jost in view of Case teaches all limitations of claim 1, above.
However, Thomson does not disclose the method, wherein the gradient is transmitted over the network and to the remote system for utilization in updating the one or more global weights of the global speech recognition model.
Case does teach the method, wherein the gradient is transmitted over the network and to the remote system for utilization in updating the one or more global weights of the global speech recognition model ([0041] - using partial/sparse weight updates such that weights. In at least one embodiment, an embedding based neural network has a sparse input domain such that only a small portion of weights is used by each step (e.g., batch or minibatch) of training. In at least one embodiment, a neural network is trained on a set of weights which are adjusted as part of training [0075] - a gradient is computed based on an error that is computed using ground truth data).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement the method, wherein the gradient is transmitted over the network and to the remote system for utilization in updating the one or more global weights of the global speech recognition model. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Regarding claim 11, Thomson in view of Jost in view of Case teaches all limitations of claim 10, above.
However, Thomson does not disclose the method, wherein one or more weights of the speech recognition model are updated based on the gradient.
Case does teach the method, wherein one or more weights of the speech recognition model are updated based on the gradient ([0041] - using partial/sparse weight updates such that weights. In at least one embodiment, an embedding based neural network has a sparse input domain such that only a small portion of weights is used by each step (e.g., batch or minibatch) of training. In at least one embodiment, a neural network is trained on a set of weights which are adjusted as part of training [0075] - a gradient is computed based on an error that is computed using ground truth data).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement the method, wherein one or more weights of the speech recognition model are updated based on the gradient. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Regarding claim 12, Thomson in view of Jost in view of Case teaches all limitations of claim 1, above.
However, Thomson does not disclose the method, wherein the user selection of the given predicted textual segment is one of: a touch selection directed to the display of the client device of the user, or a voice selection captured via one or more microphones of the client device of the user.
Jost does teach wherein the user selection of the given predicted textual segment is one of: a touch selection directed to the display of the client device of the user, or a voice selection captured via one or more microphones of the client device of the user ([0038] - System 100 is configured for performing bidirectional automatic speech recognition of voice data 102 using an external information source 140, e.g., human agent).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Case to include the teachings of Jost in order to implement wherein the user selection of the given predicted textual segment is one of: a touch selection directed to the display of the client device of the user, or a voice selection captured via one or more microphones of the client device of the user. Doing so allows for the improvement of ASR accuracy while making optimal use of the external information source (i.e., reduce the human effort that correction requires) (Jost [0024]).
Regarding claim 13, Thomson discloses a system comprising:
at least one processor (Col 6 – Rows 50-54 – each of the first device 104 and the second device 106 may include memory and at least one processor, which are configured to perform operations as described in this disclosure);
and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to (Col 6 – Rows 50-54 – each of the first device 104 and the second device 106 may include memory and at least one processor, which are configured to perform operations as described in this disclosure):
receive, via one or more microphones of the client device, audio data that captures a spoken utterance of a user of the client device (Col 7 – Rows 19-22 – The first audio may include a first voice of the first user 110…the first audio from a microphone of the first device 104);
cause the at least part of the predicted textual segment to be visually rendered at a display of the client device (Col 5 – Rows 29-30 – single transcription that is provided to a device for display to a user);
subsequent to the predicted textual segment being visually rendered at the display of the client device (Col 5 – Rows 29-30 – single transcription that is provided to a device for display to a user…[Col 144, Rows 24-27] – on a two-way communication session or conference communication session, the audio of the speaking party may be identified for the CA using visual and/or audio indicators):
However, Thomson does not teach processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance;
determining to cause at least two predicted textual segments to be visually rendered at the display of the client device;
receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance.
However, Jost does teach processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming):
determining to cause at least two predicted textual segments to be visually rendered at the display of the client device ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming [0143] - Also attached to the system bus are typically I/O device interfaces for connecting various input and output devices, e.g., keyboard, mouse, displays, printers, speakers, etc., to the computer);
receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson to include the teachings of Jost in order to implement processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance; determining to cause at least two predicted textual segments to be visually rendered at the display of the client device; receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance. Doing so allows for the improvement of ASR accuracy while making optimal use of the external information source (i.e., reduce the human effort that correction requires) (Jost [0024]).
However, Thomson in view of Jost does not teach in response to receiving the user selection of the given predicted textual segment:
utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance;
generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient;
and one or both of:
updating, based on the gradient, one or more weights of the speech recognition model;
or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments,
wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model.
Case does teach in response to receiving the user selection of the given predicted textual segment:
utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance ([0075] - a gradient is computed based on an error that is computed using ground truth data);
generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient ([0075] - a gradient is computed based on an error that is computed using ground truth data);
and one or both of:
updating, based on the gradient, one or more weights of the speech recognition model ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s));
or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s)),
wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s)).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement in response to receiving the user selection of the given predicted textual segment: utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance; generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient; and one or both of: updating, based on the gradient, one or more weights of the speech recognition model; or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments, wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
Dependent claims 14-18 are analogous in scope to claims 2-6, and are rejected according to the same reasoning.
Regarding claim 20, Thomson discloses a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to execute the instructions to (Col 6 – Rows 50-54 – each of the first device 104 and the second device 106 may include memory and at least one processor, which are configured to perform operations as described in this disclosure):
receive, via one or more microphones of the client device, audio data that captures a spoken utterance of a user of the client device (Col 7 – Rows 19-22 – The first audio may include a first voice of the first user 110…the first audio from a microphone of the first device 104);
cause the at least part of the predicted textual segment to be visually rendered at a display of the client device (Col 5 – Rows 29-30 – single transcription that is provided to a device for display to a user);
subsequent to the predicted textual segment being visually rendered at the display of the client device (Col 5 – Rows 29-30 – single transcription that is provided to a device for display to a user…[Col 144, Rows 24-27] – on a two-way communication session or conference communication session, the audio of the speaking party may be identified for the CA using visual and/or audio indicators):
However, Thomson does not teach processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance;
determining to cause at least two predicted textual segments to be visually rendered at the display of the client device;
receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance.
However, Jost does teach processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming):
determining to cause at least two predicted textual segments to be visually rendered at the display of the client device ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming [0143] - Also attached to the system bus are typically I/O device interfaces for connecting various input and output devices, e.g., keyboard, mouse, displays, printers, speakers, etc., to the computer);
receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance ([0021] - The computer code instructions, when executed by a processor, cause a device (or a system) to perform at least the following: …pre-processing an utterance in a backward direction from end to start of the utterance…presenting a word prediction based on the prediction list to an external information source to obtain a response from the external information source confirming).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson to include the teachings of Jost in order to implement processing, using a speech recognition model stored locally at the client device, the audio data to determine a predicted textual segment that is a prediction of the spoken utterance; determining to cause at least two predicted textual segments to be visually rendered at the display of the client device; receiving, via the client device, a user selection of a given predicted textual segment, from among the at least two predicted textual segments, that indicates the given predicted textual segment corresponds to the portion of the spoken utterance. Doing so allows for the improvement of ASR accuracy while making optimal use of the external information source (i.e., reduce the human effort that correction requires) (Jost [0024]).
However, Thomson in view of Jost does not teach in response to receiving the user selection of the given predicted textual segment:
utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance;
generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient;
and one or both of:
updating, based on the gradient, one or more weights of the speech recognition model;
or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments,
wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model.
Case does teach in response to receiving the user selection of the given predicted textual segment:
utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance ([0075] - a gradient is computed based on an error that is computed using ground truth data);
generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient ([0075] - a gradient is computed based on an error that is computed using ground truth data);
and one or both of:
updating, based on the gradient, one or more weights of the speech recognition model ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s));
or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s)),
wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model ([0059] - second pipeline iteration 302 illustrates a step (batch) of training that comprises: loading input data; partial/sparse weight updating by non-gradient term(s); forward propagation; backwards propagation; and partial/sparse weight updating by gradient term(s)).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Jost to include the teachings of Case in order to implement in response to receiving the user selection of the given predicted textual segment: utilizing the given predicted textual segment as a ground truth textual segment for the portion of the spoken utterance; generating, based on at least the ground truth textual segment for the portion of the spoken utterance, a gradient; and one or both of: updating, based on the gradient, one or more weights of the speech recognition model; or transmitting, over a network and to a remote system, the gradient without transmitting any of: the audio data, or the plurality of predicted textual segments, wherein the remote system utilizes the gradient to update one or more global weights of a global speech recognition model. Doing so allows users to train or perform inferencing of information directly onto hardware (Case [0098]).
7. Claims 7 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Thomson (U.S. Patent No. 10388272) in view of Jost (U.S. Publication No. 20190318731) in view of Case (U.S. Publication No. 20200380369) in view of Alpert (U.S. Patent No. 7389403).
Regarding claim 7, Thomson in view of Jost in view of Case teaches all limitations of claim 5, above.
However, Thomson in view of Jost in view of Case does not teach the method, wherein the one or more conditions comprise one or more of: that the client device is charging, that the client device has at least a threshold state of charge, that a temperature of the client device is less than a threshold, or that the client device is not being held by a user.
Alpert does teach the method, wherein the one or more conditions comprise one or more of: that the client device is charging, that the client device has at least a threshold state of charge, that a temperature of the client device is less than a threshold, or that the client device is not being held by a user ((Paragraph 91) - the switching is in response to detection of at least one of a low energy condition, a low battery condition, a high temperature condition, an insufficient performance condition, and a scheduling deadline condition).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the application to modify the teaching of Thomson in view of Case in view of Jost to include the teachings of Alpert in order to implement the method, wherein the one or more conditions comprise one or more of: that the client device is charging, that the client device has at least a threshold state of charge, that a temperature of the client device is less than a threshold, or that the client device is not being held by a user. Doing so allows improvements in performance, efficiency, cost, power tradeoffs, and utility of use (Alpert (Paragraph 3)).
Dependent claim 19 is analogous in scope to claim 7 and is rejected according to the same reasoning.
Conclusion
8. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Adams (U.S. Publication No. 20180349348) teaches generating predictive texts on an electronic device. Davis (U.S. Publication No. 20120277893) teaches channelized audio watermarks. Kumar (U.S. Patent No. 10335572) teaches systems and methods for computer assisted operation. Mirowski (U.S. Publication No. 20120150532) teaches a system and method for feature-rich continuous space language models. Ramer (U.S. Patent No. 9703892) teaches predictive text completion for a mobile communication facility. Sicconi (U.S. Publication No. 20200057287) teaches methods and systems for using artificial intelligence to evaluate, correct, and monitor user attentiveness. Tran (U.S. Patent No. 10325596) teaches voice control of appliances. Tran (U.S. Publication No. 20170312614) teaches a smart device.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ETHAN DANIEL KIM whose telephone number is (571) 272-1405. The examiner can normally be reached on Monday - Friday 9:00 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached on (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ETHAN DANIEL KIM/
Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658