DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Introduction
This office action is in response to communications filed 01/30/2025. Claims 1-20 are pending and likewise have been examined.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/30/2025 and 07/28/2025 is in compliance with the provisions of 37 CFR 1.97. References not provided by Applicant were made of record by Applicant and Examiner in Parent application 17/659,836. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 4 and 14 objected to because of the following informalities:
Claim 4 and 14 recites “the base ASR are frozen” on lines 1-2. The word “model” should be included for proper antecedent basis. The claim should read “the base ASR model are frozen”, like in claim 5.
Appropriate correction is required.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 1-20 rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 6-11 and 16-20 of U.S. Patent No. 12230258 in view of Sugaya (US 20210312930 A1). See Table below for portions of Claims that are taught by U.S. Patent No. 12230258.
Instant Application
U.S. Patent No. 12230258
Claim 1:
Claim 1:
A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:
A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:
obtaining a base automatic speech recognition (ASR) model trained on non-biased data;
obtaining a base automatic speech recognition (ASR) model trained on non-biased data;
obtaining a sub-model trained on biased data, the biased data associated with a respective domain;
obtaining a sub-model trained on biased data, the biased data representative of a particular domain;
receiving a speech recognition request comprising audio data characterizing a spoken utterance;
receiving a speech recognition request comprising audio data characterizing an utterance captured in streaming audio;
processing, using the base ASR model, the audio data to generate a first speech recognition result of the spoken utterance;
generating, using the base ASR model, a first speech recognition result of the utterance by processing the audio data;
processing, using the sub-model, the audio data to generate a second speech recognition result of the spoken utterance, the second speech recognition result biased toward one or more terms in the respective domain associated with the biased data;
generating, using the base ASR model, an encoded output by processing the audio data; biasing, using the sub-model, the base ASR model toward the particular domain; generating, using the biased base ASR model, a sub-model output by processing the audio data, the sub-model output generated in parallel with the encoded output; and generating, using a decoder of the base ASR model, a second speech recognition result of the utterance by processing the encoded output and the sub-model output, the second speech recognition result biased toward one or more terms in the particular domain
and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen.
See rejection with additional reference below table.
Claim 2: The computer-implemented method of claim 1, wherein the sub-model is disposed in a layer of the base ASR model.
Claim 6: The computer-implemented method of claim 1, wherein the sub-model is disposed in a layer of the base ASR model.
Claim 3: The computer-implemented method of claim 2, wherein the base ASR model comprises an encoder and a decoder; and the sub-model is disposed in between two layers of the encoder.
Claim 7: The computer-implemented method of claim 6, wherein: the base ASR model comprises an encoder and the decoder; and the sub-model is disposed in between two layers of the encoder.
Claim 4: The computer-implemented method of claim 1, wherein parameters of the base ASR are frozen when generating the first speech recognition result.
Claim 8: The computer-implemented method of claim 1, wherein one or more parameters of the base ASR model are frozen when generating the first and second speech recognition results.
Claim 5: The computer-implemented method of claim 1, wherein parameters of the base ASR model are frozen when using the sub-model to process the audio data to generate the second speech recognition result.
Claim 8: The computer-implemented method of claim 1, wherein one or more parameters of the base ASR model are frozen when generating the first and second speech recognition results.
Claim 6: The computer-implemented method of claim 1, wherein the sub-model comprises a residual adapter layer of the base ASR model.
See note #1 below
Claim 7: The computer-implemented method of claim 1, wherein the biased data associated with the respective domain comprises words or phrases corresponding to a software application associated with the respective domain.
See note #1 below
Claim 8: The computer-implemented method of claim 1, wherein the biased data associated with the respective domain comprises words or phrases corresponding to a geographical location associated with the respective domain.
See note #1 below
Claim 9: The computer-implemented method of claim 1, wherein: the data processing hardware resides on a user device that captured the spoken utterance; or the data processing hardware resides on a remote system in communication with the user device via a network.
Claim 9: The computer-implemented method of claim 1, wherein: the data processing hardware resides on a user device that captured the utterance in the streaming audio; or the data processing hardware resides on a remote system in communication with the user device via a network.
Claim 10: The computer-implemented method of claim 1, wherein the first speech recognition result is different from the second speech recognition result.
Claim 10: The computer-implemented method of claim 1, wherein the first speech recognition result is different from the second speech recognition result.
Claim 11: A system comprising: data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
Claim 11: A system comprising: data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
obtaining a base automatic speech recognition (ASR) model trained on non-biased data;
obtaining a base automatic speech recognition (ASR) model trained on non-biased data;
obtaining a sub-model trained on biased data, the biased data associated with a respective domain;
obtaining a sub-model trained on biased data, the biased data representative of a particular domain;
receiving a speech recognition request comprising audio data characterizing a spoken utterance;
receiving a speech recognition request comprising audio data characterizing an utterance captured in streaming audio;
processing, using the base ASR model, the audio data to generate a first speech recognition result of the spoken utterance;
generating, using the base ASR model, a first speech recognition result of the utterance by processing the audio data;
processing, using the sub-model, the audio data to generate a second speech recognition result of the spoken utterance, the second speech recognition result biased toward one or more terms in the respective domain associated with the biased data;
generating, using the base ASR model, an encoded output by processing the audio data; biasing, using the sub-model, the base ASR model toward the particular domain; generating, using the biased base ASR model, a sub-model output by processing the audio data, the sub-model output generated in parallel with the encoded output; and generating, using a decoder of the base ASR model, a second speech recognition result of the utterance by processing the audio data encoded output and the sub-model output, the second speech recognition result biased toward one or more terms in the particular domain.
and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen.
See rejection with additional reference below table.
Claim 12: The system of claim 11, wherein the sub-model is disposed in a layer of the base ASR model.
Claim 16: The system of claim 11, wherein the sub-model is disposed in a layer of the base ASR model.
Claim 13: The system of claim 12, wherein the base ASR model comprises an encoder and a decoder; and the sub-model is disposed in between two layers of the encoder.
Claim 17: The system of claim 16, wherein: the base ASR model comprises an encoder and the decoder; and the sub-model is disposed in between two layers of the encoder.
Claim 14: The system of claim 11, wherein parameters of the base ASR are frozen when generating the first speech recognition result.
Claim 18: The system of claim 11, wherein one or more parameters of the base ASR model are frozen when generating the first and second speech recognition results.
Claim 15: The system of claim 11, wherein parameters of the base ASR model are frozen when using the sub-model to process the audio data to generate the second speech recognition result.
Claim 18: The system of claim 11, wherein one or more parameters of the base ASR model are frozen when generating the first and second speech recognition results.
Claim 16: The system of claim 11, wherein the sub-model comprises a residual adapter layer of the base ASR model.
See note #1 below
Claim 17: The system of claim 11, wherein the biased data associated with the respective domain comprises words or phrases corresponding to a software application associated with the respective domain.
See note #1 below
Claim 18: The system of claim 11, wherein the biased data associated with the respective domain comprises words or phrases corresponding to a geographical location associated with the respective domain.
See note #1 below
Claim 19: The system of claim 11, wherein: the data processing hardware resides on a user device that captured the spoken utterance; or the data processing hardware resides on a remote system in communication with the user device via a network.
Claim 19: The system of claim 11, wherein: the data processing hardware resides on a user device that captured the utterance in the streaming audio; or the data processing hardware resides on a remote system in communication with the user device via a network.
Claim 20: The system of claim 11, wherein the first speech recognition result is different from the second speech recognition result.
Claim 20: The system of claim 11, wherein the first speech recognition result is different from the second speech recognition result.
Note #1
Although the claims at issue are not identical, they are not patentably distinct from each other because simply changing the statutory category from computer program product and system, to a method and removing inherent and/or unnecessary limitations/step would be within the level of one of ordinary skill in the art. It is well settled that the omission of an element, e.g. “a computer-readable storage medium”, and its function is an obvious expedient if the remaining elements perform the same function as before. In re Karlson, 136 USPQ 184 (CCPA 1963). Also note Ex parte Rainu, 168 USPQ 375 (Bd. App. 1969). Omission of a reference element or step whose function is not needed would be obvious to one of ordinary skill in the art.
Regarding Claim 1:
U.S. Patent No. 12230258 Claims the limitations shown the table above, but does not Claim and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen.
In the same field of Speech Recognition, Sugaya teaches and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen(Para [0065], Ln 1-5, The user terminal receives the recognition result data and displays the first recognition text and the second recognition text on its display unit based on the recognition result data. Abstract, Ln 1-14, The computer system acquires voice data; performs voice recognition for the acquired voice data; performs voice recognition for the acquired voice data with an algorithm or a database different from that used by the first recognition unit; and outputs both of the recognition results when the recognition results from the voice recognitions are different.).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify U.S. Patent No. 12230258 with the user feedback on recognition of Sugaya, as it provides feedback to the system allowing it to improve the quality of the results(Para [0069], Ln 1-13, & Para [0071], Ln 1-12).
Regarding Claim 11:
U.S. Patent No. 12230258 Claims the limitations shown the table above, but does not Claim and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen.
In the same field of Speech Recognition, Sugaya teaches and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen(Para [0065], Ln 1-5, The user terminal receives the recognition result data and displays the first recognition text and the second recognition text on its display unit based on the recognition result data. Abstract, Ln 1-14, The computer system acquires voice data; performs voice recognition for the acquired voice data; performs voice recognition for the acquired voice data with an algorithm or a database different from that used by the first recognition unit; and outputs both of the recognition results when the recognition results from the voice recognitions are different.).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify U.S. Patent No. 12230258 with the user feedback on recognition of Sugaya, as it provides feedback to the system allowing it to improve the quality of the results(Para [0069], Ln 1-13, & Para [0071], Ln 1-12).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6, 9-16 and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kannan et al.(US 20200380215 A1), and further in view of Sugaya (US 20210312930 A1).
Regarding Claim 1:
Kannan teaches a computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising(Para [0012], Ln 1-10, processing hardware….instructions….operations):
obtaining a base automatic speech recognition (ASR) model trained on non-biased data(Para [0011], Ln 10-18, training the multilingual E2E speech recognition model on a union of all the training data sets);
obtaining a sub-model trained on biased data, the biased data associated with a respective domain(Para [0011], Ln 27-33, during the second stage….for each of the one or more adaptor modules that are specific to the particular native language, learning values for a respective set of weights by training the multilingual E2E speech recognition model only on the training data set that is associated with the respective particular native language);
receiving a speech recognition request comprising audio data characterizing a spoken utterance(Abstract, Ln 1-5, speech recognition model includes receiving audio data for an utterance spoken in a particular native language);
processing, using the base ASR model, the audio data to generate a first speech recognition result of the spoken utterance(Para [0011], Ln 10-18, training the multilingual E2E speech recognition model on a union of all the training data sets. Para [0045], Ln 1-7, Furthermore, adapter modules 300 do not need to be employed when they are not helpful. For example, in practice, adapter modules may only be selected or used when they are most effective, while setting all other adapters to identity operations, i.e., setting the weights to zeroes. Para [0038], Ln 1-9, processes the input vector 119 associated with the utterance 106a spoken in English by passing the input vector 119 through the first LSTM layer 216a, passing an output from the first LSTM layer 216a through the appropriate adaptor module 300c for English utterance. Para [0046], Ln 1-5, effectiveness of adapters on each language is determined by measuring the word error rate (WER) of both the model augmented with language-specific adapters and the model without adapters on a test set in that language);
processing, using the sub-model, the audio data to generate a second speech recognition result of the spoken utterance, the second speech recognition result biased toward one or more terms in the respective domain associated with the biased data(Para [0038], Ln 1-9, processes the input vector 119 associated with the utterance 106a spoken in English by passing the input vector 119 through the first LSTM layer 216a, passing an output from the first LSTM layer 216a through the appropriate adaptor module 300c for English utterance. Para [0029], Ln 1-14, language vector 115 serves as an extension to the multilingual E2E model 150 by permitting the model 150 to be fed the identified language for activating appropriate language-specific adaptor modules 300 (FIGS. 2A, 2B, and 3) of an encoder network 210. In doing so, the language-specific adaptor models 218 may capture quality benefits of fine-tuning the model 150 on each individual language);
Kannan does not teach and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen.
In the same field of Speech Recognition, Sugaya teaches and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen(Para [0065], Ln 1-5, The user terminal receives the recognition result data and displays the first recognition text and the second recognition text on its display unit based on the recognition result data. Abstract, Ln 1-14, The computer system acquires voice data; performs voice recognition for the acquired voice data; performs voice recognition for the acquired voice data with an algorithm or a database different from that used by the first recognition unit; and outputs both of the recognition results when the recognition results from the voice recognitions are different.).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify Kannan with the user feedback on recognition of Sugaya, as it provides feedback to the system allowing it to improve the quality of the results(Para [0069], Ln 1-13, & Para [0071], Ln 1-12).
Regarding Claim 2:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, and Kannan teaches wherein the sub-model is disposed in a layer of the base ASR model(Para [0010], Ln 16-26, encoder network may include a plurality of stacked Long Short-Term Memory (LSTM) layers and after each LSTM layer, a respective layer that includes a respective subset of the plurality of language-specific adaptor modules. Here, each language-specific adaptor module in the respective layer is specific to a different respective native language. See Fig 2B).
Regarding Claim 3:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 2, and Kannan teaches wherein the base ASR model comprises an encoder and a decoder(Para [0010], Ln 1-5, encoder network, a prediction network, and a joint network. Para [0010], Ln 5-16, prediction network is configured to process a sequence of previously output non-blank symbols into a dense representation. The joint network configured to predict, at each of the plurality of time steps, a probability distribution over possible output labels based on the higher-order feature representation output by the encoder network and the dense representation output by the prediction network. The joint network is decoding);
and the sub-model is disposed in between two layers of the encoder(Para [0010], Ln 16-26, encoder network may include a plurality of stacked Long Short-Term Memory (LSTM) layers and after each LSTM layer, a respective layer that includes a respective subset of the plurality of language-specific adaptor modules. Here, each language-specific adaptor module in the respective layer is specific to a different respective native language. See Fig 2B).
Regarding Claim 4:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, and Kannan teaches wherein parameters of the base ASR are frozen when generating the first speech recognition result(Para [0046], Ln 1-5, effectiveness of adapters on each language is determined by measuring the word error rate (WER) of both the model augmented with language-specific adapters and the model without adapters on a test set in that language. Para [0042], Ln 11-7, In the first stage…..update model weights. Model weights are adjusted during training, not during testing).
Regarding Claim 5:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, and Kannan teaches wherein parameters of the base ASR model are frozen when using the sub-model to process the audio data to generate the second speech recognition result(Para [0046], Ln 1-5, effectiveness of adapters on each language is determined by measuring the word error rate (WER) of both the model augmented with language-specific adapters and the model without adapters on a test set in that language. Para [0042], Ln 11-7, In the first stage…..update model weights. Model weights are adjusted during training, not during testing).
Regarding Claim 6:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, and Kannan teaches wherein the sub-model comprises a residual adapter layer of the base ASR model(Para [0010], Ln 11-26, encoder network may include a plurality of stacked Long Short-Term Memory (LSTM) layers and after each LSTM layer, a respective layer that includes a respective subset of the plurality of language-specific adaptor modules. Here, each language-specific adaptor module in the respective layer is specific to a different respective native language).
Regarding Claim 9:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, and Kannan teaches wherein: the data processing hardware resides on a user device that captured the spoken utterance(Para [0024], Ln 11-17, While the example shown depicts the ASR system 100 residing on a user device 102);
or the data processing hardware resides on a remote system in communication with the user device via a network(Para [0024], Ln 11-17, While the example shown depicts the ASR system 100 residing on a user device 102, some or all components of the ASR system 100 may reside on a remote computing device 201 (e.g., one or more servers).
Regarding Claim 10:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, and Kannan teaches wherein the first speech recognition result is different from the second speech recognition result(Para [0029], Ln 1-14, language vector 115 serves as an extension to the multilingual E2E model 150 by permitting the model 150 to be fed the identified language for activating appropriate language-specific adaptor modules 300 (FIGS. 2A, 2B, and 3) of an encoder network 210. In doing so, the language-specific adaptor models 218 may capture quality benefits of fine-tuning the model 150 on each individual language).
Regarding Claim 11:
Kannan teaches a system comprising: data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising(Para [0058], Ln 1-15, The processor 510 can process instructions for execution within the computing device 500, including instructions stored in the memory 520 or on the storage device):
obtaining a base automatic speech recognition (ASR) model trained on non-biased data(Para [0011], Ln 10-18, training the multilingual E2E speech recognition model on a union of all the training data sets);
obtaining a sub-model trained on biased data, the biased data associated with a respective domain(Para [0011], Ln 27-33, during the second stage….for each of the one or more adaptor modules that are specific to the particular native language, learning values for a respective set of weights by training the multilingual E2E speech recognition model only on the training data set that is associated with the respective particular native language);
receiving a speech recognition request comprising audio data characterizing a spoken utterance(Abstract, Ln 1-5, speech recognition model includes receiving audio data for an utterance spoken in a particular native language);
processing, using the base ASR model, the audio data to generate a first speech recognition result of the spoken utterance(Para [0011], Ln 10-18, training the multilingual E2E speech recognition model on a union of all the training data sets. Para [0045], Ln 1-7, Furthermore, adapter modules 300 do not need to be employed when they are not helpful. For example, in practice, adapter modules may only be selected or used when they are most effective, while setting all other adapters to identity operations, i.e., setting the weights to zeroes. Para [0038], Ln 1-9, processes the input vector 119 associated with the utterance 106a spoken in English by passing the input vector 119 through the first LSTM layer 216a, passing an output from the first LSTM layer 216a through the appropriate adaptor module 300c for English utterance. Para [0046], Ln 1-5, effectiveness of adapters on each language is determined by measuring the word error rate (WER) of both the model augmented with language-specific adapters and the model without adapters on a test set in that language);
processing, using the sub-model, the audio data to generate a second speech recognition result of the spoken utterance, the second speech recognition result biased toward one or more terms in the respective domain associated with the biased data(Para [0038], Ln 1-9, processes the input vector 119 associated with the utterance 106a spoken in English by passing the input vector 119 through the first LSTM layer 216a, passing an output from the first LSTM layer 216a through the appropriate adaptor module 300c for English utterance. Para [0029], Ln 1-14, language vector 115 serves as an extension to the multilingual E2E model 150 by permitting the model 150 to be fed the identified language for activating appropriate language-specific adaptor modules 300 (FIGS. 2A, 2B, and 3) of an encoder network 210. In doing so, the language-specific adaptor models 218 may capture quality benefits of fine-tuning the model 150 on each individual language);
Kannan does not teach and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen.
In the same field of Speech Recognition, Sugaya teaches and displaying, by a user interface generator, the first speech recognition result and the second speech recognition results on a screen(Para [0065], Ln 1-5, The user terminal receives the recognition result data and displays the first recognition text and the second recognition text on its display unit based on the recognition result data. Abstract, Ln 1-14, The computer system acquires voice data; performs voice recognition for the acquired voice data; performs voice recognition for the acquired voice data with an algorithm or a database different from that used by the first recognition unit; and outputs both of the recognition results when the recognition results from the voice recognitions are different.).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify Kannan with the user feedback on recognition of Sugaya, as it provides feedback to the system allowing it to improve the quality of the results(Para [0069], Ln 1-13, & Para [0071], Ln 1-12).
Regarding Claim 12:
Claim 12 contains similar limitations as Claim 2 and is therefore rejected for the same reasons.
Regarding Claim 13:
Claim 13 contains similar limitations as Claim 3 and is therefore rejected for the same reasons.
Regarding Claim 14:
Claim 14 contains similar limitations as Claim 4 and is therefore rejected for the same reasons.
Regarding Claim 15:
Claim 15 contains similar limitations as Claim 5 and is therefore rejected for the same reasons.
Regarding Claim 16:
Claim 16 contains similar limitations as Claim 6 and is therefore rejected for the same reasons.
Regarding Claim 19:
Claim 19 contains similar limitations as Claim 9 and is therefore rejected for the same reasons.
Regarding Claim 20:
Claim 20 contains similar limitations as Claim 10 and is therefore rejected for the same reasons.
Claim(s) 7 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Kannan and Sugaya as applied to claim 1 above, and further in view of Deshmukh et al. (US 20210073468 A1).
Regarding Claim 7:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, but does not teach wherein the biased data associated with the respective domain comprises words or phrases corresponding to a software application associated with the respective domain.
In the same field of Speech Recognition, Deshmukh teaches wherein the biased data associated with the respective domain comprises words or phrases corresponding to a software application associated with the respective domain(Para [0028], Ln 16-22, Example product domains may include….a software application, and the like. Para [0019], Ln 1-10, digital assistant used for a multitude of domains may be referred to as an enterprise digital assistant (EDA). Domains (e.g., domains of the enterprise) may include a multiple of domains, such as a function domain, a business domain, an organization domain, a product domain. Para [0049], Ln 1-10, a plurality of terms may be received from the user. The plurality of terms may be spoken words, text, or a combination of spoken words and text. The terms may be provided in the natural language of the user. If the terms received are spoken words, the terms may be converted to text via a speech to text processing. Para [0050], Ln 1-22, one or more domain levels…may be identified. The one or more domain levels (e.g., domain layers) may be associated with the user and/or an entity related to the user. Each domain level may have a bias adjustable for the organization. The bias (e.g., adjustable bias) may be used as a weight for the confidence score of the transcript produced by the level. One or more (e.g., each) domain level may have a machine learning model trained with speech recordings. The speech recordings may be…..domain specific. One or more of the domain specific transcription models may be transcribed. The transcription and/or the confidence score may be saved…..The transcript may be provided to the intent processing component, for example, for determining of the intent of the user. Para [0052], Ln 1-8, One or more (e.g., each) domain level may have a machine learning model, such as a supervised or unsupervised machine learning model, that may be trained with training data. The training data may be specific to a domain).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Kannan and Sugaya, with the domains of Deshmukh, as it provides an application for the adaptor system of Kannan to be applied to, and improve the efficiency of(Deshmukh Para [0005], Ln 1-10, a function domain, a business domain, an organization domain, or a product domain. Kannan Para [0044], Ln 1-8, This multi-stage training results in a single model having a small number of parameters that are specific to each language, which are typically less than 10% of the original model size. Such enhancement allows for efficient parameter sharing across languages).
Regarding Claim 17:
Claim 17 contains similar limitations as Claim 7 and is therefore rejected for the same reasons.
Claim(s) 8 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Kannan and Sugaya as applied to claim 1 above, and further in view of Thomas et al. (US 20220319494 A1).
Regarding Claim 8:
The combination of Kannan and Sugaya teaches the computer-implemented method of claim 1, but does not teach wherein the biased data associated with the respective domain comprises words or phrases corresponding to a geographical location associated with the respective domain.
In the same field of Speech Recognition, Thomas teaches wherein the biased data associated with the respective domain comprises words or phrases corresponding to a geographical location associated with the respective domain(Para [0014], Ln 1-10, pre-trained ASR model is provided. The pretrained general purpose ASR model is trained with general purpose audio and verbatim transcription data. The general purpose ASR model can be tuned to a domain using audio data containing utterances in the specific domain and verbatim transcriptions for the utterances. Para [0044], Ln 1-6, Output 301 can be received by a chatbot or a question answer program (not pictured). In some embodiments, output 310 may be used to direct individuals to the correct geographic location or to answer questions relating to financial, medical, or entertainment domains).
It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Kannan and Sugaya with the domains of Thomas, as it as it provides an application for the adaptor system of Kannan to be applied to, and improve the efficiency of(Thomas, Para [0044], Ln 1-6, In some embodiments, output 310 may be used to direct individuals to the correct geographic location or to answer questions relating to financial, medical, or entertainment domains. Kannan, Para [0044], Ln 1-8, This multi-stage training results in a single model having a small number of parameters that are specific to each language, which are typically less than 10% of the original model size. Such enhancement allows for efficient parameter sharing across languages).
Regarding Claim 18:
Claim 18 contains similar limitations as Claim 8 and is therefore rejected for the same reasons.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Lee et al. (US 20220319500 A1)
Speech recognition domain adaptation using adaptor layers.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER G MARLOW whose telephone number is (571)272-4536. The examiner can normally be reached Monday - Thursday 10:00 am - 8:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richmond Dorvil can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEXANDER G MARLOW/ Assistant Examiner, Art Unit 2658
/RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658