DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
1. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
2. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
3. Claims 1-8, 10-18 & 20 are rejected under 35 U.S.C. 103 as being unpatentable over Gruber et al. (US 20150348551 A1 hereinafter, Gruber ‘551) in view of Jaber (US 20210343277 A1 hereinafter, Jaber ‘277).
Regarding claim 1; Gruber ‘551 discloses a contextual end-to-end automatic speech recognition (ASR) system (Fig. 7A, Digital Assistant System 700 i.e. STT processing module 730 can include one or more ASR systems. Paragraphs 0201-0214),
comprising: a virtual assistant (Fig. 1, System 100 i.e. Fig. 1 illustrates a block diagram of system 100 wherein system 100 can implement a digital assistant. The terms “digital assistant,” “virtual assistant,” “intelligent automated assistant,” or “automatic digital assistant” can refer to any information processing system that interprets natural language input in spoken and/or textual form to infer user intent, and performs actions based on the inferred user intent. Paragraph 0036) receiving a speech signal (i.e. At block 802, speech input can be received from a user.). Paragraph 0243) comprising a dictation (i.e. Some speech input can also mix commands with dictation, as in the example of dictating a message to be sent while requesting other actions in the same utterance. Paragraph 0245) and a command (i.e. The speech input can be directed to a virtual assistant, and can include one actionable command (e.g., “What's the weather going to be today?”) or multiple actionable commands (e.g., “Navigate to Jessica's house and send her a message that I'm on my way.”). Paragraph 0243);
and an ASR module producing an output transcription for the speech signal (i.e. Any of a variety of speech transcription approaches can be used. In addition, in some examples multiple possible transcriptions can be generated and processed in sequence or simultaneously to identify the best possible match (e.g., the most likely match). Paragraph 0246), the output transcription including a first token indicating a start of the command, and a second token indicating an end of the command (i.e. Based on the positions of the identified domain keywords, user speech 920 can be split into first candidate substring 922 beginning with the keyword “Navigate” and ending prior to the conjunction “and,” and second candidate substring 924 beginning with the keyword “send” and ending with the end of the string. Paragraph 0246-0249).
Examiner reasonably believes that Gruber ‘551 at Paragraph 0217 discloses wherein the ASR system learns to predict when the command is going to be uttered based on the first token or when the command has been uttered based on the second token. Here, Gruber ‘551 describes wherein in some examples, one of the ranked candidate pronunciations can be selected as a predicted pronunciation (e.g., the most likely pronunciation). However, Examiner cites Jaber ‘277 to cure any deficiencies of Gruber ‘551.
Jaber ‘277 discloses wherein the ASR system learns to predict when the command is going to be uttered based on the first token or when the command has been uttered based on the second token (i.e. After processing the audio data, the ASR model 206 can return the results to the inference service 202, and the inference service 202 can trigger the command included in the utterance provided by the user. The entity prediction model 208 performs entity and domain or class prediction of one or more portions of the utterance. Paragraph 0053).
Gruber ‘551 and Jaber ‘277 are combinable because they are from same field of endeavor of speech systems (Jaber ‘277 at “Technical Field”).
Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to modify the speech system as taught by Gruber ‘551 by adding wherein the ASR system learns to predict when the command is going to be uttered based on the first token or when the command has been uttered based on the second token as taught by Jaber ‘277. The motivation for doing so would have been advantageous to modernize automatic speech recognition (ASR) systems because of their ease to train when it comes to the recognition of standard language. This would enhance an ASR system and make it usable in applications that require accurate recognition of many out-of-vocabulary named entities. Therefore, it would have been obvious to combine Gruber ‘551 with Jaber ‘277 to obtain the invention as specified.
Regarding claim 2; Gruber ‘551 discloses a first attention mechanism (Fig. 8, Block 808) receiving the first token and determining whether the first token is suitable to be transcribed at a specific moment of an ongoing transcription (i.e. Referring again to process 800 of Fig. 8, at block 808, a first probability that the first candidate substring corresponds to a first actionable command and a second probability that the second candidate substring corresponds to a second actionable command can be determined. In some examples, each candidate substring can be analyzed to determine a probability that it corresponds to a valid, actionable command. Paragraph 0252)
Regarding claim 3; Gruber ‘551 discloses a second attention mechanism (Fig. 8, Block 810) producing prefix penalties and restricting, based on the prefix penalties, the first attention mechanism to only entries fitting a current transcription context (i.e. At block 810, a determination can be made as to whether the probabilities determined at block 808 exceed a threshold. For example, a minimum threshold can be determined below which a particular parse can be deemed unlikely or unacceptable. In some examples, the probability of each candidate substring corresponding to an actionable command can be compared to a threshold level, and the failure of any candidate substring of a parse to meet the level can be deemed a failure of the parse leading to the “no” branch of block 810. Paragraph 0261)
Regarding claim 4; Gruber ‘551 discloses wherein the ASR system extends the prefix penalties from only working intra-command to bookending any occurrence of the command in training data (i.e. Based on the positions of the identified domain keywords, user speech 920 can be split into first candidate substring 922 beginning with the keyword “Navigate” and ending prior to the conjunction “and,” and second candidate substring 924 beginning with the keyword “send” and ending with the end of the string. Paragraph 0246-0249).
Regarding claim 5; Jaber ‘277 discloses a label encoder encoding a current state of a transcription (i.e. Domain labels provide the inference service 202 with better understanding of which actions to perform. For example, if an utterance is determined by the ASR model 206 and the entity prediction model 208 to include “I WANT TO BUY TICKETS FOR MICHAEL JACKSON IN LAS VEGAS,” portions of the utterance can be labeled with domains to provide better contextual understanding for the utterance and provide applications to perform the command. Paragraph 0054)
wherein the prefix penalties are produced by the second attention mechanism at least in part based on the encoded current state of the transcription (i.e. As the ASR model 206 decodes portions of the utterance, the portions are provided to the entity prediction model 208 to identify the sets of nonoverlapping spans of entities and their domains. Once a label is predicted from preceding portions of an utterance using the entity prediction model 208, a DAFSA 212 is used to constrain the following ASR output to the content of the DAFSA 212. For example, the entity prediction model 208 can determine that an utterance portion or entity span could be associated with a <PLACE> domain label. In that case, a <PLACE> DAFSA 212 is accessed in the knowledge base and traversed to provide one or more candidates for an out-of-vocabulary word or phrase. Paragraph 0056)
Regarding claim 6; Jaber ‘277 discloses wherein the ASR system masks any non-fitting entry between the first token and the second token (i.e. Masking named entities in the training data, such as replacing the training data by a standard token, can be used to maximize the learning rate of the entity prediction model 508 per sample and to prevent overfitting. Paragraph 0085)
Regarding claim 7; Gruber ‘551 discloses wherein the ASR system produces prefix penalties to mask all command entries until the first token is predicted (i.e. Based on the positions of the identified domain keywords, user speech 920 can be split into first candidate substring 922 beginning with the keyword “Navigate” and ending prior to the conjunction “and,” and second candidate substring 924 beginning with the keyword “send” and ending with the end of the string. Paragraph 0246-0249).
Regarding claim 8; Gruber ‘551 discloses wherein the ASR system enables attention to the command after the first token is predicted and until the second token (i.e. Based on the positions of the identified domain keywords, user speech 920 can be split into first candidate substring 922 beginning with the keyword “Navigate” and ending prior to the conjunction “and,” and second candidate substring 924 beginning with the keyword “send” and ending with the end of the string. Paragraph 0246-0249).
Regarding claim 10; Gruber ‘551 discloses wherein the command is a multi-word command (i.e. Fig. 9 illustrates an exemplary parsed multi-part voice command. User speech 920 can include the transcription of a single utterance saying, “Navigate to Jessica's house and send her a message that I'm on my way.” Paragraph 0248)
Regarding claim 11; Claim 11 contains substantially the same subject matter as claim 1. Therefore, claim 11 is rejected on the same grounds as claim 1.
Regarding claim 12; Claim 12 contains substantially the same subject matter as claim 2. Therefore, claim 12 is rejected on the same grounds as claim 2.
Regarding claim 13; Claim 13 contains substantially the same subject matter as claim 3. Therefore, claim 13 is rejected on the same grounds as claim 3.
Regarding claim 14; Claim 14 contains substantially the same subject matter as claim 4. Therefore, claim 14 is rejected on the same grounds as claim 4.
Regarding claim 15; Claim 15 contains substantially the same subject matter as claim 5. Therefore, claim 15 is rejected on the same grounds as claim 5.
Regarding claim 16; Claim 16 contains substantially the same subject matter as claim 6. Therefore, claim 16 is rejected on the same grounds as claim 6.
Regarding claim 17; Claim 17 contains substantially the same subject matter as claim 7. Therefore, claim 17 is rejected on the same grounds as claim 7.
Regarding claim 18; Claim 18 contains substantially the same subject matter as claim 8. Therefore, claim 18 is rejected on the same grounds as claim 8.
Regarding claim 20; Claim 20 contains substantially the same subject matter as claim 10. Therefore, claim 20 is rejected on the same grounds as claim 10.
4. Claims 9 & 19 are rejected under 35 U.S.C. 103 as being unpatentable over Gruber ‘551 with Jaber ‘277 and further in view of Grosberg (US 20210165566 A1 hereinafter, Grosberg ‘566).
Regarding claim 9; Grosberg ‘566 as modified does not expressly disclose the limitation as expressed below.
Grosberg ‘566 discloses wherein the virtual assistant assists a doctor to perform verbal commands during a doctor-patient encounter (i.e. Tablet 102 may also optionally feature reporting or reporting support functions, such as a microphone (not shown) and the ability to receive dictation from the doctor or other user. Tablet 102 would receive dictation from the user and would either perform speech-to-text conversion. Optionally, tablet 102 would only act to receive the voice data and would send such data to standalone computer 110; optionally and preferably translation module 118 receives such data, along with a command indicating that it is voice data, and then forwards the data to the appropriate application (whether as a voice command that needs to be translated into an action or for voice to text conversion). Paragraph 0054)
Gruber ‘551 and Grosberg ‘566 are combinable because they are from same field of endeavor of speech systems (Grosberg ‘566 at “Field of the Invention”).
Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to modify the speech system as taught by Gruber ‘551 by adding the limitation as taught by Grosberg ‘566. The motivation for doing so would have been advantageous to provide more and better options for users of standalone computers, other than the standard keyboard and mouse, for controlling computers through input devices. Therefore, it would have been obvious to combine Gruber ‘551 with Grosberg ‘566 to obtain the invention as specified.
Regarding claim 19; Claim 19 contains substantially the same subject matter as claim 9. Therefore, claim 19 is rejected on the same grounds as claim 9.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARCUS T. RILEY, ESQ. whose telephone number is (571)270-1581. The examiner can normally be reached 9-5 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MARCUS T. RILEY, ESQ.
Primary Examiner
Art Unit 2654
/MARCUS T RILEY/Primary Examiner, Art Unit 2654