DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 2 – 8, 11 – 18 and 21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a computer-implemented method comprising: receiving a voice query; identifying a first transcription based at least in part on the voice query, wherein the first transcription includes one or more difficult to pronounce terms; determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors; generating a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors; causing submission of at least a portion of the second transcription as a search query to an electronic content search system; and causing output of one or more content search results of the search query.
The claim 2 limitations, under their broadest reasonable interpretation, cover performance of the limitations in the mind but for the recitation of generic computer components. That is, other than reciting “a computer-implemented method”, nothing in the claim elements preclude the actions from practically being performed in the mind. For example, “receiving” in the context of this claim encompasses a person hearing a voice query, “identifying” in the context of this claim encompasses a person reading a transcription of the voice query, “determining” in the context of this claim encompasses a person determining that transcription errors in the transcription of the voice query correspond to difficult to pronounce terms, and “generating” in the context of this claim encompasses a person writing a corrected transcription of the voice query. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites the additional element “a computer-implemented method”. The additional element amounts to no more than mere instructions to apply the exception using generic computer components. Examples of generic computer components can be found in paragraph 0031 of the specification, “Control circuitry 304 may be based on any suitable processing circuitry such as processing circuitry 306. As referred to herein, processing circuitry should be understood to mean circuitry based on one or more microprocessors, microcontrollers, digital signal processors, programmable logic devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., and may include a multi-core processor (e.g., dual-core, quad-core, hexa-core, or any suitable number of cores).”. Accordingly, the additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The additional limitations “causing submission of at least a portion of the second transcription as a search query to an electronic content search system” and “causing output of one or more content search results of the search query” are recited at a high level of generality and amounts to no more than insignificant post-solution activity that is incidental to the primary process of the claim. The addition of well-understood or conventional insignificant extra-solution activity does not amount to an inventive concept. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Claims 3 – 5 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 3 – 5 depend from claim 2, and thus recites the limitations of claim 2.
For the reasons discussed above for claim 2, the claim 2 limitations recite abstract ideas. The additional limitations of claims 3 – 5 do not preclude the steps of claim 2 from practically being performed in the mind. For example, a person using the method of claim 2 to identify transcription errors of difficult to pronounce terms and correct the transcription could also perform the limitations of claims 3 – 5:
Claim 3: A person could compare difficult to pronounce terms to a stored list of terms, determine a pronunciation difficulty metric for the difficult to pronounce terms, and determine that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
Claim 4: A person could compare syllables or letter combinations of difficult to pronounce terms to syllables or letter combinations of each term of a stored list of terms.
Claim 5: A person could determine difficult to pronounce terms that have an identical or similar audible sound as one or more other terms and have a different meaning than the one or more other terms.
The claims do not integrate the judicial exception into a practical application. For the reasons discussed above for claim 2, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Accordingly, these elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. For the reasons discussed above for claim 2, mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claims are not patent eligible.
Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 6 depends from claim 2, and thus recites the limitations of claim 2. For the reasons discussed above for claim 2, the claim 2 limitations recite abstract ideas. The additional element of claim 6 of using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query amounts to no more than mere instructions to apply the exception using generic computer components. There are no details about a particular machine learning model or how the machine learning model operates to determine that difficult to pronounce terms were mispronounced in the voice query. The machine learning model is used to generally apply an abstract idea (determining that difficult to pronounce terms were mispronounced in a voice query) without placing any limitation on how the machine learning model operates to make the determination. The limitation recites only the idea of determining that difficult to pronounce terms were mispronounced in a voice query using a machine learning model without details on how this is accomplished. The claim omits any details as to how the machine learning solves a technical problem, and instead recites only the idea of a solution or outcome. The claim invokes a machine learning merely as a tool for determining that difficult to pronounce terms were mispronounced in a voice query rather than purporting to improve the technology or a computer (See MPEP 2106.05(f)). Therefore, the limitation represents no more than mere instructions to apply the judicial exception on a computer.
The claims do not integrate the judicial exception into a practical application. For the reasons discussed above for claim 2, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Accordingly, these elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. For the reasons discussed above for claim 2, mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claims are not patent eligible.
Claims 7 – 8 and 11 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 7 – 8 and 11 depend from claim 2, and thus recites the limitations of claim 2.
For the reasons discussed above for claim 2, the claim 2 limitations recite abstract ideas. The additional limitations of claims 7 – 8 and 11 do not preclude the steps of claim 2 from practically being performed in the mind. For example, a person using the method of claim 2 to identify transcription errors of difficult to pronounce terms and correct the transcription could also perform the limitations of claims 7 – 8 and 11:
Claim 7: A person could determine that a transcription has transcription errors that correspond to difficult to pronounce terms included in transcription.
Claim 8: A person could write content returned by an electronic content search system for a transcription and determine that a transcription has transcription errors based on the relevance of the content returned by an electronic content search system.
Claim 11: A person could determine that a transcription has transcription errors based on the content returned by an electronic content search system for a corrected transcription of a search query not being different than the content returned by an electronic content search system for the original transcription.
The claims do not integrate the judicial exception into a practical application. For the reasons discussed above for claim 2, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Accordingly, these elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. For the reasons discussed above for claim 2, mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claims are not patent eligible.
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites a system comprising: control circuitry configured to: receive a voice query; identify a first transcription based at least in part on the voice query, wherein the first transcription includes one or more difficult to pronounce terms; determine that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors; generate a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors; cause submission of at least a portion of the second transcription as a search query to an electronic content search system; and cause output of one or more content search results of the search query.
The claim 12 limitations, under their broadest reasonable interpretation, cover performance of the limitations in the mind but for the recitation of generic computer components. That is, other than reciting “a system” and “control circuitry”, nothing in the claim elements preclude the actions from practically being performed in the mind. For example, “receive” in the context of this claim encompasses a person hearing a voice query, “identify” in the context of this claim encompasses a person reading a transcription of the voice query, “determine” in the context of this claim encompasses a person determining that transcription errors in the transcription of the voice query correspond to difficult to pronounce terms, and “generate” in the context of this claim encompasses a person writing a corrected transcription of the voice query. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites the additional elements “a system” and “control circuitry”. The additional elements amount to no more than mere instructions to apply the exception using generic computer components. Examples of generic computer components can be found in paragraph 0031 of the specification, “Control circuitry 304 may be based on any suitable processing circuitry such as processing circuitry 306. As referred to herein, processing circuitry should be understood to mean circuitry based on one or more microprocessors, microcontrollers, digital signal processors, programmable logic devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., and may include a multi-core processor (e.g., dual-core, quad-core, hexa-core, or any suitable number of cores).”. Accordingly, the additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The additional limitations “cause submission of at least a portion of the second transcription as a search query to an electronic content search system” and “cause output of one or more content search results of the search query” are recited at a high level of generality and amounts to no more than insignificant post-solution activity that is incidental to the primary process of the claim. The addition of well-understood or conventional insignificant extra-solution activity does not amount to an inventive concept. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Claims 13 – 15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 13 – 15 depend from claim 12, and thus recites the limitations of claim 12.
For the reasons discussed above for claim 12, the claim 12 limitations recite abstract ideas. The additional limitations of claims 13 – 15 do not preclude the steps of claim 12 from practically being performed in the mind. For example, a person using the method of claim 12 to identify transcription errors of difficult to pronounce terms and correct the transcription could also perform the limitations of claims 13 – 15:
Claim 13: A person could compare difficult to pronounce terms to a stored list of terms, determine a pronunciation difficulty metric for the difficult to pronounce terms, and determine that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
Claim 14: A person could compare syllables or letter combinations of difficult to pronounce terms to syllables or letter combinations of each term of a stored list of terms.
Claim 15: A person could determine difficult to pronounce terms that have an identical or similar audible sound as one or more other terms and have a different meaning than the one or more other terms.
The claims do not integrate the judicial exception into a practical application. For the reasons discussed above for claim 12, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Accordingly, these elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. For the reasons discussed above for claim 12, mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claims are not patent eligible.
Claim 16 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 16 depends from claim 12, and thus recites the limitations of claim 12. For the reasons discussed above for claim 12, the claim 12 limitations recite abstract ideas. The additional element of claim 16 of using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query amounts to no more than mere instructions to apply the exception using generic computer components. There are no details about a particular machine learning model or how the machine learning model operates to determine that difficult to pronounce terms were mispronounced in the voice query. The machine learning model is used to generally apply an abstract idea (determining that difficult to pronounce terms were mispronounced in a voice query) without placing any limitation on how the machine learning model operates to make the determination. The limitation recites only the idea of determining that difficult to pronounce terms were mispronounced in a voice query using a machine learning model without details on how this is accomplished. The claim omits any details as to how the machine learning solves a technical problem, and instead recites only the idea of a solution or outcome. The claim invokes a machine learning merely as a tool for determining that difficult to pronounce terms were mispronounced in a voice query rather than purporting to improve the technology or a computer (See MPEP 2106.05(f)). Therefore, the limitation represents no more than mere instructions to apply the judicial exception on a computer.
The claims do not integrate the judicial exception into a practical application. For the reasons discussed above for claim 12, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Accordingly, these elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. For the reasons discussed above for claim 12, mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claims are not patent eligible.
Claims 17 – 18 and 21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 17 – 18 and 21 depend from claim 12, and thus recites the limitations of claim 12.
For the reasons discussed above for claim 12, the claim 12 limitations recite abstract ideas. The additional limitations of claims 17 – 18 and 21 do not preclude the steps of claim 12 from practically being performed in the mind. For example, a person using the method of claim 12 to identify transcription errors of difficult to pronounce terms and correct the transcription could also perform the limitations of claims 17 – 18 and 21:
Claim 17: A person could determine that a transcription has transcription errors that correspond to difficult to pronounce terms included in transcription.
Claim 18: A person could write content returned by an electronic content search system for a transcription and determine that a transcription has transcription errors based on the relevance of the content returned by an electronic content search system.
Claim 21: A person could determine that a transcription has transcription errors based on the content returned by an electronic content search system for a corrected transcription of a search query not being different than the content returned by an electronic content search system for the original transcription.
The claims do not integrate the judicial exception into a practical application. For the reasons discussed above for claim 12, the additional elements amount to no more than mere instructions to apply the exception using generic computer components and insignificant extra-solution activity. Accordingly, these elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. For the reasons discussed above for claim 12, mere instructions to apply an exception using generic computer components and insignificant extra-solution activity cannot provide an inventive concept. The claims are not patent eligible.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 5, 9 – 11, 15, 19 – 21 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claims contain subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Regarding claim 5, the disclosure does not provide adequate support for the claim limitation “wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms” because the specification does not disclose difficult to pronounce terms having an identical or similar audible sound as one or more other terms and a different meaning than the one or more other terms. The specification recites, in paragraph 0010, lines 1-9, “In one embodiment, occurrence of incorrect transcriptions may be determined according to any one or more of a number of factors, including whether a transcription score or metric of the transcription is less than some predetermined value or values, whether one or more terms of the transcription have a homonym and are thus likely to be the incorrect homonym, whether one or more terms of the transcription are difficult to pronounce (e.g., have a pronunciation difficulty metric greater than some predetermined value or values), whether the audio input was excessively noisy, whether one or more terms is mispronounced, or whether the transcription is a single-word transcription or a transcription containing multiple words.”, disclosing determining the occurrence of incorrect transcriptions according to whether terms of the transcription have a homonym and whether terms of the transcription are difficult to pronounce, but not disclosing difficult to pronounce terms having an identical or similar audible sound as one or more other terms and a different meaning than the one or more other terms. The introduction of claim changes which involve narrowing the claims by introducing elements or limitations which are not supported by the as-filed disclosure is a violation of the written description requirement of 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph (see MPEP § 2163.05, subsection II).
Regarding claim 9, the disclosure does not provide adequate support for the claim limitation “wherein the one or more hint words are determined based at least in part on the one or more difficult to pronounce terms” because the specification does not disclose determining hint words based on difficult to pronounce terms. The specification recites, in paragraph 0025, lines 1-2, “In response, system 100 repeatedly attempts new translations of the audio request, with each attempt using hint words derived from different terms of the improper transcription.”, disclosing determining hint words based on terms of an improper transcription, but not disclosing determining hint words based on difficult to pronounce terms. The introduction of claim changes which involve narrowing the claims by introducing elements or limitations which are not supported by the as-filed disclosure is a violation of the written description requirement of 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph (see MPEP § 2163.05, subsection II).
Regarding claim 10, claim 10 depends from claim 9, and thus recites the limitations of claim 9, and therefore contain subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, at the time the application was filed, had possession of the claimed invention.
Regarding claim 11, the disclosure does not provide adequate support for the claim limitation "based at least in part on determining that the second content results are not different from first content results resulting from a first search query to the electronic content search system for the first transcription, determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors" because the specification does not disclose determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors based on determining that the second content results are not different from first content results resulting from a first search query to the electronic content search system for the first transcription. The specification recites, in paragraph 0061, lines 1-6, “Alternatively, if the error determination module 518 determines that the transcription was inaccurate or incorrect, e.g., the transcription has a low accuracy score, several difficult-to-pronounce terms and many potential homonyms, module 518 may instruct ASR server 220 to conduct ASR correction methods of embodiments of the disclosure, e.g., as in Step 620 above, performing successive submission of transcription terms and/or related terms as hint words to the ASR module 414 (Step 730) until the resulting transcription changes.”, disclosing determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors, but not disclosing determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors based on determining that the second content results are not different from first content results resulting from a first search query to the electronic content search system for the first transcription. The introduction of claim changes which involve narrowing the claims by introducing elements or limitations which are not supported by the as-filed disclosure is a violation of the written description requirement of 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph (see MPEP § 2163.05, subsection II).
Regarding claim 15, the disclosure does not provide adequate support for the claim limitation “wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms” because the specification does not disclose difficult to pronounce terms having an identical or similar audible sound as one or more other terms and a different meaning than the one or more other terms. The specification recites, in paragraph 0010, lines 1-9, “In one embodiment, occurrence of incorrect transcriptions may be determined according to any one or more of a number of factors, including whether a transcription score or metric of the transcription is less than some predetermined value or values, whether one or more terms of the transcription have a homonym and are thus likely to be the incorrect homonym, whether one or more terms of the transcription are difficult to pronounce (e.g., have a pronunciation difficulty metric greater than some predetermined value or values), whether the audio input was excessively noisy, whether one or more terms is mispronounced, or whether the transcription is a single-word transcription or a transcription containing multiple words.”, disclosing determining the occurrence of incorrect transcriptions according to whether terms of the transcription have a homonym and whether terms of the transcription are difficult to pronounce, but not disclosing difficult to pronounce terms having an identical or similar audible sound as one or more other terms and a different meaning than the one or more other terms. The introduction of claim changes which involve narrowing the claims by introducing elements or limitations which are not supported by the as-filed disclosure is a violation of the written description requirement of 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph (see MPEP § 2163.05, subsection II).
Regarding claim 19, the disclosure does not provide adequate support for the claim limitation “wherein the one or more hint words are determined based at least in part on the one or more difficult to pronounce terms” because the specification does not disclose determining hint words based on difficult to pronounce terms. The specification recites, in paragraph 0025, lines 1-2, “In response, system 100 repeatedly attempts new translations of the audio request, with each attempt using hint words derived from different terms of the improper transcription.”, disclosing determining hint words based on terms of an improper transcription, but not disclosing determining hint words based on difficult to pronounce terms. The introduction of claim changes which involve narrowing the claims by introducing elements or limitations which are not supported by the as-filed disclosure is a violation of the written description requirement of 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph (see MPEP § 2163.05, subsection II).
Regarding claim 20, claim 20 depends from claim 19, and thus recites the limitations of claim 19, and therefore contain subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, at the time the application was filed, had possession of the claimed invention.
Regarding claim 21, the disclosure does not provide adequate support for the claim limitation "based at least in part on determining that the second content results are not different from first content results resulting from a first search query to the electronic content search system for the first transcription, determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors" because the specification does not disclose determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors based on determining that the second content results are not different from first content results resulting from a first search query to the electronic content search system for the first transcription. The specification recites, in paragraph 0061, lines 1-6, “Alternatively, if the error determination module 518 determines that the transcription was inaccurate or incorrect, e.g., the transcription has a low accuracy score, several difficult-to-pronounce terms and many potential homonyms, module 518 may instruct ASR server 220 to conduct ASR correction methods of embodiments of the disclosure, e.g., as in Step 620 above, performing successive submission of transcription terms and/or related terms as hint words to the ASR module 414 (Step 730) until the resulting transcription changes.”, disclosing determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors, but not disclosing determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors based on determining that the second content results are not different from first content results resulting from a first search query to the electronic content search system for the first transcription. The introduction of claim changes which involve narrowing the claims by introducing elements or limitations which are not supported by the as-filed disclosure is a violation of the written description requirement of 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph (see MPEP § 2163.05, subsection II).
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 18 – 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 18 recites the limitation "the I/O circuitry " in line 1. There is insufficient antecedent basis for this limitation in the claim.
Claim 19 recites the limitation "the I/O circuitry " in line 1. There is insufficient antecedent basis for this limitation in the claim.
Claim 20 recites the limitation "the I/O circuitry " in line 1. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 7, 12 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Bakshi et al. (US Patent No. 10,354,647), hereinafter Bakshi, in view of Komissarchik et al. (US Patent Application Publication No. 2017/0345426), hereinafter Komissarchik, and Bettaglio et al. (US Patent No. 11,263,198), hereinafter Bettaglio.
Regarding claim 2, Bakshi discloses a computer-implemented method comprising: receiving a voice query (Column 3, lines 57-65, "The user devices 106 submit search queries 109 to the search system 120. In some examples, a user device 106 can include one or more input modalities. Example modalities can include a keyboard, a touchscreen and/or a microphone. For example, a user can use a keyboard and/or touchscreen to type in a search query. As another example, a user can speak a search query, the user speech being captured through a microphone, and being processed through speech recognition to provide the search query."; Capturing user speech of a user speaking a search query reads on receiving a voice query.);
identifying a first transcription based at least in part on the voice query (Column 4, lines 56-64, "In some examples, the speech recognition system 130 can process the speech data to provide text. For example, the speech recognition system 130 can process the speech data using a voice-to-text engine (also referred to as the first speech recognition engine) to provide the text. In some examples, the speech recognition system 130 provides the text to the search system 120, which processes the text as a search query to provide search results 112."; A speech recognition system processing speech data of a search query using a voice-to-text engine to provide text reads on identifying a first transcription based at least in part on the voice query.);
generating a second transcription based at least in part on the voice query [and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors] (Column 1, lines 34-50, "In general, innovative aspects of the subject matter described in this specification can be embodied in methods that include actions of providing first text for display on a computing device of a user, the first text being provided from a first speech recognition engine based on first speech received from the computing device, and being displayed as a search query, receiving a speech correction indication from the computing device, the speech correction indication indicating a portion of the first text that is to be corrected, receiving second speech from the computing device, receiving second text from a second speech recognition engine based on the second speech, the second speech recognition engine being different from the first speech recognition engine, replacing the portion of the first text with the second text to provide a combined text, and providing the combined text for display on the computing device as a revised search query."; Receiving second text from a second speech recognition engine based on the second speech reads on generating a second transcription based at least in part on the voice query.);
causing submission of at least a portion of the second transcription as a search query to an electronic content search system (Column 2, lines 55-58, "In some implementations, the portion of the first text is replaced with the second text to provide a combined text. In some examples, the combined text is a revised search query that is submitted to the search system."; Submitting a revised search query to a search system reads on causing submission of at least a portion of the second transcription as a search query to an electronic content search system.);
and causing output of one or more content search results of the search query (Column 1, lines 54-67, "These and other implementations can each optionally include one or more of the following features: the portion includes an entirety of the first text; the portion comprises less than an entirety of the first text; the second speech recognition engine includes the first speech recognition engine and at least one additional function; the at least one additional function includes selecting a potential text as the second text based on one or more entities associated with the first text; actions further include: receiving first search results based on the first text, and providing the first search results for display on the computing device; actions further include: receiving second search results based on the second text, and providing the second search results for display on the computing device in place of the first search results;"; Receiving second search results based on the second text and providing the second search results for display on a computing device reads on causing output of one or more content search results of the search query.).
Bakshi does not specifically disclose:
wherein the first transcription includes one or more difficult to pronounce terms; determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors.
Komissarchik teaches:
wherein the first transcription includes one or more difficult to pronounce terms (Paragraph 0016, lines 1-6, "In accordance with another aspect of the invention the system and methods for automatic feedback are provided to assist users to correct mispronunciation errors and to suggest alternative phrases with the same or similar meaning that are less difficult for user to pronounce correctly that lead to better recognition results."; Paragraph 0026, lines 1-8, "Referring to FIG. 1, system 10 for robust voice-based human-IoT communication is described. System 10 comprises of a number of software modules that cooperate to detect mispronunciations in a user's utterances, to detect systematic speech recognition errors caused by such mispronunciations or ASR deficiencies, provide detailed feedback to the user that enables him or her to achieve better speech recognition results."; Detecting mispronunciations in a user's utterances for phrases that are difficult for user to pronounce reads on the first transcription includes one or more difficult to pronounce terms.);
determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors (Paragraph 0016, lines 1-6, "In accordance with another aspect of the invention the system and methods for automatic feedback are provided to assist users to correct mispronunciation errors and to suggest alternative phrases with the same or similar meaning that are less difficult for user to pronounce correctly that lead to better recognition results."; Paragraph 0026, lines 1-8, "Referring to FIG. 1, system 10 for robust voice-based human-IoT communication is described. System 10 comprises of a number of software modules that cooperate to detect mispronunciations in a user's utterances, to detect systematic speech recognition errors caused by such mispronunciations or ASR deficiencies, provide detailed feedback to the user that enables him or her to achieve better speech recognition results."; Detecting systematic speech recognition errors caused by mispronunciations reads on determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors.).
Komissarchik is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi to incorporate the teachings of Komissarchik to detect mispronunciations in a user's utterances for phrases that are difficult for user to pronounce and detect systematic speech recognition errors caused by mispronunciations. Doing so would allow for detecting what is wrong with a user pronunciation and helping the user to modify his or her pronunciation to achieve better recognition results (Komissarchik; Paragraph 0011, lines 1-5).
Bakshi in view of Komissarchik does not specifically disclose: generating a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors.
Bettaglio teaches:
generating a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors (Column 3, lines 7-13, "A query response system (QRS) receives queries from a user and provides responses to those queries. One embodiment of a QRS is any virtual assistant (or any machine) that assists a user and, in most instances, that the user can control using speech. FIGS. 1A, 1B, 1C, and 1D show example embodiments of speech-enabled virtual assistants according to different embodiments of the invention."; Column 9, lines 17-25, "Furthermore, detected errors can indicate valid pronunciations missing from an ASR phonetic dictionary. The correction process can be used for data analysis, data cleansing, and for real-time query rewrites, all of which can enhance user experience. Furthermore, and in accordance with the aspects of the invention, large high-error-likelihood queries can trigger the QRS to initiate a user dialog to seek additional information in order to allow the QRS to automatically correct the error."; Detecting errors that indicate valid pronunciations are missing from an automatic speech recognition phonetic dictionary reads on determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors, and a query response system automatically correcting an error reads on generating a second transcription.).
Bettaglio is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik to incorporate the teachings of Bettaglio to detect errors that indicate valid pronunciations are missing from an automatic speech recognition phonetic dictionary and implement a query response system automatically correcting an error. Doing so would allow for systematically finding and fixing queries that often result in incorrect automatic speech recognition system errors due to missed conversion of speech to a transcription (Bettaglio; Column 1, lines 34-39).
Regarding claim 7, Bakshi in view of Komissarchik and Bettaglio discloses the method as claimed in claim 2.
Komissarchik further teaches:
further comprising determining that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription (Paragraph 0016, lines 1-6, "In accordance with another aspect of the invention the system and methods for automatic feedback are provided to assist users to correct mispronunciation errors and to suggest alternative phrases with the same or similar meaning that are less difficult for user to pronounce correctly that lead to better recognition results."; Paragraph 0026, lines 1-8, "Referring to FIG. 1, system 10 for robust voice-based human-IoT communication is described. System 10 comprises of a number of software modules that cooperate to detect mispronunciations in a user's utterances, to detect systematic speech recognition errors caused by such mispronunciations or ASR deficiencies, provide detailed feedback to the user that enables him or her to achieve better speech recognition results."; Detecting systematic speech recognition errors caused by mispronunciations reads on determining that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription.).
Komissarchik is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik and Bettaglio to further incorporate the teachings of Komissarchik to detect mispronunciations in a user's utterances for phrases that are difficult for user to pronounce and detect systematic speech recognition errors caused by mispronunciations. Doing so would allow for detecting what is wrong with a user pronunciation and helping the user to modify his or her pronunciation to achieve better recognition results (Komissarchik; Paragraph 0011, lines 1-5).
Regarding claim 12, arguments analogous to claim 2 are applicable. In addition, Bakshi discloses a system (Column 2, lines 59-60, “FIG. 1 depicts an example environment 100 in which a search system provides search results based on user queries.”) comprising: control circuitry (Column 12, lines 33-38, “Implementations of the subject matter and the operations described in this specification can be realized in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.”) configured to perform the steps of claim 2.
Regarding claim 17, arguments analogous to claim 7 are applicable.
Claims 3 – 4 and 13 – 14 are rejected under 35 U.S.C. 103 as being unpatentable over Bakshi in view of Komissarchik and Bettaglio, and further in view of Tanaka et al. (US Patent No. 9,466,291), hereinafter Tanaka.
Regarding claim 3, Bakshi in view of Komissarchik and Bettaglio discloses the method as claimed in claim 2, but does not specifically disclose: wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms; based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms; and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
Tanaka teaches:
wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms.);
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Obtaining a pronunciation difficulty of a retrieval word reads on determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms.);
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold (Column 12, line 49 - Column 13, line 4, "Generally, the lower the pronunciation difficulty of a word, the more accurately the speaker may pronounce the word. Thus, the lower the pronunciation difficulty of the retrieval word, the higher the matching score of the section including the retrieval word in the voice data. On the other hand, the higher the pronunciation difficulty of the retrieval word, the lower the matching score of the section tends to be, even if the section is the one including the retrieval word in the voice data. Therefore, it is estimated that the lower the pronunciation difficulty of the retrieval word, the higher the detection accuracy of the retrieval word. Accordingly, in this embodiment, the threshold setting section 12 increases the score threshold that is the threshold for the matching score for the retrieval word having lower pronunciation difficulty. Thus, when the pronunciation difficulty of the retrieval word is low, the number of candidate sections to be processed by the precise matching section 14 is reduced. As a result, the throughput of the entire voice retrieval processing is reduced. Meanwhile, the threshold setting section 12 may detect the section including the retrieval word as the candidate section, even if the retrieval word is not correctly pronounced, by lowering the score threshold for the retrieval word having higher pronunciation difficulty."; Determining pronunciation difficulty based on a score threshold reads on determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik and Bettaglio to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database, obtain a pronunciation difficulty of a retrieval word, and determine pronunciation difficulty based on a score threshold. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 4, Bakshi in view of Komissarchik and Bettaglio, and further in view of Tanaka, discloses the method as claimed in claim 3.
Tanaka further teaches:
wherein comparing the one or more difficult to pronounce terms to the stored list of terms comprises comparing syllables or letter combinations associated with the one or more difficult to pronounce terms to syllables or letter combinations respectively associated with each term of the stored list of terms (Column 11, line 60 - Column 12, line 5, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word. Note that the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word."; Obtaining a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word, reads on comparing syllables associated with the one or more difficult to pronounce terms to syllables respectively associated with each term of the stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik and Bettaglio, and further in view of Tanaka, to further incorporate the teachings of Tanaka to obtain a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 13, Bakshi in view of Komissarchik and Bettaglio discloses the system as claimed in claim 12, but does not specifically disclose: wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory; based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms; and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
Tanaka teaches:
wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms, and a database reads on a computer memory.);
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Obtaining a pronunciation difficulty of a retrieval word reads on determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms.);
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold (Column 12, line 49 - Column 13, line 4, "Generally, the lower the pronunciation difficulty of a word, the more accurately the speaker may pronounce the word. Thus, the lower the pronunciation difficulty of the retrieval word, the higher the matching score of the section including the retrieval word in the voice data. On the other hand, the higher the pronunciation difficulty of the retrieval word, the lower the matching score of the section tends to be, even if the section is the one including the retrieval word in the voice data. Therefore, it is estimated that the lower the pronunciation difficulty of the retrieval word, the higher the detection accuracy of the retrieval word. Accordingly, in this embodiment, the threshold setting section 12 increases the score threshold that is the threshold for the matching score for the retrieval word having lower pronunciation difficulty. Thus, when the pronunciation difficulty of the retrieval word is low, the number of candidate sections to be processed by the precise matching section 14 is reduced. As a result, the throughput of the entire voice retrieval processing is reduced. Meanwhile, the threshold setting section 12 may detect the section including the retrieval word as the candidate section, even if the retrieval word is not correctly pronounced, by lowering the score threshold for the retrieval word having higher pronunciation difficulty."; Determining pronunciation difficulty based on a score threshold reads on determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik and Bettaglio to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database, obtain a pronunciation difficulty of a retrieval word, and determine pronunciation difficulty based on a score threshold. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 14, arguments analogous to claim 4 are applicable.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Bakshi in view of Komissarchik and Bettaglio, and further in view of Jitkoff et al. (US Patent No. 10,839,805), hereinafter Jitkoff.
Regarding claim 5, Bakshi in view of Komissarchik and Bettaglio discloses the method as claimed in claim 2, but does not specifically disclose: wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms.
Jitkoff teaches:
wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms (Column 1, lines 55-59, "Ambiguous input may be particularly common for spoken input, in part because of the presence of homophones, and in part because a speech-to-text processor may have difficulty differentiating words that are pronounced differently but sound similar to each other."; Column 6, lines 49-53, "In another example, voice input that is identified as being a homophone (multiple terms with the same pronunciation but different meanings) and/or a homonym (multiple terms with the same spelling and pronunciation but different meaning) can be identified as ambiguous."; Difficulty differentiating words that are pronounced differently but sound similar to each other reads on one or more difficult to pronounce terms having an identical or similar audible sound as one or more other terms, and identifying terms with the same pronunciation but different meaning as ambiguous reads on the one or more difficult to pronounce terms having a different meaning than the one or more other terms.).
Jitkoff is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik and Bettaglio to incorporate the teachings of Jitkoff to identify terms that are pronounced differently but sound similar to each other and terms with the same pronunciation but different meaning as ambiguous. Doing so would allow for disambiguating ambiguous user inputs (Jitkoff; Column 1, lines 46-59).
Regarding claim 15, arguments analogous to claim 5 are applicable.
Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Bakshi in view of Komissarchik and Bettaglio, and further in view of Karas (US Patent No. 11,810,471).
Regarding claim 6, Bakshi in view of Komissarchik and Bettaglio discloses the method as claimed in claim 2, but does not specifically disclose: further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query.
Karas teaches:
further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query (Column 9, lines 18-22, "In step 27, a feature of the user's speech that requires direction to more accurately pronounce a particular phoneme is identified. The particular phoneme is a phoneme corresponding to the audio components from which a pattern of mispronunciation is identified in step 26."; Column 9, lines 28-32, "It is to be understood that each of the steps 21 to 28 may be implemented by one or more processors. The one more processors may include memory for storing and reading data. Furthermore, each of the steps 21 to 28 may be implemented in a neural network architecture."; Identifying a pattern of mispronunciation reads on determining that the one or more difficult to pronounce terms were mispronounced in the voice query, and a neural network reads on a machine learning model.).
Karas is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik and Bettaglio to incorporate the teachings of Karas to identify a pattern of mispronunciation using a neural network. Doing so would allow for identifying mistakes in pronunciation and providing feedback to users based on the identified mistakes (Karas; Column 2, lines 11-20).
Regarding claim 16, arguments analogous to claim 6 are applicable.
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Bakshi in view of Komissarchik and Bettaglio, and further in view of Amarilio et al. (US Patent No. 9,336,311), hereinafter Amarilio.
Regarding claim 8, Bakshi in view of Komissarchik and Bettaglio discloses the method as claimed in claim 2.
Bakshi further discloses:
further comprising: causing the first transcription to be submitted to the electronic content search system (Column 2, lines 59-60, "FIG. 1 depicts an example environment 100 in which a search system provides search results based on user queries."; Providing search results based on user queries reads on causing the first transcription to be submitted to the electronic content search system.);
providing for output representations of content returned by the electronic content search system for the first transcription (Column 1, lines 54-67, "These and other implementations can each optionally include one or more of the following features: the portion includes an entirety of the first text; the portion comprises less than an entirety of the first text; the second speech recognition engine includes the first speech recognition engine and at least one additional function; the at least one additional function includes selecting a potential text as the second text based on one or more entities associated with the first text; actions further include: receiving first search results based on the first text, and providing the first search results for display on the computing device; actions further include: receiving second search results based on the second text, and providing the second search results for display on the computing device in place of the first search results;"; Receiving search results based on text and providing the search results for display on a computing device reads on providing for output representations of content returned by the electronic content search system for the first transcription.).
Bakshi in view of Komissarchik and Bettaglio does not specifically disclose: determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value.
Amarilio teaches:
determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value (Column 11, lines 15-19, "The entity provider module 608 determines a relevancy score for each of the entities identified by associated entity identifier 606. The relevancy score is a measure of how relevant the identified entity is to the search query from which the particular entity identifier is derived from."; Column 14, lines 56-60, "If the system determines that the relevancy score does not satisfy a threshold, then the system takes no further action on the first entity identifier. If the system determines that the relevancy score satisfies a threshold (708), then the system provides the second entity (710)."; Determining a relevancy score to measure how relevant the identified entity is to the search query and determining if the relevancy score satisfies a threshold reads on determining whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value.).
Amarilio is considered to be analogous to the claimed invention because it is in the same field of search systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bakshi in view of Komissarchik and Bettaglio to incorporate the teachings of Amarilio to determine a relevancy score to measure how relevant the identified entity is to the search query and determine if the relevancy score satisfies a threshold. Doing so would allow for determining the relevance of entities to search queries and providing information about entities that are relevant to search queries (Amarilio; Column 2, lines 21-25).
Regarding claim 18, arguments analogous to claim 8 are applicable.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 2, 7 – 8, 12 and 17 – 18 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,508,354. Although the claims at issue are not identical, they are not patentably distinct from each other.
Regarding claim 2, claims 1 and 4 of U.S. Patent No. 11,508,354 claim all the limitations set forth in the application claim 2.
US Application No. 19/002,025 Claim 2
U.S. Patent No. 11,508,354 Claim 1
A computer-implemented method comprising:
A method of correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the method comprising:
receiving a voice query;
A method of correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system,
identifying a first transcription based at least in part on the voice query,
for an incorrect transcription output by one or more automated speech recognition systems,
generating a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors;
generating a transcription output by the one or more automated speech recognition systems differs from the incorrect transcription,
causing submission of at least a portion of the second transcription as a search query to an electronic content search system;
submitting the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
and causing output of one or more content search results of the search query.
submitting the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
US Application No. 19/002,025 Claim 2
U.S. Patent No. 11,508,354 Claim 4
wherein the first transcription includes one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors;
determining the presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 7, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 2. Claim 4 of U.S. Patent No. 11,508,354 claims all the additional limitations set forth in the application claim 7.
US Application No. 19/002,025 Claim 7
U.S. Patent No. 11,508,354 Claim 4
The method of claim 2,
The method of claim 2,
further comprising determining that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription.
wherein the determining that the occurrence of the improper search resulted from the incorrect transcription further comprises determining the presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 8, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 2. Claims 1 and 3 of U.S. Patent No. 11,508,354 claims all the additional limitations set forth in the application claim 8.
US Application No. 19/002,025 Claim 8
U.S. Patent No. 11,508,354 Claim 1
The method of claim 2,
A method of correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the method comprising:
further comprising: causing the first transcription to be submitted to the electronic content search system;
submitting the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
providing for output representations of content returned by the electronic content search system for the first transcription;
submitting the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
US Application No. 19/002,025 Claim 8
U.S. Patent No. 11,508,354 Claim 3
determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value.
wherein the determining an occurrence of an improper search further comprises determining the occurrence of an improper search according to one or more of: whether no electronic content was selected by the electronic content search system as a result of the search; whether electronic content selected by the electronic content search system as a result of the search has an associated relevance score less than a predetermined value; whether the terms of the transcription are not connected or weakly connected in a knowledge graph of terms of the electronic content; or whether the terms of the transcription have popularity scores below a predetermined value.
Regarding claim 12, claims 10 and 13 of U.S. Patent No. 11,508,354 claim all the limitations set forth in the application claim 12.
US Application No. 19/002,025 Claim 12
U.S. Patent No. 11,508,354 Claim 10
A system comprising: control circuitry configured to:
A system for correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the system comprising:
receive a voice query;
A system for correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system,
identify a first transcription based at least in part on the voice query,
for an incorrect transcription output by one or more automated speech recognition systems,
generate a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors;
generate a transcription output by the one or more automated speech recognition systems differs from the incorrect transcription,
cause submission of at least a portion of the second transcription as a search query to an electronic content search system;
submit the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
and cause output of one or more content search results of the search query.
submit the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
US Application No. 19/002,025 Claim 2
U.S. Patent No. 11,508,354 Claim 13
wherein the first transcription includes one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors;
determining the presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 17, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 12. Claim 13 of U.S. Patent No. 11,508,354 claims all the additional limitations set forth in the application claim 17.
US Application No. 19/002,025 Claim 17
U.S. Patent No. 11,508,354 Claim 13
The system of claim 12,
The system of claim 11,
wherein the control circuitry is further configured determine that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription.
wherein the determining that the occurrence of the improper search resulted from the incorrect transcription further comprises determining the presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 18, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 12. Claims 10 and 12 of U.S. Patent No. 11,508,354 claims all the additional limitations set forth in the application claim 18.
US Application No. 19/002,025 Claim 18
U.S. Patent No. 11,508,354 Claim 10
The system of claim 12, wherein the I/O circuitry is further configured to:
A system for correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the system comprising:
cause the first transcription to be submitted to the electronic content search system;
submit the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
provide for output representations of content returned by the electronic content search system for the first transcription;
submit the generated transcription to an electronic content search system for selecting electronic content corresponding to the resulting transcription;
US Application No. 19/002,025 Claim 8
U.S. Patent No. 11,508,354 Claim 12
determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value.
wherein the determining an occurrence of an improper search further comprises determining the occurrence of an improper search according to one or more of: whether no electronic content was selected by the electronic content search system as a result of the search; whether electronic content selected by the electronic content search system as a result of the search has an associated relevance score less than a predetermined value; whether the terms of the transcription are not connected or weakly connected in a knowledge graph of terms of the electronic content; or whether the terms of the transcription have popularity scores below a predetermined value.
Claims 3 – 4 and 13 – 14 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,508,354 in view of Tanaka.
Regarding claim 3, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 2. Claim 4 of U.S. Patent No. 11,508,354 claims the additional limitations set forth in the application claim 3, expect for the limitation “wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms”.
US Application No. 19/002,025 Claim 3
U.S. Patent No. 11,508,354 Claim 4
The method of claim 2,
The method of claim 2,
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Tanaka teaches:
wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 4, U.S. Patent No. 11,508,354 in view of Tanaka claims all the limitations set forth in the application claim 3.
Tanaka further teaches:
wherein comparing the one or more difficult to pronounce terms to the stored list of terms comprises comparing syllables or letter combinations associated with the one or more difficult to pronounce terms to syllables or letter combinations respectively associated with each term of the stored list of terms (Column 11, line 60 - Column 12, line 5, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word. Note that the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word."; Obtaining a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word, reads on comparing syllables associated with the one or more difficult to pronounce terms to syllables respectively associated with each term of the stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 in view of Tanaka to further incorporate the teachings of Tanaka to obtain a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 13, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 12. Claim 13 of U.S. Patent No. 11,508,354 claims the additional limitations set forth in the application claim 13, expect for the limitation “wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory”.
US Application No. 19/002,025 Claim 13
U.S. Patent No. 11,508,354 Claim 13
The system of claim 12,
The system of claim 11,
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Tanaka teaches:
wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms, and a database reads on a computer memory.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 14, U.S. Patent No. 11,508,354 in view of Tanaka claims all the limitations set forth in the application claim 13.
Tanaka further teaches:
wherein comparing the one or more difficult to pronounce terms to the stored list of terms comprises comparing syllables or letter combinations associated with the one or more difficult to pronounce terms to syllables or letter combinations respectively associated with each term of the stored list of terms (Column 11, line 60 - Column 12, line 5, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word. Note that the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word."; Obtaining a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word, reads on comparing syllables associated with the one or more difficult to pronounce terms to syllables respectively associated with each term of the stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 in view of Tanaka to further incorporate the teachings of Tanaka to obtain a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Claims 5 and 15 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,508,354 in view of Jitkoff.
Regarding claim 5, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 2.
Jitkoff teaches:
wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms (Column 1, lines 55-59, "Ambiguous input may be particularly common for spoken input, in part because of the presence of homophones, and in part because a speech-to-text processor may have difficulty differentiating words that are pronounced differently but sound similar to each other."; Column 6, lines 49-53, "In another example, voice input that is identified as being a homophone (multiple terms with the same pronunciation but different meanings) and/or a homonym (multiple terms with the same spelling and pronunciation but different meaning) can be identified as ambiguous."; Difficulty differentiating words that are pronounced differently but sound similar to each other reads on one or more difficult to pronounce terms having an identical or similar audible sound as one or more other terms, and identifying terms with the same pronunciation but different meaning as ambiguous reads on the one or more difficult to pronounce terms having a different meaning than the one or more other terms.).
Jitkoff is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 to incorporate the teachings of Jitkoff to identify terms that are pronounced differently but sound similar to each other and terms with the same pronunciation but different meaning as ambiguous. Doing so would allow for disambiguating ambiguous user inputs (Jitkoff; Column 1, lines 46-59).
Regarding claim 15, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 12.
Jitkoff teaches:
wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms (Column 1, lines 55-59, "Ambiguous input may be particularly common for spoken input, in part because of the presence of homophones, and in part because a speech-to-text processor may have difficulty differentiating words that are pronounced differently but sound similar to each other."; Column 6, lines 49-53, "In another example, voice input that is identified as being a homophone (multiple terms with the same pronunciation but different meanings) and/or a homonym (multiple terms with the same spelling and pronunciation but different meaning) can be identified as ambiguous."; Difficulty differentiating words that are pronounced differently but sound similar to each other reads on one or more difficult to pronounce terms having an identical or similar audible sound as one or more other terms, and identifying terms with the same pronunciation but different meaning as ambiguous reads on the one or more difficult to pronounce terms having a different meaning than the one or more other terms.).
Jitkoff is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 to incorporate the teachings of Jitkoff to identify terms that are pronounced differently but sound similar to each other and terms with the same pronunciation but different meaning as ambiguous. Doing so would allow for disambiguating ambiguous user inputs (Jitkoff; Column 1, lines 46-59).
Claims 6 and 16 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,508,354 in view of Karas.
Regarding claim 6, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 2.
Karas teaches:
further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query (Column 9, lines 18-22, "In step 27, a feature of the user's speech that requires direction to more accurately pronounce a particular phoneme is identified. The particular phoneme is a phoneme corresponding to the audio components from which a pattern of mispronunciation is identified in step 26."; Column 9, lines 28-32, "It is to be understood that each of the steps 21 to 28 may be implemented by one or more processors. The one more processors may include memory for storing and reading data. Furthermore, each of the steps 21 to 28 may be implemented in a neural network architecture."; Identifying a pattern of mispronunciation reads on determining that the one or more difficult to pronounce terms were mispronounced in the voice query, and a neural network reads on a machine learning model.).
Karas is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 to incorporate the teachings of Karas to identify a pattern of mispronunciation using a neural network. Doing so would allow for identifying mistakes in pronunciation and providing feedback to users based on the identified mistakes (Karas; Column 2, lines 11-20).
Regarding claim 16, U.S. Patent No. 11,508,354 claims all the limitations set forth in the application claim 12.
Karas teaches:
further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query (Column 9, lines 18-22, "In step 27, a feature of the user's speech that requires direction to more accurately pronounce a particular phoneme is identified. The particular phoneme is a phoneme corresponding to the audio components from which a pattern of mispronunciation is identified in step 26."; Column 9, lines 28-32, "It is to be understood that each of the steps 21 to 28 may be implemented by one or more processors. The one more processors may include memory for storing and reading data. Furthermore, each of the steps 21 to 28 may be implemented in a neural network architecture."; Identifying a pattern of mispronunciation reads on determining that the one or more difficult to pronounce terms were mispronounced in the voice query, and a neural network reads on a machine learning model.).
Karas is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,508,354 to incorporate the teachings of Karas to identify a pattern of mispronunciation using a neural network. Doing so would allow for identifying mistakes in pronunciation and providing feedback to users based on the identified mistakes (Karas; Column 2, lines 11-20).
Claims 2, 7 – 8, 12 and 17 – 18 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,948,551. Although the claims at issue are not identical, they are not patentably distinct from each other.
Regarding claim 2, claims 1 and 4 of U.S. Patent No. 11,948,551 claim all the limitations set forth in the application claim 2.
US Application No. 19/002,025 Claim 2
U.S. Patent No. 11,948,551 Claim 1
A computer-implemented method comprising:
A method of correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the method comprising:
receiving a voice query;
A method of correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system,
identifying a first transcription based at least in part on the voice query,
from incorrect transcriptions by an automated speech recognition system,
generating a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors;
generating a second transcription output based on the second subset of terms that does not match the first transcription output;
causing submission of at least a portion of the second transcription as a search query to an electronic content search system;
receiving second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
and causing output of one or more content search results of the search query.
receiving second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
US Application No. 19/002,025 Claim 2
U.S. Patent No. 11,948,551 Claim 4
wherein the first transcription includes one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors;
determining a presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 7, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 2. Claim 4 of U.S. Patent No. 11,948,551 claims all the additional limitations set forth in the application claim 7.
US Application No. 19/002,025 Claim 7
U.S. Patent No. 11,948,551 Claim 4
The method of claim 2,
The method of claim 2,
further comprising determining that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription.
wherein the determining that the occurrence of the improper search resulted from the incorrect transcription further comprises determining a presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 8, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 2. Claims 1 and 3 of U.S. Patent No. 11,948,551 claims all the additional limitations set forth in the application claim 8.
US Application No. 19/002,025 Claim 8
U.S. Patent No. 11,948,551 Claim 1
The method of claim 2,
A method of correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the method comprising:
further comprising: causing the first transcription to be submitted to the electronic content search system;
receiving second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
providing for output representations of content returned by the electronic content search system for the first transcription;
receiving second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
US Application No. 19/002,025 Claim 8
U.S. Patent No. 11,948,551 Claim 3
determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value.
wherein the determining the occurrence of the improper search further comprises determining the occurrence of the improper search according to one or more of: whether no electronic content was selected by the electronic content search system as a result of the search; whether electronic content selected by the electronic content search system as a result of the search has an associated relevance score less than a predetermined value; whether the terms of the transcription are not connected or weakly connected in a knowledge graph of terms of the electronic content; or whether the terms of the transcription have popularity scores below a predetermined value.
Regarding claim 12, claims 11 and 14 of U.S. Patent No. 11,948,551 claim all the limitations set forth in the application claim 12.
US Application No. 19/002,025 Claim 12
U.S. Patent No. 11,948,551 Claim 11
A system comprising: control circuitry configured to:
A system for correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the system comprising: memory; and control circuitry configured to:
receive a voice query;
A system for correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system,
identify a first transcription based at least in part on the voice query,
from incorrect transcriptions by an automated speech recognition system,
generate a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors;
generate a second transcription output based on the second subset of terms that does not match the first transcription output;
cause submission of at least a portion of the second transcription as a search query to an electronic content search system;
receive second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
and cause output of one or more content search results of the search query.
receive second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
US Application No. 19/002,025 Claim 2
U.S. Patent No. 11,948,551 Claim 14
wherein the first transcription includes one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors;
determining the presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 17, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 12. Claim 14 of U.S. Patent No. 11,948,551 claims all the additional limitations set forth in the application claim 17.
US Application No. 19/002,025 Claim 17
U.S. Patent No. 11,948,551 Claim 14
The system of claim 12,
The system of claim 12,
wherein the control circuitry is further configured determine that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription.
wherein the determining that the occurrence of the improper search resulted from the incorrect transcription further comprises determining a presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 18, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 12. Claims 11 and 13 of U.S. Patent No. 11,948,551 claims all the additional limitations set forth in the application claim 18.
US Application No. 19/002,025 Claim 18
U.S. Patent No. 11,948,551 Claim 11
The system of claim 12, wherein the I/O circuitry is further configured to:
A system for correcting errors in searches for electronic content resulting from incorrect transcriptions by an automated speech recognition system, the system comprising: memory; and control circuitry configured to:
cause the first transcription to be submitted to the electronic content search system;
receive second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
provide for output representations of content returned by the electronic content search system for the first transcription;
receive second results from the electronic content search system based on the second transcription output that do not match the incorrect results from the electronic content search system based on the incorrect transcription output.
US Application No. 19/002,025 Claim 8
U.S. Patent No. 11,948,551 Claim 13
determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value.
wherein the determining the occurrence of the improper search further comprises determining the occurrence of the improper search according to one or more of: whether no electronic content was selected by the electronic content search system as a result of the search; whether electronic content selected by the electronic content search system as a result of the search has an associated relevance score less than a predetermined value; whether the terms of the transcription are not connected or weakly connected in a knowledge graph of terms of the electronic content; or whether the terms of the transcription have popularity scores below a predetermined value.
Claims 3 – 4 and 13 – 14 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,948,551 in view of Tanaka.
Regarding claim 3, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 2. Claim 4 of U.S. Patent No. 11,948,551 claims the additional limitations set forth in the application claim 3, expect for the limitation “wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms”.
US Application No. 19/002,025 Claim 3
U.S. Patent No. 11,948,551 Claim 4
The method of claim 2,
The method of claim 2,
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Tanaka teaches:
wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 4, U.S. Patent No. 11,948,551 in view of Tanaka claims all the limitations set forth in the application claim 3.
Tanaka further teaches:
wherein comparing the one or more difficult to pronounce terms to the stored list of terms comprises comparing syllables or letter combinations associated with the one or more difficult to pronounce terms to syllables or letter combinations respectively associated with each term of the stored list of terms (Column 11, line 60 - Column 12, line 5, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word. Note that the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word."; Obtaining a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word, reads on comparing syllables associated with the one or more difficult to pronounce terms to syllables respectively associated with each term of the stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 in view of Tanaka to further incorporate the teachings of Tanaka to obtain a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 13, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 12. Claim 14 of U.S. Patent No. 11,948,551 claims the additional limitations set forth in the application claim 13, expect for the limitation “wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory”.
US Application No. 19/002,025 Claim 13
U.S. Patent No. 11,948,551 Claim 14
The system of claim 12,
The system of claim 12,
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms;
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
whether one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Tanaka teaches:
wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms, and a database reads on a computer memory.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 14, U.S. Patent No. 11,948,551 in view of Tanaka claims all the limitations set forth in the application claim 13.
Tanaka further teaches:
wherein comparing the one or more difficult to pronounce terms to the stored list of terms comprises comparing syllables or letter combinations associated with the one or more difficult to pronounce terms to syllables or letter combinations respectively associated with each term of the stored list of terms (Column 11, line 60 - Column 12, line 5, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word. Note that the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word."; Obtaining a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word, reads on comparing syllables associated with the one or more difficult to pronounce terms to syllables respectively associated with each term of the stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 in view of Tanaka to further incorporate the teachings of Tanaka to obtain a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Claims 5 and 15 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,948,551 in view of Jitkoff.
Regarding claim 5, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 2.
Jitkoff teaches:
wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms (Column 1, lines 55-59, "Ambiguous input may be particularly common for spoken input, in part because of the presence of homophones, and in part because a speech-to-text processor may have difficulty differentiating words that are pronounced differently but sound similar to each other."; Column 6, lines 49-53, "In another example, voice input that is identified as being a homophone (multiple terms with the same pronunciation but different meanings) and/or a homonym (multiple terms with the same spelling and pronunciation but different meaning) can be identified as ambiguous."; Difficulty differentiating words that are pronounced differently but sound similar to each other reads on one or more difficult to pronounce terms having an identical or similar audible sound as one or more other terms, and identifying terms with the same pronunciation but different meaning as ambiguous reads on the one or more difficult to pronounce terms having a different meaning than the one or more other terms.).
Jitkoff is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 to incorporate the teachings of Jitkoff to identify terms that are pronounced differently but sound similar to each other and terms with the same pronunciation but different meaning as ambiguous. Doing so would allow for disambiguating ambiguous user inputs (Jitkoff; Column 1, lines 46-59).
Regarding claim 15, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 12.
Jitkoff teaches:
wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms (Column 1, lines 55-59, "Ambiguous input may be particularly common for spoken input, in part because of the presence of homophones, and in part because a speech-to-text processor may have difficulty differentiating words that are pronounced differently but sound similar to each other."; Column 6, lines 49-53, "In another example, voice input that is identified as being a homophone (multiple terms with the same pronunciation but different meanings) and/or a homonym (multiple terms with the same spelling and pronunciation but different meaning) can be identified as ambiguous."; Difficulty differentiating words that are pronounced differently but sound similar to each other reads on one or more difficult to pronounce terms having an identical or similar audible sound as one or more other terms, and identifying terms with the same pronunciation but different meaning as ambiguous reads on the one or more difficult to pronounce terms having a different meaning than the one or more other terms.).
Jitkoff is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 to incorporate the teachings of Jitkoff to identify terms that are pronounced differently but sound similar to each other and terms with the same pronunciation but different meaning as ambiguous. Doing so would allow for disambiguating ambiguous user inputs (Jitkoff; Column 1, lines 46-59).
Claims 6 and 16 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 11,948,551 in view of Karas.
Regarding claim 6, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 2.
Karas teaches:
further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query (Column 9, lines 18-22, "In step 27, a feature of the user's speech that requires direction to more accurately pronounce a particular phoneme is identified. The particular phoneme is a phoneme corresponding to the audio components from which a pattern of mispronunciation is identified in step 26."; Column 9, lines 28-32, "It is to be understood that each of the steps 21 to 28 may be implemented by one or more processors. The one more processors may include memory for storing and reading data. Furthermore, each of the steps 21 to 28 may be implemented in a neural network architecture."; Identifying a pattern of mispronunciation reads on determining that the one or more difficult to pronounce terms were mispronounced in the voice query, and a neural network reads on a machine learning model.).
Karas is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 to incorporate the teachings of Karas to identify a pattern of mispronunciation using a neural network. Doing so would allow for identifying mistakes in pronunciation and providing feedback to users based on the identified mistakes (Karas; Column 2, lines 11-20).
Regarding claim 16, U.S. Patent No. 11,948,551 claims all the limitations set forth in the application claim 12.
Karas teaches:
further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query (Column 9, lines 18-22, "In step 27, a feature of the user's speech that requires direction to more accurately pronounce a particular phoneme is identified. The particular phoneme is a phoneme corresponding to the audio components from which a pattern of mispronunciation is identified in step 26."; Column 9, lines 28-32, "It is to be understood that each of the steps 21 to 28 may be implemented by one or more processors. The one more processors may include memory for storing and reading data. Furthermore, each of the steps 21 to 28 may be implemented in a neural network architecture."; Identifying a pattern of mispronunciation reads on determining that the one or more difficult to pronounce terms were mispronounced in the voice query, and a neural network reads on a machine learning model.).
Karas is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 11,948,551 to incorporate the teachings of Karas to identify a pattern of mispronunciation using a neural network. Doing so would allow for identifying mistakes in pronunciation and providing feedback to users based on the identified mistakes (Karas; Column 2, lines 11-20).
Claims 2, 7 – 8, 12 and 17 – 18 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 12,217,737. Although the claims at issue are not identical, they are not patentably distinct from each other.
Regarding claim 2, claims 1 and 6 of U.S. Patent No. 12,217,737 claim all the limitations set forth in the application claim 2.
US Application No. 19/002,025 Claim 2
U.S. Patent No. 12,217,737 Claim 1
A computer-implemented method comprising:
A method of correcting errors in searches for content resulting from incorrect transcriptions by an automated speech recognition (ASR) system, the method comprising:
receiving a voice query;
A method of correcting errors in searches for content resulting from incorrect transcriptions by an automated speech recognition (ASR) system,
identifying a first transcription based at least in part on the voice query,
receiving an incorrect transcription output by an ASR module,
generating a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors;
receiving from the ASR module a plurality of resulting transcriptions, based on the successively submitted different subsets of the plurality of transcription terms with the one or more hint words, that are each different from the incorrect transcription;
causing submission of at least a portion of the second transcription as a search query to an electronic content search system;
submitting the most common transcription as a search query of a database;
and causing output of one or more content search results of the search query.
and providing for display one or more representations of content returned by the search query of the database using the most common transcription.
US Application No. 19/002,025 Claim 2
U.S. Patent No. 12,217,737 Claim 6
wherein the first transcription includes one or more difficult to pronounce terms;
whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors;
determining a presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 7, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 2. Claim 6 of U.S. Patent No. 12,217,737 claims all the additional limitations set forth in the application claim 7.
US Application No. 19/002,025 Claim 7
U.S. Patent No. 12,217,737 Claim 6
The method of claim 2,
The method of claim 5,
further comprising determining that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription.
wherein the determining that the occurrence of the improper search resulted from the incorrect transcription further comprises determining a presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 8, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 2. Claims 1 and 5 of U.S. Patent No. 12,217,737 claims all the additional limitations set forth in the application claim 8.
US Application No. 19/002,025 Claim 8
U.S. Patent No. 12,217,737 Claim 1
The method of claim 2,
A method of correcting errors in searches for content resulting from incorrect transcriptions by an automated speech recognition (ASR) system, the method comprising:
further comprising: causing the first transcription to be submitted to the electronic content search system;
submitting the most common transcription as a search query of a database;
providing for output representations of content returned by the electronic content search system for the first transcription;
and providing for display one or more representations of content returned by the search query of the database using the most common transcription.
US Application No. 19/002,025 Claim 8
U.S. Patent No. 12,217,737 Claim 5
determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value.
further comprising determining an occurrence of an improper search by determining one or more of: whether no content was selected by the search query of the database as a result of a search; whether the content returned by the search query of the database as the result of the search has an associated relevance score less than a predetermined value; whether terms of a transcription are not connected or weakly connected in a knowledge graph of terms of the content; or whether the terms of the transcription have popularity scores below a predetermined value.
Regarding claim 12, claims 11 and 16 of U.S. Patent No. 12,217,737 claim all the limitations set forth in the application claim 12.
US Application No. 19/002,025 Claim 12
U.S. Patent No. 12,217,737 Claim 11
A system comprising: control circuitry configured to:
A system of correcting errors in searches for content resulting from incorrect transcriptions by an automated speech recognition (ASR) system, the system comprising: a control circuitry; and an input/output (I/O) circuitry configured to:
receive a voice query;
A system of correcting errors in searches for content resulting from incorrect transcriptions by an automated speech recognition (ASR) system,
identify a first transcription based at least in part on the voice query,
receive an incorrect transcription output by an ASR module,
generate a second transcription based at least in part on the voice query and determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors;
receive from the ASR module a plurality of resulting transcriptions, based on the successively submitted different subsets of the plurality of transcription terms with the one or more hint words, that are each different from the incorrect transcription;
cause submission of at least a portion of the second transcription as a search query to an electronic content search system;
submit the most common transcription as a search query of a database;
and cause output of one or more content search results of the search query.
and provide for display one or more representations of content returned by the search query of the database using the most common transcription.
US Application No. 19/002,025 Claim 2
U.S. Patent No. 12,217,737 Claim 16
wherein the first transcription includes one or more difficult to pronounce terms;
whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
determining that the one or more difficult to pronounce terms of the first transcription is associated with one or more transcription errors;
determining a presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 17, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 12. Claim 16 of U.S. Patent No. 12,217,737 claims all the additional limitations set forth in the application claim 17.
US Application No. 19/002,025 Claim 17
U.S. Patent No. 12,217,737 Claim 16
The system of claim 12,
The system of claim 15,
wherein the control circuitry is further configured determine that the first transcription is associated with the one or more transcription errors based at least in part on a number of the one or more difficult to pronounce terms included in the first transcription.
wherein the control circuitry is further configured to determine that the occurrence of the improper search resulted from the incorrect transcription by determining a presence of the incorrect transcription according to one or more of: whether a transcription score of the transcription corresponding to the improper search is less than a predetermined value; whether one or more terms of the transcription corresponding to the improper search have a homonym; whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Regarding claim 18, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 12. Claims 11 and 15 of U.S. Patent No. 12,217,737 claims all the additional limitations set forth in the application claim 18.
US Application No. 19/002,025 Claim 18
U.S. Patent No. 12,217,737 Claim 11
The system of claim 12, wherein the I/O circuitry is further configured to:
A system of correcting errors in searches for content resulting from incorrect transcriptions by an automated speech recognition (ASR) system, the system comprising: a control circuitry; and an input/output (I/O) circuitry configured to:
cause the first transcription to be submitted to the electronic content search system;
submit the most common transcription as a search query of a database;
provide for output representations of content returned by the electronic content search system for the first transcription;
and provide for display one or more representations of content returned by the search query of the database using the most common transcription.
US Application No. 19/002,025 Claim 8
U.S. Patent No. 12,217,737 Claim 15
determining that the first transcription is associated with the one or more transcription errors further based at least in part on determining: whether none of the representations of content were selected by a user; whether the representations of content returned by the electronic content search system is associated with a relevance score that is less than a predetermined value; whether the one or more difficult to pronounce terms of the first transcription are not connected or weakly connected in a knowledge graph of terms associated with the representations of content; or whether the one or more difficult to pronounce terms are associated with popularity scores below a predetermined value.
wherein the control circuitry is further configured to determine an occurrence of an improper search by determining one or more of: whether no content was selected by the search query of the database as a result of a search; whether the content returned by the search query of the database as the result of the search has an associated relevance score less than a predetermined value; whether terms of a transcription are not connected or weakly connected in a knowledge graph of terms of the content; or whether the terms of the transcription have popularity scores below a predetermined value.
Claims 3 – 4 and 13 – 14 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 12,217,737 in view of Tanaka.
Regarding claim 3, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 2. Claim 6 of U.S. Patent No. 12,217,737 claims the additional limitations set forth in the application claim 3, expect for the limitation “wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms”.
US Application No. 19/002,025 Claim 3
U.S. Patent No. 12,217,737 Claim 6
The method of claim 2,
The method of claim 5,
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms;
whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Tanaka teaches:
wherein determining that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors comprises: comparing the one or more difficult to pronounce terms to a stored list of terms (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 4, U.S. Patent No. 12,217,737 in view of Tanaka claims all the limitations set forth in the application claim 3.
Tanaka further teaches:
wherein comparing the one or more difficult to pronounce terms to the stored list of terms comprises comparing syllables or letter combinations associated with the one or more difficult to pronounce terms to syllables or letter combinations respectively associated with each term of the stored list of terms (Column 11, line 60 - Column 12, line 5, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word. Note that the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word."; Obtaining a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word, reads on comparing syllables associated with the one or more difficult to pronounce terms to syllables respectively associated with each term of the stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 in view of Tanaka to further incorporate the teachings of Tanaka to obtain a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 13, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 12. Claim 16 of U.S. Patent No. 12,217,737 claims the additional limitations set forth in the application claim 13, expect for the limitation “wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory”.
US Application No. 19/002,025 Claim 13
U.S. Patent No. 12,217,737 Claim 16
The system of claim 12,
The system of claim 15,
based at least in part on the comparing, determining a pronunciation difficulty metric corresponding to the one or more difficult to pronounce terms;
whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
and determining that the pronunciation difficulty metric exceeds a predetermined difficulty threshold.
whether the one or more terms of the transcription corresponding to the improper search have a corresponding pronunciation difficulty metric greater than a predetermined value;
Tanaka teaches:
wherein the control circuitry is configured to determine that the one or more difficult to pronounce terms of the first transcription is associated with the one or more transcription errors by: comparing the one or more difficult to pronounce terms to a stored list of terms, wherein the stored list of terms is stored in a computer memory (Column 11, line 60 - Column 12, line 3, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word."; Determining if a word matches a word registered in a pronunciation difficulty database reads on comparing the one or more difficult to pronounce terms to a stored list of terms, and a database reads on a computer memory.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 to incorporate the teachings of Tanaka to determine if a word matches a word registered in a pronunciation difficulty database. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Regarding claim 14, U.S. Patent No. 12,217,737 in view of Tanaka claims all the limitations set forth in the application claim 13.
Tanaka further teaches:
wherein comparing the one or more difficult to pronounce terms to the stored list of terms comprises comparing syllables or letter combinations associated with the one or more difficult to pronounce terms to syllables or letter combinations respectively associated with each term of the stored list of terms (Column 11, line 60 - Column 12, line 5, "In this embodiment, the threshold setting section 12 obtains a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database recording a pronunciation difficulty of each word prestored in the storage unit 5. For example, the threshold setting section 12 detects a word that matches text data of a retrieval word specified through the user interface unit 6 from among the words registered in the pronunciation difficulty database, and sets a pronunciation difficulty corresponding to the detected word as the pronunciation difficulty of the retrieval word. Note that the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word."; Obtaining a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word, reads on comparing syllables associated with the one or more difficult to pronounce terms to syllables respectively associated with each term of the stored list of terms.).
Tanaka is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 in view of Tanaka to further incorporate the teachings of Tanaka to obtain a pronunciation difficulty of a retrieval word by referring to a pronunciation difficulty database, where the pronunciation difficulty is expressed as the ratio of the number of difficult pronunciation points to the number of syllables of the word. Doing so would allow for precise voice retrieval processing having relatively high detection accuracy despite relatively high throughput (Tanaka; Column 2, lines 24-39).
Claims 5 and 15 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 12,217,737 in view of Jitkoff.
Regarding claim 5, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 2.
Jitkoff teaches:
wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms (Column 1, lines 55-59, "Ambiguous input may be particularly common for spoken input, in part because of the presence of homophones, and in part because a speech-to-text processor may have difficulty differentiating words that are pronounced differently but sound similar to each other."; Column 6, lines 49-53, "In another example, voice input that is identified as being a homophone (multiple terms with the same pronunciation but different meanings) and/or a homonym (multiple terms with the same spelling and pronunciation but different meaning) can be identified as ambiguous."; Difficulty differentiating words that are pronounced differently but sound similar to each other reads on one or more difficult to pronounce terms having an identical or similar audible sound as one or more other terms, and identifying terms with the same pronunciation but different meaning as ambiguous reads on the one or more difficult to pronounce terms having a different meaning than the one or more other terms.).
Jitkoff is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 to incorporate the teachings of Jitkoff to identify terms that are pronounced differently but sound similar to each other and terms with the same pronunciation but different meaning as ambiguous. Doing so would allow for disambiguating ambiguous user inputs (Jitkoff; Column 1, lines 46-59).
Regarding claim 15, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 12.
Jitkoff teaches:
wherein the one or more difficult to pronounce terms has an identical or similar audible sound as one or more other terms, and wherein the one or more difficult to pronounce terms has a different meaning than the one or more other terms (Column 1, lines 55-59, "Ambiguous input may be particularly common for spoken input, in part because of the presence of homophones, and in part because a speech-to-text processor may have difficulty differentiating words that are pronounced differently but sound similar to each other."; Column 6, lines 49-53, "In another example, voice input that is identified as being a homophone (multiple terms with the same pronunciation but different meanings) and/or a homonym (multiple terms with the same spelling and pronunciation but different meaning) can be identified as ambiguous."; Difficulty differentiating words that are pronounced differently but sound similar to each other reads on one or more difficult to pronounce terms having an identical or similar audible sound as one or more other terms, and identifying terms with the same pronunciation but different meaning as ambiguous reads on the one or more difficult to pronounce terms having a different meaning than the one or more other terms.).
Jitkoff is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 to incorporate the teachings of Jitkoff to identify terms that are pronounced differently but sound similar to each other and terms with the same pronunciation but different meaning as ambiguous. Doing so would allow for disambiguating ambiguous user inputs (Jitkoff; Column 1, lines 46-59).
Claims 6 and 16 are rejected on the ground of nonstatutory double patenting as being unpatentable over U.S. Patent No. 12,217,737 in view of Karas.
Regarding claim 6, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 2.
Karas teaches:
further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query (Column 9, lines 18-22, "In step 27, a feature of the user's speech that requires direction to more accurately pronounce a particular phoneme is identified. The particular phoneme is a phoneme corresponding to the audio components from which a pattern of mispronunciation is identified in step 26."; Column 9, lines 28-32, "It is to be understood that each of the steps 21 to 28 may be implemented by one or more processors. The one more processors may include memory for storing and reading data. Furthermore, each of the steps 21 to 28 may be implemented in a neural network architecture."; Identifying a pattern of mispronunciation reads on determining that the one or more difficult to pronounce terms were mispronounced in the voice query, and a neural network reads on a machine learning model.).
Karas is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 to incorporate the teachings of Karas to identify a pattern of mispronunciation using a neural network. Doing so would allow for identifying mistakes in pronunciation and providing feedback to users based on the identified mistakes (Karas; Column 2, lines 11-20).
Regarding claim 16, U.S. Patent No. 12,217,737 claims all the limitations set forth in the application claim 12.
Karas teaches:
further comprising using one or more machine learning models to determine that the one or more difficult to pronounce terms were mispronounced in the voice query (Column 9, lines 18-22, "In step 27, a feature of the user's speech that requires direction to more accurately pronounce a particular phoneme is identified. The particular phoneme is a phoneme corresponding to the audio components from which a pattern of mispronunciation is identified in step 26."; Column 9, lines 28-32, "It is to be understood that each of the steps 21 to 28 may be implemented by one or more processors. The one more processors may include memory for storing and reading data. Furthermore, each of the steps 21 to 28 may be implemented in a neural network architecture."; Identifying a pattern of mispronunciation reads on determining that the one or more difficult to pronounce terms were mispronounced in the voice query, and a neural network reads on a machine learning model.).
Karas is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified U.S. Patent No. 12,217,737 to incorporate the teachings of Karas to identify a pattern of mispronunciation using a neural network. Doing so would allow for identifying mistakes in pronunciation and providing feedback to users based on the identified mistakes (Karas; Column 2, lines 11-20).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Lewis (US Patent No. 11,848,000) teaches a method for efficient correction of a transcription output of an automatic speech recognition system.
Raghunathan et al. (US Patent No. 11,482,213) teaches a method for correcting transcriptions created through automatic speech recognition.
Prabhavalkar et al. (US Patent No. 11,217,231) teaches a method of biasing speech recognition including receiving audio data encoding an utterance and obtaining a set of one or more biasing phrases corresponding to a context of the utterance.
Norouzi et al. (US Patent No. 11,043,213) teaches a method for identifying an incorrect pronunciation for a word in a segment of speech audio.
Kapralova et al. (US Patent No. 10,019,986) teaches a speech recognition system that receives an utterance of one or more terms from a user, provides a transcription of the utterance to a user device, receives user input to correct a particular term or terms of the transcription when the provided transcription is not correct, and uses the user input to correct the particular term or terms and audio data corresponding to the particular term or terms.
Chen (US Patent No. 8,762,156) teaches a speech control system that can recognize a spoken command and associated words and can cause a selected application to execute the command to cause a data processing system to perform an operation based on the command, where the speech control system can use a set of interpreters to repair recognized text from a speech recognition system.
Rao et al. ("Talking to Your TV: Context-Aware Voice Search with Hierarchical Recurrent Neural Networks") teaches a method for navigational voice queries for an entertainment system.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Boggs whose telephone number is (571)272-2968. The examiner can normally be reached M-F 8:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAMES BOGGS/Examiner, Art Unit 2657