DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 13-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent claims 13, 19, and 23-24 recite “select a first speech recognition dictionary …”, “generate a first speech recognition text …”, “select … a second speech recognition dictionary …”, “generate a second speech recognition text …”and “providing …”. These limitations, under its broadest reasonable interpretation, cover performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “processor”. For example, but for the “processor” language, these steps in the context of this claim encompasses the user manually selecting a first person to transcribe speech into a first text and then selecting a second person to transcribe a second speech to a second text. All these steps can be performed in the mind and/or using a pen and paper. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements - using a processor to perform these steps. The use of a processor is recited at a high-level of generality (i.e., as a generic computer device performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element/step of “display” is merely for the purpose of data gathering and/or insignificant extra-solution activity that amount to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible.
Regarding claims 14-18 and 20-22, the steps in these claims, under broadest reasonable interpretation, cover performance of the limitation in the mind but for the recitation of generic computer components in the context of this claim encompasses the user manually performing these steps. All these steps can be performed in the mind and/or using a pen and paper. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 13-16, 19, and 23-24 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Basson et al. (USPG 2003/0125940, hereinafter Basson).
Regarding claim 13, Basson discloses an information processing system comprising:
a processor (paragraph 19, “processor”); and a memory storing computer-executable instructions that, when executed by the processor (paragraph 38, memory), cause the information processing system to at least:
select a first speech recognition dictionary for use in speech recognition from among a plurality of speech recognition dictionaries (paragraph 21, “the multicaster 204 initially transmits the captured audio signals (voice data) to a speaker independent automatic speech recognition system 205 that decodes the audio data and returns the decoded data back to the controller 220”, speaker independent ASR utilizes general recognition dictionaries; also see process in figure 7);
generate first speech recognition text by converting voices uttered during a voice call with a customer, into the first speech recognition text, by speech recognition using the first speech recognition dictionary (paragraph 21, “the multicaster 204 initially transmits the captured audio signals (voice data) to a speaker independent automatic speech recognition system 205 that decodes the audio data and returns the decoded data back to the controller 220”; also see process in figure 7);
select, subsequent to the selection of the first speech recognition dictionary, a second speech recognition dictionary from among the plurality of speech recognition dictionaries (paragraph 22, “the multicaster 204 also transmits the captured audio signals (voice data) to a speaker identification system 207 if the speaker has not yet been identified. The speaker identification system 207 identifies the speaker and provides a speaker identifier to a speaker specific ASR 206, which is used by the speaker specific ASR 206 to access the appropriate speaker model from a database 209; also see process in figure 7); and
generate a second speech recognition text using the second speech recognition dictionary by converting at least a part of the voices having been converted into the first speech recognition text using the first speech recognition dictionary (paragraphs 22 and 25, “Once the identity of the speaker is known, the speaker specific ASR 206 decodes the speech using the proper model 209 and provides the decoded speech output back to the controller 220”; also see process in figure 7).
Regarding claim 19, Basson discloses an information processing system comprising: a processor (see claim 13 above); and a memory storing computer-executable instructions that, when executed by the processor (see claim 13 above), cause the information processing system to at least:
select a first speech recognition dictionary for use in speech recognition from among a plurality of speech recognition dictionaries (see claim 13 above);
generate first speech recognition text by converting voices uttered during a voice call with a customer, into the first speech recognition text, by speech recognition using the first speech recognition dictionary (see claim 13 above);
display the first speech recognition text on a screen (paragraphs 20, 22, and 25, displaying both speaker-independent and speaker-dependent recognition results of a display);
select, subsequent to the selection of the first speech recognition dictionary, a second speech recognition dictionary, from among the plurality of speech recognition dictionaries (see claim 13 above);
generate second speech recognition text by converting voices uttered before the second speech recognition dictionary is selected and voices uttered after the second speech recognition dictionary is selected, among the voices uttered during the voice call with the customer, into the second speech recognition text, by speech recognition using the second speech recognition dictionary (see claim 13 above; obvious that speaker-specific model must be selected before the speech recognizer can perform speech recognition on speech utters by the user); and
display the second speech recognition text on the screen (paragraphs 20, 22, and 25, displaying both speaker-independent and speaker-dependent recognition results of a display).
Regarding claims 23-24, Basson discloses an information processing method and non-transitionary CRM for causing a computer to perform steps including:
selecting a first speech recognition dictionary for use in speech recognition from among a plurality of speech recognition dictionaries (see claim 13 above);
generating first speech recognition text by converting voices uttered during a voice call with a customer, into the first speech recognition text, by speech recognition using the first speech recognition dictionary (see claim 13 above); and
generating, upon a switchover from the first speech recognition dictionary to a second speech recognition dictionary, second speech recognition text by converting voices uttered before the switchover is made, among the voices uttered during the voice call with the customer, into the second speech recognition text, by speech recognition using the second speech recognition dictionary (see claim 13 above, switching over to second dictionary is merely a process of using second model to recognize speech of second user).
Regarding claim 14, Basson further discloses the information processing system according to claim 13, wherein, when the second speech recognition dictionary is selected, the computer-executable program instructions further cause the information processing system to convert a plurality of voices, uttered before the second speech recognition dictionary is selected, into the second speech recognition text, by speech recognition using the second speech recognition dictionary, to generate the second speech recognition text (process in figure 7, steps 704-706, identify speaker as speech is coming in and use speaker dependent models for each identified speaker).
Regarding claim 15, Basson further discloses the information processing system according to claim 13, wherein, when the second speech recognition dictionary is selected, the computer-executable program instructions further cause the information processing system to convert only voices uttered by the customer, into the second speech recognition text, among a plurality of voices uttered before the second speech recognition dictionary is selected, by speech recognition using the second speech recognition dictionary, to generate the second speech recognition text (process in figure 7, steps 704-706, identify speaker as speech is coming in and use speaker dependent models for each identified speaker; “customer” can be one of the identified speaker).
Regarding claim 16, Basson further discloses the information processing system according to claim 13, wherein, when the second speech recognition dictionary is selected, the computer-executable program instructions further cause the information processing system to convert voices uttered after the second speech recognition dictionary is selected, into the second speech recognition text, by speech recognition using the second speech recognition dictionary, to generate the second speech recognition text (process in figure 7, steps 704-706, identify speaker as speech is coming in and use speaker dependent models for each identified speaker; speech uttered after the second dictionary is selected can still be from the same person).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 17 is rejected under 35 U.S.C. 103 as being unpatentable over Basson in view of Khalil et al. (USPG 2021/0217410, hereinafter Khalil).
Regarding claim 17, Basson fails to explicitly disclose, however, Khalil teaches the information processing system according to claim 13, wherein the computer-executable program instructions further cause the information processing system to carry out the speech recognition only after a predetermined period of time elapses from a beginning of the voice call, and generate the speech recognition text using a predetermined speech recognition dictionary when no speech recognition dictionary is selected before the predetermined period of time elapses (paragraphs 45 and/or 54, after a “batch size” has been established, it can remain fixed length or change depending on various conditions; “batch size maybe predetermined”).
Since Basson and Khalil are analogous in the art because they are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to use the known technique of processing speech at a predetermined time duration or patch size. One of ordinary skill in the art would have recognized that the results of the combination were predictable since the use of that known technique provides the rationale to arrive at a conclusion of obviousness. See KSR International Co. v. Teleflex Inc., 82 USPQ2d 1385 (U.S. 2007).
Claims 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Basson in view of Furukawa et al. (USPG 2020/0311354, hereinafter Furukawa).
Regarding claims 20-21, Basson fails to explicitly disclose, however, Furukawa teaches the information processing system according to claim 19, wherein, when the second speech recognition dictionary is selected, the computer-executable program instructions further cause the information processing system to at least: divide the screen into a first screen and a second screen (see figures 1A-C, display both texts produced by both users in two sections of a display); display, on the first screen, the second speech recognition text obtained by converting the voices uttered after the second speech recognition dictionary is selected, into the second speech recognition text, by the speech recognition using the second speech recognition dictionary; and display, on the second screen, the first speech recognition text or the second speech recognition text obtained by converting the voices uttered before the second speech recognition dictionary is selected, into the second speech recognition text, by the speech recognition using the second speech recognition dictionary (see figures 1A-C, display both texts produced by both users in two sections of a display); and further cause the information processing system to display, on the first screen, the second speech recognition text, obtained by converting a latest voice uttered after the second speech recognition dictionary is selected, by the speech recognition using the second speech recognition dictionary (see figures 1A-C, display both texts produced by both users in two sections of a display).
Since Basson and Furukawa are analogous in the art because they are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to use the known technique of displaying texts produced by respective speakers in separate sections of a display. One of ordinary skill in the art would have recognized that the results of the combination were predictable since the use of that known technique provides the rationale to arrive at a conclusion of obviousness. See KSR International Co. v. Teleflex Inc., 82 USPQ2d 1385 (U.S. 2007).
Allowable Subject Matter
Claims 18 and 22 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and resolving the 101 issue.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Thomas et al. (USPG 2020/0243094) teach a method of switching between speech recognition system that is considered pertinent to the claimed invention.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUYEN X VO whose telephone number is (571)272-7631. The examiner can normally be reached M-F, 8-4.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HUYEN X VO/Primary Examiner, Art Unit 2656