Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This office action is in response to application 19/007,920, which was filed 01/02/25 and is a continuation of application 18/309,754, now U.S. Patent 12,198,671, which is a continuation of application 17/153,463, now U.S. Patent 11,670,281, which is a continuation of application 16/573,492, now U.S. Patent 10,923,100, which is a continuation of application 16/135,885, now U.S. Patent 10,453,441, which is a continuation of application 15/653,872, now U.S. Patent 10,109,270, which is a continuation of application 15/477,360, now U.S. Patent 9,886,942, which is a continuation of application 15/009,432, now U.S. Patent 9,799,324. Claims 1-20 are pending in the application and have been considered.
Double Patenting Considerations
Claims 1-20 of the present application have been compared to the claims of the issued U.S. Patent 12,198,671, U.S. Patent 11,670,281, U.S. Patent 10,923,100, U.S. Patent 10,453,441, U.S. Patent 10,109,270, U.S. Patent 9,886,942 and U.S. Patent 9,799,324, and are considered to be dissimilar to the claims of each of the above-mentioned patents to the extent that no double patenting rejections would be proper.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 6, 10-12, 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Talwar et al. (US 20130080173) in view of Buryak et al. (US 20140335483).
Consider claim 1, Talwar discloses a computer-implemented method when executed on data processing hardware (processor executing computer code, [0089]) causes the data processing hardware to perform operations comprising:
receiving a user input (user utters speech which is received by microphone 32, [0083]);
a level of language proficiency for the user (communication ability of the user as novice, expert, native speaker, or non-native speaker is identified, [0084]);
receiving a query input by the user to a client device (“what?”, [0083]); and
generating a text-to-speech response to the query and based on the level of language proficiency (generating text for subsequent synthesized speech processing, [0085], by producing different verbiage for the non-native and native speakers, [0086]), the text-to-speech response comprising one of:
a first text segment when the language proficiency determined for the user comprises a first level of language proficiency, the first text segment comprising primary information responsive to the query (the system produces a less verbose response for an expert or native speaker, e.g. “Number Please”, [0084-0086]); or
a second text segment when the language proficiency determined for the user comprises a second level of language proficiency, the second text segment comprising additional information responsive to the query that is not included in the first text segment (the system produces a more verbose response for a notice or non-native speaker, e.g. “Please Say a Contact Name For The Person You Are Trying To Call”, [0084-0086]); and
generating audio data comprising a synthesized utterance of the text-to-speech response to the query (the text generated in response to the query is processed into synthesized speech, [0085-0086]).
Talwar does not specifically mention during a registration process, receiving a user input that specifies a level of language proficiency for the user.
Buryak discloses during a registration process, receiving a user input that specifies a level of language proficiency for the user (user specifies level of language proficiency, e.g. 50% for Hindi and 75% for Chinese using a sliding scale interface, [0019]; this is considered a registration process since the user explicitly provides the level, which is stored in a user profile, [0033]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Talwar by receiving a user input that specifies a level of language proficiency for the user during a registration process as in Buryak, and by generating a text-to-speech response to the query as in Talwar based on the level of language proficiency specified by the user input during the registration process of Buryak in order to better accommodate multilingual users, as suggested by Buryak ([0012]), predictably enhancing user experience, as suggested by Buryak ([0012]). The cited references are analogous art in the same field of providing user assistance.
Consider claim 11, Talwar discloses a system comprising: data processing hardware (one or more processors, [0089]); and memory hardware (memory, [0090]) in communication with the data processing hardware and storing instructions, that when executed by the data processing hardware (processor executing computer code, [0089]), cause the data processing hardware to perform operations comprising:
receiving a user input (user utters speech which is received by microphone 32, [0083]);
a level of language proficiency for the user (communication ability of the user as novice, expert, native speaker, or non-native speaker is identified, [0084]);
receiving a query input by the user to a client device (“what?”, [0083]); and
generating a text-to-speech response to the query and based on the level of language proficiency (generating text for subsequent synthesized speech processing, [0085], by producing different verbiage for the non-native and native speakers, [0086]), the text-to-speech response comprising one of:
a first text segment when the language proficiency determined for the user comprises a first level of language proficiency, the first text segment comprising primary information responsive to the query (the system produces a less verbose response for an expert or native speaker, e.g. “Number Please”, [0084-0086]); or
a second text segment when the language proficiency determined for the user comprises a second level of language proficiency, the second text segment comprising additional information responsive to the query that is not included in the first text segment (the system produces a more verbose response for a notice or non-native speaker, e.g. “Please Say a Contact Name For The Person You Are Trying To Call”, [0084-0086]); and
generating audio data comprising a synthesized utterance of the text-to-speech response to the query (the text generated in response to the query is processed into synthesized speech, [0085-0086]).
Talwar does not specifically mention during a registration process, receiving a user input that specifies a level of language proficiency for the user.
Buryak discloses during a registration process, receiving a user input that specifies a level of language proficiency for the user (user specifies level of language proficiency, e.g. 50% for Hindi and 75% for Chinese using a sliding scale interface, [0019]; this is considered a registration process since the user explicitly provides the level, which is stored in a user profile, [0033]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Talwar by receiving a user input that specifies a level of language proficiency for the user during a registration process as in Buryak, and by generating a text-to-speech response to the query as in Talwar based on the level of language proficiency specified by the user input during the registration process of Buryak for reasons similar to those for claim 1.
Consider claim 2, Talwar discloses providing the audio data for audible output from the client device (the subsequent synthesized speech is output via a loudspeaker, [0087]).
Consider claim 6, Talwar discloses, prior to generating the text-to-speech response: identifying multiple candidate text segments that are responsive to the query, each candidate text segment associated with a different level of language complexity (e.g. “number please” and “Please Say a Contact Name For The Person You Are Trying To Call”, [0084-0086]); and selecting, from among the multiple candidate text segments, the text-to-speech response to the query based on the language proficiency designated to the user (producing different verbiage for the non-native and native speakers, [0086]).
Consider claim 10, Talwar discloses: the second level of language proficiency comprises a higher level of language proficiency than the first level of language proficiency (native speaker versus non-native speaker, [0084]); and the second text segment is associated with a grammatical structure that is more complex than a grammatical structure associated with the first text segment (the prompt “Number Please” is fairly considered more complex than the sentence “Please Say A Contact Name For The Person You Are Trying To Call” since Talwar considers the latter “simpler to understand” and having “greater context and understanding”, [0085]).
Consider claim 12, Talwar discloses: providing the audio data for audible output from the client device (the subsequent synthesized speech is output via a loudspeaker, [0087]).
Consider claim 16, Talwar discloses, prior to generating the text-to-speech response: identifying multiple candidate text segments that are responsive to the query, each candidate text segment associated with a different level of language complexity (e.g. “number please” and “Please Say a Contact Name For The Person You Are Trying To Call”, [0084-0086]); and selecting, from among the multiple candidate text segments, the text-to-speech response to the query based on the language proficiency designated to the user (producing different verbiage for the non-native and native speakers, [0086]).
Consider claim 20, Talwar discloses: the second level of language proficiency comprises a higher level of language proficiency than the first level of language proficiency (native speaker versus non-native speaker, [0084]); and the second text segment is associated with a grammatical structure that is more complex than a grammatical structure associated with the first text segment (the prompt “Number Please” is fairly considered more complex than the sentence “Please Say A Contact Name For The Person You Are Trying To Call” since Talwar considers the latter “simpler to understand” and having “greater context and understanding”, [0085]).
Allowable Subject Matter
Claims 3-5, 7-9, 13-15, and 17-19 are objected to as being dependent on a rejected base claim, but would be allowable if rewritten in independent form including all limitations of the base and any intervening claims.
With regard to claim 3, the prior art does not fairly teach or suggest: “…the first text segment comprises a respective independent clause conveying the primary information responsive to the query; and the second text segment comprises a respective independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text segment conveying the additional information responsive to the query that is not included in the first text segment.“ Claim 13 recites similar limitations. Claims 4, 5, 14, and 15 further limit the allowable subject matter of respective parent claims 3 and 13.
With regard to claims 7 and 17, the prior art does not fairly teach or suggest: “…wherein selecting from among the multiple candidate text segments comprises: determining a language complexity score for each of the multiple candidate text segments; and selecting the text segment associated with the language complexity score that best matches a reference score that describes the language proficiency designated to the user as the particular text segment.”
With regard to claims 8 and 18, the prior art does not fairly teach or suggest: “…wherein the operations further comprise, prior to generating text-to-speech response: obtaining a baseline text segment responsive to the voice query; and generating the particular text segment by increasing a complexity level of the baseline text segment based on the language proficiency designated to the user.”
With regard to claims 9 and 19, the prior art does not fairly teach or suggest: “…wherein the operations further comprise, prior to generating the text-to-speech response: obtaining a baseline text segment responsive to the voice query; and generating the particular text segment by decreasing a complexity level of the baseline text segment based on the language proficiency designated to the user.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20140282098 McConnell discloses generating a user profile based on skill information, such as evaluating programming language proficiency based on queries
US 20130238580 D’Orazio Pedo De Matos discloses automatic adaptive content delivery based on language, skill level, etc. of the user, see [0029]
US 20050175970 Dunlap discloses interactive teaching and practicing of language listening and speaking skills
Janarthanam et al. (“Adaptive Generation in Dialogue Systems using Dynamic User Modeling”. Computational Linguistics, MIT Press, Volume 40, Number 4, 2014) discloses adapting text generation in a TTS dialogue system based on user domain expertise
US 6029156 Lannert discloses a goal based tutoring system with behavior to tailor to characteristics of a particular user
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Jesse S Pullias/
Primary Examiner, Art Unit 2655 09/09/24