Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This office action is in response to application 18/879,694, which was filed 12/27/24. Claims 1-15 are pending in the application and have been considered.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
The following title is suggested: Determining whether a User Utterance has Ended.
Claim Objections
Claim 9 is objected to because of the following informalities: in line 3, the examiner assumes “a speech” should be “a speech input”, since line 6 refers to “the speech input”. Also, in line 7, “a utterance” should be “an utterance” (similarly to line 6 of claim 1). Appropriate correction is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 5, 6, 8, 9, and 13-15 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Liu et al. (US 9437186).
Consider claim 1, Liu discloses a display device (ASR device having a display, Col 7 lines 1-6), comprising:
a network interface communicating with a first server and a second server (portions of the speech recognition process implemented on remote servers as part of a distributed computing system, Col 2 lines 54-60, ASR module, considered a “first server” connected via network interface to ASR device 302, Col 7 lines 63-66, as well as NLU unit, a “second server”, Col 11 lines 13-19); and
a controller configured to obtain a speech input of a user (controller receives speech from user from local device over network, Col 6 lines 20-48),
transmit a speech signal corresponding to the obtained speech input to the first server (controller sends audio captured by local device to ASR module over network interface, Col 7-8 lines 63-3, e.g. “play songs by Coldplay”, Fig. 6, Col 12 lines 20-40),
receive an energy level of the speech signal, a text corresponding to the speech input, and utterance endpoint information for the speech input from the first server (ASR module determines energy levels of the audio input in one or more spectral bands, Col 10 lines 39-51, text from transcribing the audio data, Col 7 lines 26-30, and semantic information, which includes information regarding the likelihood that a particular utterances has ended, Col 13 lines 44-48, Col 11 lines 20-33, lines 54-60), and
determine whether an utterance of the user has ended based on the energy level and the utterance endpoint information (controller uses semantic information and pause length, which is determined based on energy level, to determine whether the utterance has ended, Col 11-12 lines 54-5, Col 13 lines 26-36, Col 8 lines 60-67).
Consider claim 9, Liu discloses a display device (ASR device having a display, Col 7 lines 1-6), comprising:
a network interface communicating with a first server and a second server (portions of the speech recognition process implemented on remote servers as part of a distributed computing system, Col 2 lines 54-60, ASR module, considered a “first server” connected via network interface to ASR device 302, Col 7 lines 63-66, as well as NLU unit, a “second server”, Col 11 lines 13-19); and
a controller configured to obtain a speech of a user (controller receives speech from user from local device over network, Col 6 lines 20-48),
transmit a speech signal corresponding to the obtained speech input to the first server (controller sends audio captured by local device to ASR module over network interface, Col 7-8 lines 63-3, e.g. “play songs by Coldplay”, Fig. 6, Col 12 lines 20-40),
receive an energy level of the speech signal and a text corresponding to the speech input from the first server, transmit the text to the second server, receive utterance endpoint information for the speech input from the second server (ASR module determines energy levels of the audio input in one or more spectral bands, Col 10 lines 39-51, text from transcribing the audio data, Col 7 lines 26-30, which is transmitted to NLU unit, Fig. 3, Col 11 lines 13-15, orchestrator/controller receives semantic information from NLU unit, which includes information regarding the likelihood that a particular utterances has ended, Col 13 lines 44-48, Col 11 lines 20-33, lines 54-60), and
determine whether an utterance of the user has ended based on the energy level and the utterance endpoint information (controller uses semantic information and pause length, which is determined based on energy level, to determine whether the utterance has ended, Col 11-12 lines 54-5, Col 13 lines 26-36, Col 8 lines 60-67).
Consider claim 5, Liu discloses when it is determined that the utterance of the user has ended, the controller is configured to transmit the text to the second server, receive analysis result information indicating an intent analysis result of the text from the second server, and output the received analysis result information (when a certain about of non-speech is detected, a speculative endpoint is reached and the recognized text is transmitted to NLU module which provides NLU output to orchestrator, Col 9 lines 1-17, which includes an intent classification, Col 11 lines 16-19).
Consider claim 6, Liu discloses when it is determined that the utterance of the user has ended, the controller is configured to ignore an additional speech input until outputting a speech recognition result for the speech input (at speculative endpoints, the orchestrator waits for text from early recognition results to determine how long to wait for more input, Col 9-10 lines 63-9, e.g. the recognizer outputs “play songs by Michael” and ignores “Jackson” until the text for “play songs by Michael” is received and the system determines it should wait longer for more speech, Fig 6C-6D, Col 12-13 lines 41-3).
Consider claim 8, Liu discloses the first server is a Speech To Text (STT) server configured to convert a speech into a text, and the second server is a Natural Language Processing (NLP) server (ASR module 314 and NLU unit 326, Fig. 3, portions of the speech recognition process implemented on remote servers as part of a distributed computing system, Col 2 lines 54-60, ASR module, considered a “first server” connected via network interface to ASR device 302, Col 7 lines 63-66, as well as NLU unit, a “second server”, Col 11 lines 13-19).
Consider claim 13, Liu discloses the controller is further configured to receive analysis result information indicating intent analysis for the speech input from the second server through the network interface (the recognized text is transmitted to NLU module which provides NLU output to orchestrator, Col 9 lines 1-17, which includes an intent classification, Col 11 lines 16-19).
Consider claim 14, Liu discloses when the utterance of the user has ended when the controller determines that the user's utterance is ended, the controller is configured to ignore additional speech input until outputting a speech recognition result for the speech input (at speculative endpoints, the orchestrator waits for text from early recognition results to determine how long to wait for more input, Col 9-10 lines 63-9, e.g. the recognizer outputs “play songs by Michael” and ignores “Jackson” until the text for “play songs by Michael” is received and the system determines it should wait longer for more speech, Fig 6C-6D, Col 12-13 lines 41-3).
Consider claim 15, Liu discloses the first server is a Speech To Text (STT) server configured to convert a speech into a text, and the second server is a Natural Language Processing (NLP) server (ASR module 314 and NLU unit 326, Fig. 3, portions of the speech recognition process implemented on remote servers as part of a distributed computing system, Col 2 lines 54-60, ASR module, considered a “first server” connected via network interface to ASR device 302, Col 7 lines 63-66, as well as NLU unit, a “second server”, Col 11 lines 13-19).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-4 and 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 9437186) in view of Choi et al. (US 20210074290), which is U.S. Patent Publication reference 2 on Applicant’s 12/27/24 IDS.
Consider claim 2, Liu discloses the controller is configured to determine that the utterance has ended when the utterance endpoint information includes a value indicating that the utterance is an endpoint (the utterance is ended when semantic tags indicate the end, Col 13 lines 32-36).
Liu does not specifically mention determining that the utterance has ended when the energy level is less than a preset first level.
Choi discloses determining that the utterance has ended when the energy level is less than a preset first level (energy threshold for identifying an end time, [0023]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Liu by determining that the utterance has ended when the energy level is less than a preset first level in order to improve the accuracy and efficiency of voice activity detection, as suggested by Choi ([0187]). Doing so would have led to predictable results of avoiding unnecessary speech recognition processing of audio portions not containing speech, as suggested by Choi ([0004]). The references cited are analogous art in the same field of speech processing.
Consider claim 3, Liu discloses the controller is configured to determine that the utterance has ended when a reliability score indicating the utterance endpoint included in the utterance endpoint information is equal to or higher than a preset score (thresholding techniques for likelihood that the audio does not contain speech, i.e. that the speech has ended Col 10 lines 28-45).
Liu does not specifically mention determining that the utterance has ended when the energy level is lower than a preset first level.
Choi discloses determining that the utterance has ended when the energy level is lower than a preset first level (identifying end of speech when energy is less than a threshold, [0055]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Liu by determining that the utterance has ended when the energy level is lower than a preset first level for reasons similar to those for claim 2.
Consider claim 4, Liu discloses the controller is configured to determine that the utterance has not ended when a reliability score indicating the utterance endpoint included in the utterance endpoint information is equal to or higher than a preset score (thresholding techniques for likelihood that the audio contains speech, i.e. that the speech has not ended, Col 10 lines 28-45).
Liu does not specifically mention determining that the utterance has ended when the energy level is higher than a preset second level.
Choi discloses determining that the utterance has not ended when the energy level is higher than a preset second level (identifying voice section, i.e. that the speech has not ended, when energy is greater than a threshold, [0055]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Liu by determining that the utterance has ended when the energy level is higher than a preset second level for reasons similar to those for claim 2.
Consider claim 10, Liu discloses the controller is configured to determine that the utterance has ended when the utterance endpoint information includes a value indicating that the utterance is an endpoint (the utterance is ended when semantic tags indicate the end, Col 13 lines 32-36).
Liu does not specifically mention determining that the utterance has ended when the energy level is less than a preset first level.
Choi discloses determining that the utterance has ended when the energy level is less than a preset first level (energy threshold for identifying an end time, [0023]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Liu by determining that the utterance has ended when the energy level is less than a preset first level for reasons similar to those for claim 2.
Consider claim 11, Liu discloses the controller is configured to determine that the utterance has ended when a reliability score indicating the utterance endpoint included in the utterance endpoint information is equal to or higher than a preset score (thresholding techniques for likelihood that the audio does not contain speech, i.e. that the speech has ended Col 10 lines 28-45).
Liu does not specifically mention determining that the utterance has ended when the energy level is lower than a preset first level.
Choi discloses determining that the utterance has ended when the energy level is lower than a preset first level (identifying end of speech when energy is less than a threshold, [0055]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Liu by determining that the utterance has ended when the energy level is lower than a preset first level for reasons similar to those for claim 2.
Consider claim 12, Liu discloses the controller is configured to determine that the utterance has not ended when a reliability score indicating the utterance endpoint included in the utterance endpoint information is equal to or higher than a preset score (thresholding techniques for likelihood that the audio contains speech, i.e. that the speech has not ended, Col 10 lines 28-45).
Liu does not specifically mention determining that the utterance has ended when the energy level is higher than a preset second level.
Choi discloses determining that the utterance has not ended when the energy level is higher than a preset second level (identifying voice section, i.e. that the speech has not ended, when energy is greater than a threshold, [0055]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Liu by determining that the utterance has ended when the energy level is higher than a preset second level for reasons similar to those for claim 2.
Claims 7 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 9437186) in view of Parthasarathi et al. (US 20200035231).
Consider claim 7, Liu discloses the controller is configured to transmit the signal to the first server through the network interface (controller sends audio captured by local device to ASR module over network interface, Col 7-8 lines 63-3, e.g. “play songs by Coldplay”, Fig. 6, Col 12 lines 20-40).
Liu does not specifically mention converting the speech signal into a pulse code modulation (PCM) signal and transmit the converted PCM signal.
Parthasarathi discloses converting the speech signal into a pulse code modulation (PCM) signal and transmit the converted PCM signal (the mechanical sound wave comprising the audio is converted to PCM data which is sent to a downstream remote device for further processing, [0028]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Liu by converting the speech signal into a pulse code modulation (PCM) signal and transmit the converted PCM signal in order to facilitate processing with a distributed computing environment, as suggested by Parthasarathi ([0028]), predictably alleviating the known computational expense of ASR and NLU, as suggested by Parthasarathi ([0028]). The references cited are analogous art in the same field of speech processing.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20190130907 Joh discloses using context to determine end points of user speech within a vehicle
US 20190198012 Zhang discloses determining whether spoken words parsed are a complete parse
US 20180350395 Simko discloses using confidence measure to detect the end point of a spoken query
US 6782363 Lee discloses performing real-time endpoint detection in automatic speech recognition
US 20060241948 Abrash discloses augmenting audio to locate starting and ending speech points for obtaining complete speech signals
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Jesse S Pullias/
Primary Examiner, Art Unit 2655 08/31/26