DETAILED ACTION
DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 21-29 and 31-39 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Helwani et al. (US 10,878,812).
The applied reference has a common assignee with the instant application. Based upon the earlier effectively filed date of the reference, it constitutes prior art under 35 U.S.C. 102(a)(2). This rejection under 35 U.S.C. 102(a)(2) might be overcome by: (1) a showing under 37 CFR 1.130(a) that the subject matter disclosed in the reference was obtained directly or indirectly from the inventor or a joint inventor of this application and is thus not prior art in accordance with 35 U.S.C. 102(b)(2)(A); (2) a showing under 37 CFR 1.130(b) of a prior public disclosure under 35 U.S.C. 102(b)(2)(B) if the same invention is not being claimed; or (3) a statement pursuant to 35 U.S.C. 102(b)(2)(C) establishing that, not later than the effective filing date of the claimed invention, the subject matter disclosed in the reference and the claimed invention were either owned by the same person or subject to an obligation of assignment to the same person or subject to a joint research agreement.
Regarding Claim 21, Helwant et al discloses a computer-implemented method, comprising:
capturing, by a first device, first audio representing a first spoken input by a first user (At an operation 502, the remote system receives a first audio signal generated by a first device in an environment, with the first audio signal corresponding to first user speech) (col. 19, lines 17-28);
determining, by the first device, first audio data representing the first audio (At an operation 502, the remote system receives a first audio signal generated by a first device in an environment, with the first audio signal corresponding to first user speech) (col. 19, lines 17-28);
causing speech processing to be performed using the first audio data (FIG. 7 illustrates a block diagram of an example architecture of a remote system 110 which receives audio data 112 and metadata from voice-enabled devices 108, and performs processing techniques to determine which of the voice-enabled devices 108 is to respond to a voice command associated of a user 104 represented in the audio data 112) (col. 21, lines 29-34);
determining a second device associated with the first device is to capture further audio corresponding to further spoken inputs by the first user (At an operation 614, the remote system determines whether the speaker IDs are the same for each of the audio signals) (col. 20, lines 29-42);
capturing, by a second device, second audio representing a second spoken input by the first user (At an operation 504, the remote system receives a second audio signal generated by a second device in the environment, with the second audio signal also corresponding to the first user speech) (col. 19, lines 17-28);
determining, by the second device, second audio data representing the second audio (At an operation 504, the remote system receives a second audio signal generated by a second device in the environment, with the second audio signal also corresponding to the first user speech) (col. 19, lines 17-28); and
causing speech processing to be performed using the second audio data (FIG. 7 illustrates a block diagram of an example architecture of a remote system 110 which receives audio data 112 and metadata from voice-enabled devices 108, and performs processing techniques to determine which of the voice-enabled devices 108 is to respond to a voice command associated of a user 104 represented in the audio data 112) (col. 21, lines 29-34).
Regarding Claim 22, Helwant et al discloses the computer-implemented method, further comprising: determining network data corresponding to a network connection of the second device, wherein determining the second device is to capture further audio is based at least in part on the network data (Each of the voice-enabled devices may access the speech-processing system through a communications network, such as the internet, to provide the speech-processing system with the captured audio data and the various types of contextual information detected, determined, etc., by the voice-enabled devices. In various examples, the voice-enabled devices may receive a “wake” trigger (e.g., wake word, button input, etc.) which indicates to the voice-enabled devices that a user is speaking a command, and the voice-enabled devices begin sending audio data representing the spoken command to the network-based speech service) (col. 8, lines 8-25).
Regarding Claim 23, Helwant et al discloses the computer-implemented method, further comprising: determining position data corresponding to the second device (he remote system 110 may proceed to determine at an operation 608 whether each received audio signal is associated with a common location) (col. 20, lines 7-42), wherein determining the second device is to capture further audio is based at least in part on the position data (For example, a user's house or building may include multiple levels, rooms, or the like, and each of these locations may be associated with different models of the respective environment. Therefore, if the received audio signals are associated with different locations (e.g., a first signal is from a first floor of a building and a second signal is from a second floor of the building), then at an operation 610 the remote system 110 may process the audio signals independently) (col. 20, lines 7-42).
Regarding Claim 24, Helwant et al discloses the computer-implemented method, further comprising: before determining the second device is to capture further audio, determining the first device corresponds to a voice call (performing a telephone call) (col. 2, lines 9-30), wherein determining the second device is to capture further audio is based at least in part on the first device corresponding to a voice call (the voice-enabled devices 108 may detect a predefined trigger expression or word (e.g., “awake”), which may be followed by instructions or directives (e.g., “please end my phone call,” “please turn off the alarm,” etc.).) (col. 8, line 61-col. 9, line 29).
Regarding Claim 25, Helwant et al discloses the computer-implemented method, wherein:
the first device is associated with a first profile (At an operation 614, the remote system determines whether the speaker IDs are the same for each of the audio signals) (col. 20, lines 29-43);
the second device is associated with the first profile (At an operation 614, the remote system determines whether the speaker IDs are the same for each of the audio signals) (col. 20, lines 29-43); and
determining the second device is to capture further audio is based at least in part on the first device and the second device being associated with the first profile (If, however, each audio signal is associated with a common speaker ID, then at an operation 616 the remote system may retrieve a current topology of the user environment. At an operation 618, the remote system 110 generates one or more features of the audio signals as described above) (col. 20, lines 29-43).
Regarding Claim 26, Helwant et al discloses the computer-implemented method, further comprising:
receiving, by the second device, a model corresponding to the first user and configured to be used to process audio data (In some instances, the device-selection component 126 may utilize a model-generation component 130 and a user-localization component 146 to localize the user 104 within the environment for the purpose of selecting a device to respond to the user speech 106) ( col. 11, lines 15-33); and
processing, by the second device, the second audio data using the model (Thereafter, a device associated with that particular region may be selected for responding to the user speech) (col. 11, lines 40-49).
Regarding Claim 27, Helwant et al discloses the computer-implemented method, wherein causing speech processing to be performed using the second audio data comprises causing speech processing to be performed using the second audio data, as the second audio data originated with the first device () (col. 19, lines 17-28).
Regarding Claim 28, Helwant et al discloses the computer-implemented method, further comprising, prior to capturing, by the second device, the second audio: receiving, by the second device, an instruction to process audio corresponding to a detected spoken input (At an operation 504, the remote system receives a second audio signal generated by a second device in the environment, with the second audio signal also corresponding to the first user speech.) (col. 19, lines 17-28).
Regarding Claim 29, Helwant et al discloses the computer-implemented method, further comprising, prior to capturing, by the second device, the second audio: outputting a notification indicating that the second device is configured for speech capture (For example, envision that a user states “wakeup, what time is my next meeting?” within a user environment that includes multiple voice-enabled devices configured to receive data representing the appropriate answer and output (e.g., visually, audibly, etc.)) (col. 2, lines 37-43).
Regarding Claim 31, Helwant et al discloses a system comprising:
at least one processor (As used herein, a processor, such as processor(s) 116 and/or 800, may include multiple processors and/or a processor having multiple cores) (col. 31, lines 5-26); and
at least one memory comprising instructions that, when executed by the at least one processor (Additionally, each of the processor(s) 116 and/or 800 may possess its own local memory, which also may store program components, program data, and/or one or more operating systems) (col. 31, lines 5-26), cause the system to:
capture, by a first device, first audio representing a first spoken input by a first user (At an operation 502, the remote system receives a first audio signal generated by a first device in an environment, with the first audio signal corresponding to first user speech) (col. 19, lines 17-28);
determine, by the first device, first audio data representing the first audio (At an operation 502, the remote system receives a first audio signal generated by a first device in an environment, with the first audio signal corresponding to first user speech) (col. 19, lines 17-28);
cause speech processing to be performed using the first audio data (FIG. 7 illustrates a block diagram of an example architecture of a remote system 110 which receives audio data 112 and metadata from voice-enabled devices 108, and performs processing techniques to determine which of the voice-enabled devices 108 is to respond to a voice command associated of a user 104 represented in the audio data 112) (col. 21, lines 29-34);
determine a second device associated with the first device is to capture further audio corresponding to further spoken inputs by the first user (At an operation 614, the remote system determines whether the speaker IDs are the same for each of the audio signals) (col. 20, lines 29-42);
capture, by a second device, second audio representing a second spoken input by the first user (At an operation 504, the remote system receives a second audio signal generated by a second device in the environment, with the second audio signal also corresponding to the first user speech) (col. 19, lines 17-28);
determine, by the second device, second audio data representing the second audio (At an operation 504, the remote system receives a second audio signal generated by a second device in the environment, with the second audio signal also corresponding to the first user speech) (col. 19, lines 17-28); and
cause speech processing to be performed using the second audio data (FIG. 7 illustrates a block diagram of an example architecture of a remote system 110 which receives audio data 112 and metadata from voice-enabled devices 108, and performs processing techniques to determine which of the voice-enabled devices 108 is to respond to a voice command associated of a user 104 represented in the audio data 112) (col. 21, lines 29-34).
Claim 32 is rejected for the same reason as claim 22.
Claim 33 is rejected for the same reason as claim 23.
Claim 34 is rejected for the same reason as claim 24.
Claim 35 is rejected for the same reason as claim 25.
Claim 36 is rejected for the same reason as claim 26.
Claim 37 is rejected for the same reason as claim 27.
Claim 38 is rejected for the same reason as claim 28.
Claim 39 is rejected for the same reason as claim 29.
Allowable Subject Matter
Claims 30 and 40 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Cited Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Fineberg et al. (US 9,729,821) discloses sensor fusion for location based device grouping.
Sundaram et al. (US 10,121,494) discloses user presence detection
Roy et al. (US 10,504,520) discloses voice controlled communication requests and responses.
You et al. (US 2008/0144788) discloses performing voice communication in mobile terminal.
Seshadri et al. (US 2009/0079622) discloses sharing of GPS information in between mobile devices.
Faaborg (US 2015/0370531) discloses device designation for audio input monitoring.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SATWANT K SINGH whose telephone number is (571)272-7468. The examiner can normally be reached Monday thru Friday 9:00 AM to 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D Shah can be reached at (571}270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SATWANT K SINGH/Primary Examiner, Art Unit 2653