Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 20 July 2026 has been entered.
Response to Arguments
Applicant’s arguments with respect to the rejection(s) of claim(s) 1 under 35 U.S.C. 103 have been fully considered and are, in part, persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Lien et al.
Applicant’s arguments with respect to the limitation of “creating a wireless communications link between the device and the smart assistant device” have been fully considered, but they are not persuasive.
Applicant argues (see page 10 of Applicant’s reply) that Yakishyn discloses a single electronic device that receives voice input using a microphone and acquires a user’s gesture. Thus, as alleged by Applicant, because both inputs are acquired internally by this single electronic device, Yakishyn fails to disclose “creating a wireless communications link between the device and the smart assistant device”.
However, Yakishyn discloses wireless communication unit 1510 (paragraph [101]) may “transmit/receive various data required to process a voice input based on a gesture. For example, when a gesture corresponding to a voice input is input by an external device (not shown), the communications unit 1500 may receive the gesture input from an external device (not shown).” (paragraph [104], emphasis added). Contrary to Applicant’s assertions, Yakishyn expressly discloses creating a wireless communication link between the device (the external gesture input device) and the smart assistant (through wireless communication unit 1510 of device 1000).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 7-14, 16-23, and 25-30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yakishyn et al. (WIPO WO2021/187653, hereinafter “Yakishyn”), in view of Lein et al. (U.S. Patent Application Pub. No. 2020/0393912, hereinafter “Lein”).
In regard to claim 1, Yakishyn discloses a method for instructing a smart assistant device to perform an action, the method comprising:
receiving, by a microphone, an utterance including a trigger word from a user (Fig. 6, 610, an electronic device receives a voice input, paragraph [120]; via a microphone, paragraph [108]; the voice including a trigger word such as “this” which is ambiguous, paragraphs [124-125]);
determining that the user performed a gesture while making the utterance received by the microphone (a user’s gesture is acquired, paragraph [121]);
in response to determining that the utterance includes the trigger word, determining a direction of the gesture (if an ambiguous word is found in text recognized from the input speech, additional information is acquired according to the gesture, paragraphs [125-126]; the additional information including a gesture pointing to an object in a particular direction, paragraphs [144-146]);
determining an object external to the smart assistant device and associated with the gesture based on the direction (the object being pointed to is determined, paragraphs [144-146]);
creating a wireless communications link between the device and the smart assistant device (a wireless communications unit receives gesture input from an external device, paragraphs [101] and [104]); and
transmitting an enhanced directive to an application programming interface (API) of the smart assistance assistant device, the enhanced directive based on the object, the gesture, and the utterance,
wherein the enhanced directive causes the smart assistant device to perform an action to control the object (the voice command is enhanced with unambiguous location information for the object referred to by the gesture, and a device is sent instructions to perform an operation, paragraphs [125-127]).
Yakishyn discloses the gesture may be captured by any type of sensor (see, e.g., paragraph [51]), but does not expressly disclose using radio frequency sensing.
Lien discloses determining, based on radio frequency sensing at an antenna array, that the user performed a gesture (see Fig. 4, an array of antennas 406a-d of a gesture sensor component 106 detect an input gesture, paragraph [0043]); and
determining a direction of the gesture based on the radio frequency sensing via the antenna array (the gesture sensor component 106 detects directional gestures, paragraph [0023]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize radio frequency sensing to capture the input gestures, because it would capture gestures that could not be captured by a camera, as taught by Lien (paragraph [0038]).
In regard to claim 2, Yakishyn discloses the trigger word indicates an association of the gesture with the object (location information of an object corresponding to the trigger word, paragraph [126]).
In regard to claim 3, Yakishyn discloses determining a motion associated with the gesture (gestures include motions, paragraph [50]).
In regard to claim 4, Yakishyn discloses determining a motion associated with the gesture (gestures include motions, paragraph [50]);
determining a relative amount associated with the motion (a relative distance of the object, paragraphs [130-131]);
converting the relative amount to an amount that is understood by the object (identification information of the object at a relative distance determined according to the trigger word, paragraphs [130-131]); and
including the amount in the enhanced directive (the identified object is included in the request information, paragraphs [130-131]).
In regard to claim 5, Yakishyn discloses determining the relative amount associated with the motion comprises one of:
determining a first distance between a thumb and a forefinger of a hand of the user;
determining a second distance between a left palm and a right palm of the user; or
determining a third distance between a starting position of the gesture and an ending position of the gesture (see Fig. 10, gestures 1070, 1080, and 1090, used as gestures for indicating a quantity between forefinger and thumb, etc., paragraph [155]).
In regard to claim 7, Yakishyn discloses the gesture comprises pointing or gesturing towards the object (gestures point at the objects, paragraphs [153-154]); and
the utterance comprises the action associated with the object (the voice input comprises an operation request associated with the object, paragraphs [126-127]).
In regard to claim 8, Yakishyn discloses the action comprises on, off, dim, brighten, increase, decrease, play, stop, pause, positioning of an audio object, or any combination thereof (e.g., increase or decrease volume, paragraphs [144-146]; turn lights on, paragraph [285]; etc.).
In regard to claim 9, Yakishyn discloses the object comprises a light source, a media playback device, a set of blinds or shutters, a controllable object, a heating ventilation air conditioning (HVAC) controller, or any combination thereof (a controllable TV, paragraphs [144-146]; lights, paragraph [285], etc.).
In regard to claim 10, Yakishyn discloses a device (Fig. 5, 1000) comprising:
a memory (memory 1700, paragraph [83]);
at least one transceiver (communications unit 1500, paragraph [100]); and
at least one processor communicatively coupled to the memory and the at least one transceiver (processor 1300, paragraph [92]), the at least one processor configured to:
receive, by a microphone, an utterance including a trigger word from a user (Fig. 6, 610, an electronic device receives a voice input, paragraph [120]; via a microphone, paragraph [108]; the voice including a trigger word such as “this” which is ambiguous, paragraphs [124-125]);
determine that the user performed a gesture while making the utterance received by the microphone (a user’s gesture is acquired, paragraph [121]);
in response to determining that the utterance includes the trigger word, determine a direction of the gesture (if an ambiguous word is found in text recognized from the input speech, additional information is acquired according to the gesture, paragraphs [125-126]; the additional information including a gesture pointing to an object in a particular direction, paragraphs [144-146]);
determine an object external to a smart assistant device and associated with the gesture based on the direction (the object being pointed to is determined, paragraphs [144-146]);
create a wireless communications link between the device and the smart assistant device (a wireless communications unit receives gesture input from an external device, paragraphs [101] and [104]); and
wirelessly transmit, in cooperation with the at least one transceiver, an enhanced directive to an application programming interface (API) of the smart assistant device, the enhanced directive based on the object, the gesture, and the utterance, wherein the enhanced directive causes the smart assistant device to perform an action to control the object (the voice command is enhanced with unambiguous location information for the object referred to by the gesture, and a device is sent instructions to perform an operation, paragraphs [125-127]; via wireless transmission, paragraphs [100-101]).
Yakishyn discloses the gesture may be captured by any type of sensor (see, e.g., paragraph [51]), but does not expressly disclose using radio frequency sensing.
Lien discloses determining, based on radio frequency sensing at an antenna array, that the user performed a gesture (see Fig. 4, an array of antennas 406a-d of a gesture sensor component 106 detect an input gesture, paragraph [0043]); and
determining a direction of the gesture based on the radio frequency sensing via the antenna array (the gesture sensor component 106 detects directional gestures, paragraph [0023]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize radio frequency sensing to capture the input gestures, because it would capture gestures that could not be captured by a camera, as taught by Lien (paragraph [0038]).
In regard to claim 11, Yakishyn discloses the trigger word indicates an association of the gesture with the object (location information of an object corresponding to the trigger word, paragraph [126]).
In regard to claim 12, Yakishyn discloses determining a motion associated with the gesture (gestures include motions, paragraph [50]).
In regard to claim 13, Yakishyn discloses determining a motion associated with the gesture (gestures include motions, paragraph [50]);
determining a relative amount associated with the motion (a relative distance of the object, paragraphs [130-131]);
converting the relative amount to an amount that is understood by the object (identification information of the object at a relative distance determined according to the trigger word, paragraphs [130-131]); and
including the amount in the enhanced directive (the identified object is included in the request information, paragraphs [130-131]).
In regard to claim 14, Yakishyn discloses determining the relative amount associated with the motion comprises one of:
determining a first distance between a thumb and a forefinger of a hand of the user;
determining a second distance between a left palm and a right palm of the user; or
determining a third distance between a starting position of the gesture and an ending position of the gesture (see Fig. 10, gestures 1070, 1080, and 1090, used as gestures for indicating a quantity between forefinger and thumb, etc., paragraph [155]).
In regard to claim 16, Yakishyn discloses the gesture comprises pointing or gesturing towards the object (gestures point at the objects, paragraphs [153-154]); and
the utterance comprises the action associated with the object (the voice input comprises an operation request associated with the object, paragraphs [126-127]).
In regard to claim 17, Yakishyn discloses the action comprises on, off, dim, brighten, increase, decrease, play, stop, pause, positioning of an audio object, or any combination thereof (e.g., increase or decrease volume, paragraphs [144-146]; turn lights on, paragraph [285]; etc.).
In regard to claim 18, Yakishyn discloses the object comprises a light source, a media playback device, a set of blinds or shutters, a controllable object, a heating ventilation air conditioning (HVAC) controller, or any combination thereof (a controllable TV, paragraphs [144-146]; lights, paragraph [285], etc.).
In regard to claim 19, Yakishyn discloses an apparatus comprising:
means for receiving an utterance from a user, the utterance including a trigger word (Fig. 6, 610, an electronic device receives a voice input, paragraph [120]; via a microphone, paragraph [108]; the voice including a trigger word such as “this” which is ambiguous, paragraphs [124-125]);
means for determining that the user performed a gesture while making the utterance received by the microphone (a user’s gesture is acquired, paragraph [121]);
means for determining a direction of the gesture in response to determining that the utterance includes the trigger word (if an ambiguous word is found in text recognized from the input speech, additional information is acquired according to the gesture, paragraphs [125-126]; the additional information including a gesture pointing to an object in a particular direction, paragraphs [144-146]);
means for determining an object external to the smart assistant device and associated with the gesture based on the direction (the object being pointed to is determined, paragraphs [144-146]);
means for creating a wireless communications link between the device and the smart assistant device (a wireless communications unit receives gesture input from an external device, paragraphs [101] and [104]); and
means for transmitting an enhanced directive to an application programming interface (API) of the smart assistance assistant device, the enhanced directive based on the object, the gesture, and the utterance, wherein the enhanced directive causes the smart assistant device to perform an action to control the object (the voice command is enhanced with unambiguous location information for the object referred to by the gesture, and a device is sent instructions to perform an operation, paragraphs [125-127]).
Yakishyn discloses the gesture may be captured by any type of sensor (see, e.g., paragraph [51]), but does not expressly disclose using radio frequency sensing.
Lien discloses determining, based on radio frequency sensing at an antenna array, that the user performed a gesture (see Fig. 4, an array of antennas 406a-d of a gesture sensor component 106 detect an input gesture, paragraph [0043]); and
determining a direction of the gesture based on the radio frequency sensing via the antenna array (the gesture sensor component 106 detects directional gestures, paragraph [0023]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize radio frequency sensing to capture the input gestures, because it would capture gestures that could not be captured by a camera, as taught by Lien (paragraph [0038]).
In regard to claim 20, Yakishyn discloses means for determining that the utterance includes the trigger word (Fig. 6, 610, an electronic device receives a voice input, paragraph [120]; via a microphone, paragraph [108]; the voice including a trigger word such as “this” which is ambiguous, paragraphs [124-125]).
In regard to claim 21, Yakishyn discloses means for determining a motion associated with the gesture (gestures include motions, paragraph [50]).
In regard to claim 22, Yakishyn discloses means for determining a motion associated with the gesture (gestures include motions, paragraph [50]);
means for determining a relative amount associated with the motion (a relative distance of the object, paragraphs [130-131]);
means for converting the relative amount to an amount that is understood by the object (identification information of the object at a relative distance determined according to the trigger word, paragraphs [130-131]); and
means for including the amount in the enhanced directive (the identified object is included in the request information, paragraphs [130-131]).
In regard to claim 23, Yakishyn discloses determining the relative amount associated with the motion comprises one of:
means for determining a first distance between a thumb and a forefinger of a hand of the user;
means for determining a second distance between a left palm and a right palm of the user; or
means for determining a third distance between a starting position of the gesture and an ending position of the gesture (see Fig. 10, gestures 1070, 1080, and 1090, used as gestures for indicating a quantity between forefinger and thumb, etc., paragraph [155]).
In regard to claim 25, Yakishyn discloses the gesture comprises pointing or gesturing towards the object (gestures point at the objects, paragraphs [153-154]); and
the utterance comprises the action associated with the object (the voice input comprises an operation request associated with the object, paragraphs [126-127]).
In regard to claim 26, Yakishyn discloses the action comprises on, off, dim, brighten, increase, decrease, play, stop, pause, positioning of an audio object, or any combination thereof (e.g., increase or decrease volume, paragraphs [144-146]; turn lights on, paragraph [285]; etc.).
In regard to claim 27, Yakishyn discloses the object comprises a light source, a media playback device, a set of blinds or shutters, a controllable object, a heating ventilation air conditioning (HVAC) controller, or any combination thereof (a controllable TV, paragraphs [144-146]; lights, paragraph [285], etc.).
In regard to claim 28, Yakishyn discloses a non-transitory computer-readable storage medium to store instructions executable by one or more processors (paragraph [292]) to:
receive, by a microphone, an utterance including a trigger word from a user (Fig. 6, 610, an electronic device receives a voice input, paragraph [120]; via a microphone, paragraph [108]; the voice including a trigger word such as “this” which is ambiguous, paragraphs [124-125]);
determine that the user performed a gesture while making the utterance received by the microphone (a user’s gesture is acquired, paragraph [121]);
in response to determining that the utterance includes the trigger word, determine a direction of the gesture (if an ambiguous word is found in text recognized from the input speech, additional information is acquired according to the gesture, paragraphs [125-126]; the additional information including a gesture pointing to an object in a particular direction, paragraphs [144-146]);
determine an object external to the device and the smart assistant device and associated with the gesture based on the direction (the object being pointed to is determined, paragraphs [144-146]); and
transmit an enhanced directive to an application programming interface (API) of the smart assistance assistant device, the enhanced directive based on the object, the gesture, and the utterance, wherein the enhanced directive causes the smart assistant device to perform an action to control the object (the voice command is enhanced with unambiguous location information for the object referred to by the gesture, and a device is sent instructions to perform an operation, paragraphs [125-127]).
Yakishyn discloses the gesture may be captured by any type of sensor (see, e.g., paragraph [51]), but does not expressly disclose using radio frequency sensing.
Lien discloses determining, based on radio frequency sensing at an antenna array, that the user performed a gesture (see Fig. 4, an array of antennas 406a-d of a gesture sensor component 106 detect an input gesture, paragraph [0043]); and
determining a direction of the gesture based on the radio frequency sensing via the antenna array (the gesture sensor component 106 detects directional gestures, paragraph [0023]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize radio frequency sensing to capture the input gestures, because it would capture gestures that could not be captured by a camera, as taught by Lien (paragraph [0038]).
In regard to claim 29, Yakishyn discloses determining a motion associated with the gesture (gestures include motions, paragraph [50]).
In regard to claim 30, Yakishyn discloses the gesture comprises pointing or gesturing towards the object (gestures point at the objects, paragraphs [153-154]); and
the utterance comprises the action associated with the object (the voice input comprises an operation request associated with the object, paragraphs [126-127]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIAN LOUIS ALBERTALLI whose telephone number is (571)272-7616. The examiner can normally be reached M-F 8AM-3PM, 4PM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
BLA 7/31/26
/BRIAN L ALBERTALLI/ Primary Examiner, Art Unit 2656