DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1 - 20 are pending and claims 1, 10 and 19 are independent claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, and 6-9, 10, and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over Srinivasa Pat No. US 12626697 B2 (Srinivasa) in view of Sreepathihalli et al. Pat App No. US 20200053257 A1 (Sreepathihalli), and further in view of Milleret al. Pat App No US 20240403081 A1 (Miller).
Regarding Claim 1. Srinivasa discloses a method of activating a device (col 9, ln 25-27, The electronic device 101 further includes …an activation state of the electronic device 101), comprising:
detecting, by a hardware processor of the device, a first user utterance specifying a first keyword of a multi-keyword phrase from audio data (Srinivasa, col 1, ln 47-48, keyword detection model configured to predict a first likelihood that the audio data includes speech; Srinivasa, col 7, ln 50 - col 8, ln 17, FIG. 1 illustrates an example network configuration 100 including an electronic device… an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, or a sensor 180… the processor 120 may receive and process inputs (such as audio inputs or data received from an audio input device like a microphone) and perform keyword detection and/or automated speech recognition tasks using the inputs.);
in response to the detecting the first user utterance (Srinivasa, col 23, ln 51-52, in response to the first likelihood exceeding a first threshold,),
monitoring, by the hardware processor, the audio data for a second user utterance specifying a second keyword of the multi-keyword phrase (Srinivasa, col 24, ln 4-9, Each audio data sample can be annotated with a speech label indicating whether the audio data sample includes speech, a keyword-like label indicating whether the audio data sample includes keyword-like speech, and a keyword label identifying which keyword, if any, is in the audio data sample); and
Srinivasa does not specifically disclose monitoring sensor data generated by a user attention sensor of the device for an indication of user attention directed to the device.
However, Sreepathihalli, in the same field of endeavor, discloses monitoring sensor data generated by a user attention sensor of the device for an indication of user attention directed to the device (Sreepathihalli, para 0020-0022, the electronic device may differentiate when a user is present versus when an animal, such as a pet, runs in front of the depth sensor (e.g., as opposed to a motion sensor or a depth sensor that monitors a single zone). Compared to a camera, the resolution of the depth sensor data (e.g., 4×4, 8×8, 16×16, or 24×24) is significantly lower, such that most physical features of a user may not be identifiable in the depth sensor data… the electronic device may determine whether the user's attention is on the electronic device); and
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Sreepathihalli in the method of Srinivasa because this would enable monitoring a user’s attention using depth sensors which in turn enables, for instance, if the user is determined to be present (i.e., if it is determined that the user's attention is on the device), then the device causes a display to enter an operational mode, or otherwise, the device causes the display to enter a standby mode (Sreepathihalli, Abstract).
Srinivasa in view of Sreepathihalli do not specifically disclose in response to detecting the second keyword and detecting the indication of user attention directed to the device, initiating, by the hardware processor, a selected operation of the device.
However, Miller, in the same field of endeavor, discloses in response to detecting the second keyword and detecting the indication of user attention directed to the device (Miller, para 0508, Device 900 detects input 2350m (e.g., a touch input, an air gesture, a mouse click, a gaze, and/or a speech input) directed at select account affordance 2338; [i.e., “gaze” as “detecting the indication of user attention”; and “speech input “ as “utterance specifying a keyword/keywords”]), initiating, by the hardware processor, a selected operation of the device (Miller, para 0508-0509, Device 900 detects input 2350m (e.g., a touch input, an air gesture, a mouse click, a gaze, and/or a speech input) directed at select account affordance 2338. At FIG. 23N, in response to detecting input 2350m, device 900 displays account selection interface 2310n; [i.e., for this particular application, the device operation initiated is to display the account selection interface, which would not be displayed unless those two criteria (i.e., the detection of the gaze/attention and keywords/speech input) are satisfied] ).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Miller in the method of Srinivasa in view of Sreepathihalli because this would enable providing electronic devices with faster, more efficient methods and interfaces for displaying resource transfer user interfaces through the detection of a touch input, an air gesture, a mouse click, a gaze, and/or a speech input (Miller, para 0005 and 0509).
Regarding Claim 6. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 1, wherein the user attention sensor is a red-green-blue (RGB) camera (Srinivasa, col 9, ln 25 - 58, The electronic device 101 further includes one or more sensors 180 that can meter a physical quantity or detect an activation state of the electronic device 101 and convert metered or detected information into an electrical signal. The sensor(s) 180 can also include one or more buttons for touch input, one or more microphones, a gesture sensor, a gyroscope or gyro sensor, an air pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (such as an RGB sensor) … Any of these sensor(s) 180 can be located within the electronic device 101... The electronic device 101 can also be an augmented reality wearable device, such as eyeglasses, that include one or more cameras).
Regarding Claim 7. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 6, wherein the user attention sensor is an infrared camera (Srinivasa, col 9, ln 25 - 58, The electronic device 101 further includes one or more sensors 180 that can meter a physical quantity or detect an activation state of the electronic device 101 and convert metered or detected information into an electrical signal. The sensor(s) 180 can also include… an infrared (IR) sensor… Any of these sensor(s) 180 can be located within the electronic device 101... The electronic device 101 can also be an augmented reality wearable device, such as eyeglasses, that include one or more cameras).
Regarding Claim 8. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 1, wherein the selected operation includes waking the device from a low power mode (Srinivasa, col 11, ln 11-20, As shown in FIG. 2, the system 200 includes the electronic device 101, which includes the processor 120. The processor 120 is operatively coupled to or otherwise configured to use one or more machine learning models, such as one or more keyword detection models 202. The keyword detection model 202 can be trained to recognize one or more keywords or phrases. For example, keywords can include wake words or phrases (e.g., “hey GOOGLE” or “hey BIXBY”), which are used to wake up a voice assistant of an electronic device from a low power or sleep mode).
Regarding Claim 9. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 1, wherein the selected operation includes responding to a further user utterance specifying a command (Srinivasa, col 11, ln 11 - 28, As shown in FIG. 2, the system 200 includes the electronic device 101, which includes the processor 120. The processor 120 is operatively coupled to or otherwise configured to use one or more machine learning models, such as one or more keyword detection models 202. The keyword detection model 202 can be trained to recognize one or more keywords or phrases…Additionally or alternatively, the keyword detection model 202 can detect predetermined keywords or phrases other than wake words or phrases, such as specific device commands that serve to both wake up the voice assistant and trigger further action, such as a preset command to “call Mom,” a preset command to play a song, a preset command to set a TV device to a particular setting or channel, a preset command to set an oven to a given temperature, and so on).
Regarding Claim 10. Srinivasa discloses a device (Srinivasa, col 1, ln 60 - 61, an electronic device includes at least one processing device), comprising:
a microphone capable of detecting sound (Srinivasa, col 8, ln 13 - 15, the processor 120 may receive and process inputs (such as audio inputs or data received from an audio input device like a microphone));
a hardware processor coupled to the microphone and the user attention sensor, wherein the hardware processor is capable of executing operations (Srinivasa, col 18, ln 9 - 15, As shown in FIG. 11, the architecture 1100 includes multiple audio sources 1101 (e.g., audio input devices of the electronic device 101 such as microphones) used to source audio data. The information received from the multiple audio sources 1101 is provided to the preprocessing operation 1102. This can include the processor executing the preprocessing operation) including:
detecting, from audio data generated by the microphone, a first user utterance specifying a keyword phrase (Srinivasa, col 1, ln 47-48, keyword detection model configured to predict a first likelihood that the audio data includes speech; Srinivasa, col 7, ln 50 - col 8, ln 17, FIG. 1 illustrates an example network configuration 100 including an electronic device… an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, or a sensor 180… the processor 120 may receive and process inputs (such as audio inputs or data received from an audio input device like a microphone) and perform keyword detection and/or automated speech recognition tasks using the inputs);
Srinivasa does not specifically disclose in response to detecting at least a portion of the keyword phrase, monitoring sensor data generated by the user attention sensor for an indication of user attention directed to the device.
However, Sreepathihalli, in the same field of endeavor, discloses:
a user attention sensor capable of detecting user attention directed to the device (Sreepathihalli, para 0022, the electronic device may determine whether the user's attention is on the electronic device); and
in response to detecting at least a portion of the keyword phrase, monitoring sensor data generated by the user attention sensor for an indication of user attention directed to the device (Sreepathihalli, para 0020-0022, the electronic device may differentiate when a user is present versus when an animal, such as a pet, runs in front of the depth sensor (e.g., as opposed to a motion sensor or a depth sensor that monitors a single zone). Compared to a camera, the resolution of the depth sensor data (e.g., 4×4, 8×8, 16×16, or 24×24) is significantly lower, such that most physical features of a user may not be identifiable in the depth sensor data… the electronic device may determine whether the user's attention is on the electronic device).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Sreepathihalli in the method of Srinivasa because this would enable monitoring a user’s attention using depth sensors which in turn enables, for instance, if the user is determined to be present (i.e., if it is determined that the user's attention is on the device), then the device causes a display to enter an operational mode, or otherwise, the device causes the display to enter a standby mode (Sreepathihalli, Abstract).
Srinivasa in view of Sreepathihalli do not specifically disclose in response to detecting a remainder of the keyword phrase and detecting the indication of user attention directed to the device, initiating a selected operation of the device.
However, Miller, in the same field of endeavor, discloses in response to detecting a remainder of the keyword phrase and detecting the indication of user attention directed to the device (Miller, para 0508, Device 900 detects input 2350m (e.g., a touch input, an air gesture, a mouse click, a gaze, and/or a speech input) directed at select account affordance 2338; [i.e., “gaze” as “detecting the indication of user attention”; and “speech input “ as “utterance specifying a keyword/keywords”]), initiating a selected operation of the device (Miller, para 0508-0509, Device 900 detects input 2350m (e.g., a touch input, an air gesture, a mouse click, a gaze, and/or a speech input) directed at select account affordance 2338. At FIG. 23N, in response to detecting input 2350m, device 900 displays account selection interface 2310n; [i.e., for this particular application, the device operation initiated is to display the account selection interface, which would not be displayed unless those two criteria (i.e., the detection of the gaze/attention and keywords/speech input) are satisfied]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Miller in the method of Srinivasa in view of Sreepathihalli because this would enable providing electronic devices with faster, more efficient methods and interfaces for displaying resource transfer user interfaces through the detection of a touch input, an air gesture, a mouse click, a gaze, and/or a speech input (Miller, para 0005 and 0509).
Regarding Claim 15. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 10, wherein the user attention sensor is a red-green-blue (RGB) camera (Srinivasa, col 9, ln 25 - 58, The electronic device 101 further includes one or more sensors 180 that can meter a physical quantity or detect an activation state of the electronic device 101 and convert metered or detected information into an electrical signal. The sensor(s) 180 can also include one or more buttons for touch input, one or more microphones, a gesture sensor, a gyroscope or gyro sensor, an air pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (such as an RGB sensor) … Any of these sensor(s) 180 can be located within the electronic device 101... The electronic device 101 can also be an augmented reality wearable device, such as eyeglasses, that include one or more cameras).
Regarding Claim 16. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 15, wherein the user attention sensor is an infrared camera (Srinivasa, col 9, ln 25 - 58, The electronic device 101 further includes one or more sensors 180 that can meter a physical quantity or detect an activation state of the electronic device 101 and convert metered or detected information into an electrical signal. The sensor(s) 180 can also include… an infrared (IR) sensor…… Any of these sensor(s) 180 can be located within the electronic device 101... The electronic device 101 can also be an augmented reality wearable device, such as eyeglasses, that include one or more cameras).
Regarding Claim 17. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 10, wherein the selected operation includes waking the device from a low power mode (Srinivasa, col 11, ln 11-20, As shown in FIG. 2, the system 200 includes the electronic device 101, which includes the processor 120. The processor 120 is operatively coupled to or otherwise configured to use one or more machine learning models, such as one or more keyword detection models 202. The keyword detection model 202 can be trained to recognize one or more keywords or phrases. For example, keywords can include wake words or phrases (e.g., “hey GOOGLE” or “hey BIXBY”), which are used to wake up a voice assistant of an electronic device from a low power or sleep mode).
Regarding Claim 18. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 10, wherein the selected operation includes responding to a further user utterance specifying a command (Srinivasa, col 11, ln 11 - 28, As shown in FIG. 2, the system 200 includes the electronic device 101, which includes the processor 120. The processor 120 is operatively coupled to or otherwise configured to use one or more machine learning models, such as one or more keyword detection models 202. The keyword detection model 202 can be trained to recognize one or more keywords or phrases…Additionally or alternatively, the keyword detection model 202 can detect predetermined keywords or phrases other than wake words or phrases, such as specific device commands that serve to both wake up the voice assistant and trigger further action, such as a preset command to “call Mom,” a preset command to play a song, a preset command to set a TV device to a particular setting or channel, a preset command to set an oven to a given temperature, and so on).
Regarding Claim 19. Srinivasa discloses acomputer program product comprising a computer readable storage medium having program instructions embodied therewith (Srinivasa, col 3, ln 3-5, computer readable program code and embodied in a computer readable medium… for implementation in a suitable computer readable program code), wherein the program instructions are executable by computer hardware of a device to cause the computer hardware to execute operations (Srinivasa, col 2, ln 48-50, instructions that when executed cause the at least one processor to generate instructions to perform an action based at least in part on the identified keyword) comprising:
detecting a first user utterance specifying a first keyword of a multi-keyword phrase from audio data (Srinivasa, col 1, ln 47-48, keyword detection model configured to predict a first likelihood that the audio data includes speech; Srinivasa, col 7, ln 50 - col 8, ln 17, FIG. 1 illustrates an example network configuration 100 including an electronic device… an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, or a sensor 180… the processor 120 may receive and process inputs (such as audio inputs or data received from an audio input device like a microphone) and perform keyword detection and/or automated speech recognition tasks using the inputs);
in response to the detecting the first user utterance (Srinivasa, col 23, ln 51-52, in response to the first likelihood exceeding a first threshold), monitoring the audio data for a second user utterance specifying a second keyword of the multi-keyword phrase (Srinivasa, col 1, ln 47-48, keyword detection model configured to predict a first likelihood that the audio data includes speech; Srinivasa, col 7, ln 50 - col 8, ln 17, FIG. 1 illustrates an example network configuration 100 including an electronic device… an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, or a sensor 180… the processor 120 may receive and process inputs (such as audio inputs or data received from an audio input device like a microphone) and perform keyword detection and/or automated speech recognition tasks using the inputs);
Srinivasa does not specifically disclose monitoring sensor data generated by a user attention sensor of the device for an indication of user attention directed to the device.
However, Sreepathihalli, in the same field of endeavor, discloses monitoring sensor data generated by a user attention sensor of the device for an indication of user attention directed to the device (Sreepathihalli, para 0020-0022, the electronic device may differentiate when a user is present versus when an animal, such as a pet, runs in front of the depth sensor (e.g., as opposed to a motion sensor or a depth sensor that monitors a single zone). Compared to a camera, the resolution of the depth sensor data (e.g., 4×4, 8×8, 16×16, or 24×24) is significantly lower, such that most physical features of a user may not be identifiable in the depth sensor data… the electronic device may determine whether the user's attention is on the electronic device).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Sreepathihalli in the method of Srinivasa because this would enable monitoring a user’s attention using depth sensors which in turn enables, for instance, if the user is determined to be present (i.e., if it is determined that the user's attention is on the device), then the device causes a display to enter an operational mode, or otherwise, the device causes the display to enter a standby mode (Sreepathihalli, Abstract).
Srinivasa in view of Sreepathihalli do not specifically disclose in response to detecting the second keyword and detecting the indication of user attention directed to the device, initiating a selected operation of the device.
However, Miller, in the same field of endeavor, discloses in response to detecting the second keyword and detecting the indication of user attention directed to the device (Miller, para 0508, Device 900 detects input 2350m (e.g., a touch input, an air gesture, a mouse click, a gaze, and/or a speech input) directed at select account affordance 2338; [i.e., “gaze” as “detecting the indication of user attention”; and “speech input “ as “utterance specifying a keyword/keywords”]), initiating a selected operation of the device (Miller, para 0508-0509, Device 900 detects input 2350m (e.g., a touch input, an air gesture, a mouse click, a gaze, and/or a speech input) directed at select account affordance 2338. At FIG. 23N, in response to detecting input 2350m, device 900 displays account selection interface 2310n; [i.e., for this particular application, the device operation initiated is to display the account selection interface, which would not be displayed unless those two criteria (i.e., the detection of the gaze/attention and keywords/speech input) are satisfied]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Miller in the method of Srinivasa in view of Sreepathihalli because this would enable providing electronic devices with faster, more efficient methods and interfaces for displaying resource transfer user interfaces through the detection of a touch input, an air gesture, a mouse click, a gaze, and/or a speech input (Miller, para 0005 and 0509).
Claims 2, 11 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Srinivasa in view of Sreepathihalli, further in view of Miller, and further in view of Finkelstein et al Pat No US 11100384 B2 (Finkelstein).
Regarding Claim 2. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 1, further comprising:
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose in response to detecting the first keyword, activating the user attention sensor of the device.
However, Finkelstein, in the same field of endeavor, discloses in response to detecting the first keyword, activating the user attention sensor of the device (Finkelstein, col 51, ln 27-41, the intelligent assistant system 20 may be activated upon detection of one or more keywords that are spoken by a user… An attention activator 32 in parser 40 may identify the keyword phrase “Hey computer” in the text. In response, the parser 40 may activate or modify other components and functionality of the intelligent assistant system 20; [i.e., “the intelligent assistant system 20” as “attention sensor” ] ).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Finkelstein in the method of Srinivasa in view of Sreepathihalli and Miller because this would ensure that a user's requests and intentions are fully captured, by allowing the device to wait for a period of time after the keyword is received and recognized, capturing additional user input after the keyword is detected, and verbal queries or commands following the keyword are captured for further processing (Finkelstein, col 51, ln 45-51, ).
Regarding Claim 11. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 10, wherein the hardware processor is capable of executing operations further comprising:
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose in response to detecting at least the portion of the keyword phrase, activating the user attention sensor of the device.
However, Finkelstein, in the same field of endeavor, discloses in response to detecting at least the portion of the keyword phrase, activating the user attention sensor of the device (Finkelstein, col 51, ln 27-41, the intelligent assistant system 20 may be activated upon detection of one or more keywords that are spoken by a user… An attention activator 32 in parser 40 may identify the keyword phrase “Hey computer” in the text. In response, the parser 40 may activate or modify other components and functionality of the intelligent assistant system 20; [i.e., “the intelligent assistant system 20” as “attention sensor” ]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Finkelstein in the method of Srinivasa in view of Sreepathihalli and Miller because this would ensure that a user's requests and intentions are fully captured, by allowing the device to wait for a period of time after the keyword is received and recognized, capturing additional user input after the keyword is detected, and verbal queries or commands following the keyword are captured for further processing (Finkelstein, col 51, ln 45-51.
Regarding Claim 20. Srinivasa in view of Sreepathihalli and Miller disclose the computer program product of claim 19, wherein the program instructions are executable by the computer hardware to execute operations further comprising:
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose in response to detecting the first keyword, activating the user attention sensor of the device.
However, Finkelstein, in the same field of endeavor, discloses in response to detecting the first keyword, activating the user attention sensor of the device (Finkelstein, col 51, ln 27-41, the intelligent assistant system 20 may be activated upon detection of one or more keywords that are spoken by a user… An attention activator 32 in parser 40 may identify the keyword phrase “Hey computer” in the text. In response, the parser 40 may activate or modify other components and functionality of the intelligent assistant system 20; [i.e., “the intelligent assistant system 20” as “attention sensor”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Finkelstein in the method of Srinivasa in view of Sreepathihalli and Miller because this would ensure that a user's requests and intentions are fully captured, by allowing the device to wait for a period of time after the keyword is received and recognized, capturing additional user input after the keyword is detected, and verbal queries or commands following the keyword are captured for further processing (Finkelstein, col 51, ln 45-51).
Claims 3-5 and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Srinivasa in view of Sreepathihalli, further in view of Miller, and further in view of Masahiro et al. Pat App No. JP 2015125632 A (Masahiro).
Regarding Claim 3. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 1.
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose wherein the detecting the attention of the user comprises detecting a body position of the user that matches a predetermined body position.
However, Masahiro, in the same field of endeavor, discloses wherein the detecting the attention of the user comprises detecting a body position of the user that matches a predetermined body position (Masahiro, 10th page, 8th para, the technique of the attention level estimation device described…This is the following technique. That is, a person who is viewing video content is photographed, and the skeleton position information of the person is obtained from the photographed image by measuring with motion capture. Then, the skeleton position information of the person is input in time series, and the amount of change per unit time at a predetermined skeleton position is measured as one of the body feature quantities of the person; [“predetermined skeleton position” as “predetermined body position”; “the attention level estimation” as “detecting the attention”; “measure” as “match”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Masahiro in the method of Srinivasa in view of Sreepathihalli and Miller because this would enable extracting attention keyword information that can be a key for retrieving video content or related information (Masahiro, 1st page, 2nd para).
Regarding Claim 4. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 1.
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose wherein the detecting the attention of the user comprises detecting at least one of a head orientation or a face of the user facing toward the device.
However, Masahiro, in the same field of endeavor, discloses wherein the detecting the attention of the user comprises detecting at least one of a head orientation or a face of the user facing toward the device (Masahiro, 10th page, 9th para, the technique of the interest level estimation apparatus described in JP2013-109537Acan be used. This is the following technique. That is, a person is photographed, and an image of the face area is automatically extracted. Based on the extracted face area image, face direction detection processing is performed. Depending on whether or not the face of the person is facing a display device (for example, a television receiver) that displays the content, the content can be viewed by the person. The presence or absence of is determined; [“the interest level estimation” as “the attention of the user”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Masahiro in the method of Srinivasa in view of Sreepathihalli and Miller because this would enable extracting attention keyword information that can be a key for retrieving video content or related information (Masahiro, 1st page, 2nd para).
Regarding Claim 5. Srinivasa in view of Sreepathihalli and Miller disclose the method of claim 1.
Srinivasa in view of Sreepathihalli and Miller do not specifically wherein the detecting the attention of the user comprises detecting that an eye gaze of the user is directed toward the device.
However, Masahiro, in the same field of endeavor, discloses wherein the detecting the attention of the user comprises detecting that an eye gaze of the user is directed toward the device (Masahiro, 10th page, 8th para, the technique of the attention level estimation device described…This is the following technique. That is, a person who is viewing video content is photographed, and the skeleton position information of the person is obtained from the photographed image by measuring with motion capture. Then, the skeleton position information of the person is input in time series, and the amount of change per unit time at a predetermined skeleton position is measured as one of the body feature quantities of the person. Then, in the camera images input in time series, the human eye area is detected, and the gaze fluctuation amount per unit time is measured as one of the body feature values based on the luminance of the left and right areas that divide the eye area; [i.e., “predetermined skeleton position” as “predetermined body position”; “the attention level estimation” as “detecting the attention”; “measure” as “match”; the “human eye area is detected, and the gaze fluctuation” as “detecting that an eye gaze of the user is directed toward the device”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Masahiro in the method of Srinivasa in view of Sreepathihalli and Miller because this would enable extracting attention keyword information that can be a key for retrieving video content or related information (Masahiro, 1st page, 2nd para).
Regarding Claim 12. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 10.
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose wherein the detecting the attention of the user comprises detecting a body position of the user that matches a predetermined body position.
However, Masahiro, in the same field of endeavor, discloses wherein the detecting the attention of the user comprises detecting a body position of the user that matches a predetermined body position (Masahiro, 10th page, 8th para, the technique of the attention level estimation device described…This is the following technique. That is, a person who is viewing video content is photographed, and the skeleton position information of the person is obtained from the photographed image by measuring with motion capture. Then, the skeleton position information of the person is input in time series, and the amount of change per unit time at a predetermined skeleton position is measured as one of the body feature quantities of the person; [“predetermined skeleton position” as “predetermined body position”; “the attention level estimation” as “detecting the attention”; “measure” as “match”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Masahiro in the method of Srinivasa in view of Sreepathihalli and Miller because this would enable extracting attention keyword information that can be a key for retrieving video content or related information (Masahiro, 1st page, 2nd para).
Regarding Claim 13. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 10.
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose wherein the detecting the attention of the user comprises detecting at least one of a head orientation or a face of the user facing toward the device.
However, Masahiro, in the same field of endeavor, discloses wherein the detecting the attention of the user comprises detecting at least one of a head orientation or a face of the user facing toward the device (Masahiro, 10th page, 9th para, the technique of the interest level estimation apparatus described in JP2013-109537Acan be used. This is the following technique. That is, a person is photographed, and an image of the face area is automatically extracted. Based on the extracted face area image, face direction detection processing is performed. Depending on whether or not the face of the person is facing a display device (for example, a television receiver) that displays the content, the content can be viewed by the person. The presence or absence of is determined; [“the interest level estimation” as “the attention of the user”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Masahiro in the method of Srinivasa in view of Sreepathihalli and Miller because this would enable extracting attention keyword information that can be a key for retrieving video content or related information (Masahiro, 1st page, 2nd para).
Regarding Claim 14. Srinivasa in view of Sreepathihalli and Miller disclose the device of claim 10.
Srinivasa in view of Sreepathihalli and Miller do not specifically disclose wherein the detecting the attention of the user comprises detecting that an eye gaze of the user is directed toward the device.
However, Masahiro, in the same field of endeavor, discloses wherein the detecting the attention of the user comprises detecting that an eye gaze of the user is directed toward the device (Masahiro, 10th page, 8th para, the technique of the attention level estimation device described…This is the following technique. That is, a person who is viewing video content is photographed, and the skeleton position information of the person is obtained from the photographed image by measuring with motion capture. Then, the skeleton position information of the person is input in time series, and the amount of change per unit time at a predetermined skeleton position is measured as one of the body feature quantities of the person. Then, in the camera images input in time series, the human eye area is detected, and the gaze fluctuation amount per unit time is measured as one of the body feature values based on the luminance of the left and right areas that divide the eye area; [i.e., “predetermined skeleton position” as “predetermined body position”; “the attention level estimation” as “detecting the attention”; “measure” as “match”; the “human eye area is detected, and the gaze fluctuation” as “detecting that an eye gaze of the user is directed toward the device”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Masahiro in the method of Srinivasa in view of Sreepathihalli and Miller because this would enable extracting attention keyword information that can be a key for retrieving video content or related information (Masahiro, 1st page, 2nd para).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MULUGETA T. DUGDA whose telephone number is (703)756-1106. The examiner can normally be reached Mon - Fri, 4:30am - 7:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MULUGETA TUJI DUGDA/Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
06/26/2026