DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In response to the office action from 5/7/2026, the applicant has submitted an amendment, filed 6/23/2026, amending claims 2, 5-8, 10-12, while arguing to traverse the prior art and 101 rejections. Applicant’s arguments have been fully considered but the rejections are maintained for the reasons explained in the response to arguments and further in view of Klein et al.
Response to Arguments
Following a broad overview of the latest amendments and the examiner interview of 6/10/2026, the 35 U.S.C. 101 is discussed.
Page 7 paragraph 2 it is recited: “the human mind is not trained to receive natural language input with a second sensor in response to detecting a particular context from a first sensor…”
Respectfully a human can receive a spoken content by his ear (a second sensor) following observation of an event by his eyes (a first sensor).
On page 7 the last paragraph last 2 lines it is recited: “For instance, amended claim 2 enhances the ability of a digital assistant to determine a specific task (e.g, by receiving input form one sensor when detecting input from different sensor”.
Respectfully how does the “first sensor” run faster and/or require less memory and/or engage in a function not performed by another sensor before upon detecting input from the “second sensor”?
Page 9 the last paragraph provides arguments directed at the latest amendments.
Please visit the new office action for further details.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 2 (11-12) are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
The claim amendment as drafted requires a specific “sensor” (“second sensor”) to be assigned for reception of “a natural language input”, while a specific but different “sensor” (“first sensor”) to be associated with “reference” to the “perform a task” limitation.
The only written description support for this claim are the original claims 2 (9) which merely mentioned “one or more sensors” in each case and thus were not specific to assignment of one “second” “sensor” to the “natural language input” and another different “first” “sensor” associated with the “reference” to the “perform a task” limitation.
Regarding claims 3-10 as they depend on claim 2 and as they do not obviate the problems noted in claim 2, they are thus rejected under similar rationale.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 2, 4-11 stand rejected:
The independent claims 2, 11 and 12, correspond to “non-transitory computer-readable storage medium”, “method” and “electronic device” respectively and are rejected under 35 U.S.C. 101 because they are directed to an abstract idea without significantly more.
The claims are basically about interaction between a user and a multi-mode “electronic device”, by virtue of what is defined as “a context of the electronic device” (Sp. ¶ 0249 S. before last: “In some examples, a context associated with a user includes an electronic device gesture. Exemplary electronic device gestures include a user pointing towards an object with the electronic device, lifting the electronic device to their ear”). The “electronic device” is also always supposed to “receive a natural language input” following e.g., the “gesture” to perform a “task”, so it is always by default supposed to have the voice mode, and the “context” cannot really be interpreted to be “voice”, since it allows reception of the “natural language input” only after “detect[ion]” of the “context”; i.e., the reception of “context” acts like a form of enabling access of the “electronic device”. And afterwards in response to the “natural language input” the device “perform[s] a task” (Sp. Par. 0302 S. before last: “e.g. audibly asking the user to keep turning a particular direction”).
These limitations, under their broadest reasonable interpretation, cover performance of the limitations in the mind but for the recitation of generic computer components. That is, other than reciting “by one or more processors” (Claims 2, and 12), nothing in the claims precludes their limitations from practically being performed in the mind. For example, it is very common that one person providing a gesture such as nodding his head or providing a hand gesture before verbally uttering something that he needs done as a task; e.g., a person providing direction to a destination to another person. If a claim limitation or limitations, under their broadest reasonable interpretation, cover performance of the limitation in the mind but for the recitation of generic computer components, then it falls withing the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claims recite only two additional elements , namely “one or more processors” and the “sensors” to “detect a context …”, “receive a natural language input”, “perform a task …”, “provide an output of the task”. The “one or more processors” are thus recited at a high-level of generality (i.e., as a generic processor performing all the above limitations) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea; see also ¶ 0245 S2: “In some examples, input processor 810 is implemented as part of an input/output component of a user device, server, or a combination of the two (e.g., I/O Subsystem 206, I/O interface 430, and I/O processing module 728)”. As a result, the claims are directed to an abstract idea.
Likewise, regarding the “sensors”, the claim limitations do not impose any limitations on their practice and they are introduced as a high level of generality; see also Sp. ¶ 0246 lines 2-5: “exemplary sensors include” “cameras” “microphones”. Furthermore these “sensors” are used for pre-solution data activity.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of the “one or more processors” to do “detect a context …”, “receive a natural language input”, “perform a task …”, “provide an output of the task” amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are thus not patent eligible.
Regarding claim 4, in the example above, the ability of making gestures is certainly an attribute of the person providing direction, e.g., towards a location and/or a device to perform an action.
Regarding claim 5, in the example above, the ability of making gestures is constitutes a context associated with the person providing direction. The sensors here are well known additional elements; just as discussed in the parent claim, the claim limitations do not impose any limitations on their practice and they are introduced as a high level of generality; see also Sp. ¶ 0246 lines 2-5: “exemplary sensors include” “cameras” “microphones”.
Regarding claim 6, in the example above, if the direction given by the provider of the direction turns out to be valid, the recipient could memorize it for future reference, and otherwise he will discard it.
Regarding claim 7, in the example above, the gestures by the person providing direction are received by the eye (a sensor) of the recipient.
Regarding claim 8, in the example above, the gestures provided are independent of the verbal part of providing direction and they are not associated with any verbal natural language input.
Regarding claim 9, in the example above, the verbal part of providing directions are received by the recipient’s ear (one or more sensors of the recipient tailored to hearing). The sensors here are well known additional elements; just as discussed in the parent claim, the claim limitations do not impose any limitations on their practice and they are introduced as a high level of generality; see also Sp. ¶ 0246 lines 2-5: “exemplary sensors include” “cameras” “microphones”.
Regarding claim 10, in the example above, the recipient of the direction could record the direction by his smart phone (an application of his). The sensors here are well known additional elements; just as discussed in the parent claim, the claim limitations do not impose any limitations on their practice and they are introduced as a high level of generality; see also Sp. ¶ 0246 lines 2-5: “exemplary sensors include” “cameras” “microphones”.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 2-12 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Klein et al. (US 2011/0313768).
Regarding claim 2, Klein et al. do teach a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device (¶ 0079: “Computing system” (an electronic device) “220 comprises a computer 241, which typically includes a variety of computer readable media” (comprising a non-transitory computer-readable medium) “contains data and/or program” (storing one or more programs) “modules that are immediately accessible to and/or presently being operated on by processing unit 259”),
cause the electronic device (as shown in Fig. 2 the “computing system” (the electronic device) couples to or includes the “Capture device” which comprises of “camera” and “microphone” (one or more sensors))
to:
detect a context of the electronic device with a first sensor of the electronic device (¶ 0088 lines 5+: “in step 412, the capture device 20” (at the electronic device) “captures” (detect) “a user movement” “as a defined command gesture” (a context by e.g., the “camera” (a first sensor) because according to ¶ 0034 S1: “camera” “monitors” “one or more user” “movement” “gestures”);
in response to detecting the context of the electronic device with the first sensor of the electronic device (¶ 0088 lines 8+: “Having recognized the gesture” (in response to the detecting of the context by the “camera” (the first sensor))):
receive a natural language input with a second sensor of the electronic device, wherein the second sensor is a different type of sensor from the first sensor (¶ 0089 S1: “In step 420, the microphone 30” (a second different sensor used to) “captures speech input” (receive a natural language input));
in accordance with a determination that the natural language input includes a reference to an input of the first sensor of the electronic device (¶ 0089 S1: “In step 420, the microphone 30” “captures speech input” (receive the natural language input, and this follows and is in response to capturing of the “gesture” (which was in reference to an input to the “camera” (the first sensor))):
perform a task associated with the natural language input based on the input of the first sensor of the electronic device (¶ 0094 last S: “In step 514, the system performs the action” (perform a task) “associated with the combination of recognized gesture” (in reference to input to “camera” (first sensor of the electronic device)) “and speech command”);
and provide an output of the task (¶ 0089 last S: “In step 424, the system performs the action” (providing an output) “associated with the recognized speech command” (corresponding to the task), e.g., ¶ 0104 lines 6+: “the user says” “PLAY” (in response to a natural language input) “the system changes state to a movie playback” (a movie is played as output))).
Regarding claim 3, Klein et al. do teach the non-transitory computer-readable storage medium of claim 2, wherein the one or more programs further comprise instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:
in accordance with the determination that the context corresponds to the predetermined type of event (¶ 0088 lines 7+: “movement” is a “defined” (predetermined) “command gesture” (event); see also another example ¶0104 P. 10 lines 6-7: “predefined gesture” (predetermined type of event)):
initiate a session of a digital assistant (¶ 0106 lines 5+ referring to Fig. 9C : “Having recognized the gesture, the system selects a voice library” (initiate a session with the “system” (electronic device which also functions as a digital assistant), where Fig. 9C according to ¶ 0024 describes “a process for user interaction” (a session between the user) “with a computing system” (with the electronic device))); and
wherein the natural language input is received in accordance with initiating the session of the digital assistant (¶ 0107 S1: “microphone” “captures speech input” (a natural language input is received) “Using the voice library that has been loaded” (in response to the initiating)).
Regarding claim 4, Klein et al. do teach the non-transitory computer-readable storage medium of claim 2, wherein the context of the electronic device includes at least one of a location of the electronic device, a setting of the electronic device, and an attribute of the electronic device (¶ 0051 last S: “the computing environment 12 may use the gesture” (the context is an attribute of the “computing” system” (electronic device)) “recognizer engine 54”).
Regarding claim 5, Klein et al. do teach the non-transitory computer-readable storage medium of claim 2, wherein detecting the context of the electronic device with the first sensor of the electronic device further comprises detecting a user gesture with the first sensor of the electronic device (¶ 0088 lines 5+: “in step 412, the capture device 20” (at the electronic device) “captures” (detect) “a user movement” “as a defined command gesture” (the context being a gesture by e.g., the “camera” (the first sensor) because according to ¶ 0034 S1: “camera” “monitors” “one or more user” “movement” “gestures”).
Regarding claim 6, Klein et al. do teach the non-transitory computer-readable storage medium of claim 2, wherein detecting the context of the electronic device with the first sensor of the electronic device further comprises:
in accordance with a determination that the context of the electronic device corresponds to a predetermined type of event, storing data from the first sensor to determine the task associated with the natural language input (¶ 0088 lines 8+: “Having recognized the gesture” (if the “gesture” (context) is a predetermined one and therefore a stored one associated with the “camera” (first sensor)) “the system selects a voice library (such as voice library 70, 72 . . . 76 shown in FIG. 2) having a limited set of voice commands that correspond to the gesture in step 416. The voice commands” “corresponding to the recognized gesture are then loaded” “into the voice recognizer” “engine 56 in step 418”); and
in accordance with a determination that the context of the electronic device does not correspond to the predetermined type of event, discarding the data from the first sensor (¶ 0092 last S.: “If the gesture” (if the context) “is not recognized” (does not correspond to a predetermined event) “it is not reported” (it is discarded from the “gesture recognizer” (i.e., the “camera” (the first sensor))) “to the application in step 480”).
Regarding claim 7, Klein et al. do teach the non-transitory computer-readable storage medium of claim 2, wherein the input from the first sensor includes at least one of an image, a RFID signal, and a location (¶ 0034 S1: “camera” (the first sensor) “monitors” (receives as input) “one or more user” “movement” “gestures” (images)).
Regarding claim 8, Klein et al. do teach the non-transitory computer-readable storage medium of claim 2, wherein the input from the first sensor includes data not determined from the natural language input (¶ 0034 S1: “camera” (the first sensor) “monitors” (receives as input) “one or more user” “movement” “gestures” (data not determined from the natural language input)).
Regarding claim 9, Klein et al. do teach the non-transitory computer-readable storage medium of claim 2, wherein the one or more programs further comprise instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:
in response to receiving the natural language input: engaging the first sensor of the electronic device; and receiving the input from the first sensor (¶ 0084 lines 15-17: “a user's speech can be recognized” (in response to a natural language input) “as a voice command to take action with regard to the specific item being pointed to” (engaging the first sensor to receive an input)).
Regarding claim 10, Klein et al. the non-transitory computer-readable storage medium of claim 2, wherein the input from the first sensor is received by an application of the electronic device (¶ 0050 lines 11+: “Application 52” (an application of the electronic device) “provides the tracking information, visual image data” (receives input from the first sensor) “to gesture recognizer engine 54 and the audio data to voice recognizer engine 56”).
Regarding claim 11, Klein et al. do teach a method (Title, Abstract),
Comprising:
At an electronic device (as shown in Fig. 2 the “computing system” (a electronic device) couples to or includes the “Capture device” which comprises of “camera” and “microphone” (one or more sensors)):
detecting a context of the electronic device with a first sensor of the electronic device (¶ 0088 lines 5+: “in step 412, the capture device 20” (at the electronic device) “captures” (detect) “a user movement” “as a defined command gesture” (a context by e.g., the “camera” (a first sensor) because according to ¶ 0034 S1: “camera” “monitors” “one or more user” “movement” “gestures”);
in response to detecting the context of the electronic device with the first sensor of the electronic device (¶ 0088 lines 8+: “Having recognized the gesture” (in response to the detecting of the context by the “camera” (the first sensor))):
receiving a natural language input with a second sensor of the electronic device wherein the second sensor is a different type of sensor from the first sensor (¶ 0089 S1: “In step 420, the microphone 30” (using a second different sensor) “captures speech input” (receive a natural language input));
in accordance with a determination that the natural language input includes a reference to an input of the first sensor of the electronic device performing a task associated with the natural language input based on the input of the first sensor of the electronic device (¶ 0094 last S: “In step 514, the system performs the action” (perform a task) “associated with the combination of recognized gesture” (in reference to input to “camera” (first sensor of the electronic device)) “and speech command”);
and providing an output of the task (¶ 0089 last S: “In step 424, the system performs the action” (providing an output) “associated with the recognized speech command” (corresponding to the task), e.g., ¶ 0104 lines 6+: “the user says” “PLAY” (in response to a natural language input) “the system changes state to a movie playback” (a movie is played as output))).
Regarding claim 12, Klein et al. do teach an electronic device, comprising: one or more processors; memory; and one or more programs stored in the memory (¶ 0079: “Computing system” (an electronic device) “220 comprises a computer 241, which typically includes a variety of computer readable media” (comprising a memory) “contains data and/or program” (and one or more programs stored in the memory) “modules that are immediately accessible to and/or presently being operated on by processing unit 259” (and one or more processors)),
the one or more programs including instructions for:
detecting a context of the electronic device with a first sensor of the electronic device (¶ 0088 lines 5+: “in step 412, the capture device 20” (at the electronic device) “captures” (detect) “a user movement” “as a defined command gesture” (a context by e.g., the “camera” (one or more sensors) because according to ¶ 0034 S1: “camera” “monitors” “one or more user” “movement” “gestures”);
in response to detecting the context of the electronic device with the first sensor of the electronic device (¶ 0088 lines 8+: “Having recognized the gesture” (in response to the detecting of the context by the “camera” (the first sensor))):
receiving a natural language input with a second sensor of the electronic device, wherein the second sensor is a different type of sensor from the first sensor (¶ 0089 S1: “In step 420, the microphone 30” (using a second different sensor )“captures speech input” (receive a natural language input));
in accordance with a determination that the natural language input includes a reference to an input of the first sensor of the electronic device performing a task associated with the natural language input based on the input of the first sensor of the electronic device (¶ 0094 last S: “In step 514, the system performs the action” (perform a task) “associated with the combination of recognized gesture” (in reference to input to “camera” (first sensor of the electronic device)) “and speech command”);
and providing an output of the task (¶ 0089 last S: “In step 424, the system performs the action” (providing an output) “associated with the recognized speech command” (corresponding to the task), e.g., ¶ 0104 lines 6+: “the user says” “PLAY” (in response to a natural language input) “the system changes state to a movie playback” (a movie is played as output))).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARZAD KAZEMINEZHAD whose telephone number is (571)270-5860. The examiner can normally be reached 10:30 am to 11:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Farzad Kazeminezhad/
Art Unit 2653
July 25th 2026.