Prosecution Insights
Last updated: August 06, 2026
Application No. 19/085,436

MULTIMODAL ROBOT-HUMAN INTERACTION VIA TEXT, VOICE, AND VIDEO FOR ROBOT CONTROLS

Non-Final OA §102§103§112§Other
Filed
Mar 20, 2025
Priority
Mar 27, 2024 — provisional 63/570,692
Examiner
RAMIREZ, ELLIS B
Art Unit
Tech Center
Assignee
Blue Hill Tech Inc.
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
1y 8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
179 granted / 220 resolved
+21.4% vs TC avg
Strong +18% interview lift
Without
With
+18.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
28 currently pending
Career history
243
Total Applications
across all art units

Statute-Specific Performance

§101
6.9%
-33.1% vs TC avg
§103
63.2%
+23.2% vs TC avg
§102
18.9%
-21.1% vs TC avg
§112
6.9%
-33.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 220 resolved cases

Office Action

§102 §103 §112 §Other
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims This is in response to applicant’s filing date of March 20, 2025. Claims 1-21 are currently pending. Information Disclosure Statement The information disclosure statement (IDS) submitted on June 09, 2025, is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Priority Regarding U.S. Provisional Patent Application Nos. 63/570,692, filed on 3/27/2024, Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120 is acknowledged. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: control module; gesture recognition module; facial recognition module; communication interface; set of gesture-based inputs; capturing, processing, interpreting, identifying, and enabling in claims 1-21. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections --35 U.S.C. § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-21are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Simon A. I. STENT (US-20190317594-A1)(“Stent”), provided by Applicant in the IDS filed on 6/09/2025. As per claim 1, Stent teaches a system for multi modal robot-human interaction, comprising: a robotic system equipped with a camera mounted on an automated service robot, wherein the camera is configured to capture visual data for a hand gesture recognition and a facial profiling (para [0020]. FIG. l is a conceptual diagram illustrating an unconstrained environment l 00 or context for implementing the system and method for detecting a human eye gaze and/or gesture movement by an intelligent machine according to an embodiment of the disclosure. A gaze and gesture detection system may be embodied in a computing machine, such as a robot 105 equipped with an image capturing device 110 .... "); a memory operatively connected to the robotic system to store a set of program instructions pertaining to the hand gesture recognition and the facial profiling (para [0054]. "The memory 408 may store information within the computing machine 405. The memory 408 may include volatile or non-volatile memory or other computer-readable medium, including without limitation a random access memory (RAM). a flash memory. a read-only memory (ROM), a programmable read-only memory (PROM). an erasable PROM (EPROM), registers. and so forth. The memory 408 may store program instructions, program data. executables. and other software and data useful for controlling operation of the computing machine 405 and configuring the computing machine 405 to perform functions for a user .... "); a computer processor coupled to the memory to execute the set of program instructions (para [0052], "In general. the computing machine 405 may include a processor 406, a memory 408, a network interface 410. a data store 412. and one or more input/output interfaces 414. The computing machine 405 may further include or be in communication with wired or wireless peripherals 416. or other external input/output devices, such as remote controls, communication devices. or the like, that might connect to the input/output interfaces 414."), wherein memory comprises: a control module to process the visual data to identify the hand gestures and the facial profiles of a user (para [0063]. " A controller 418 may serve to control motion or movement control of the computing machine 405. For example, the computing machine 405 may be implemented as a moving robot, in which the controller 418 may be responsible for monitoring and controlling all motion and movement functions of the computing machine 405. The controller 418 may be tasked with controlling motors, wheels, gears, arms or articulating limbs or digits associated with the motion or movement of the moving components of the computing machine. The controller may interface with the input/output interface 414 and the peripheral 416, where the peripheral 416 is a remote control capable of directing the movement of the computing machine."); a gesture recognition module to interpret the hand gestures to control a set of operations of the robotic system (para [0068], "Turning now to FIG. 5, a flowchart for a method 500 for locating and identifying an object using human gaze and gesture detection is shown according to an embodiment of the present disclosure. As detailed above, the systems described herein may use environment image data, human eye gaze information, and gesture information to locate and identify objects that are of interest to an individual. As shown in block 505, the system may receive a request to identify an object. For example, an individual may initiate a location and detection request by issuing a verbal command or a tactile command on a remote control or other interface in communication with the system "): and a facial recognition module to identify the user based on the facial profile and facilitate one or more personalized interactions (para [0068]); and a communication interface to enable interaction via one or more of: text. voice. and video (para [0069]:" system may, upon receiving a request, verbal or otherwise, parse the request using natural language processing, learned behavior, or the like to further refine the location and identification process. As detailed above, the system may be trained or programmed with a library of commands, objects and affordances that can be used refine saliency maps and object location and detection. The language of the request may provide the system with additional contextual information that may allow the system to narrow its search area, or environment. "). wherein the control module is configured to execute a set of predefined actions in response to recognized hand gestures and facial profiles ( para [0083]: "the system may be able to use contextual information and affordances extracted from the original request, at block 605, to further limit the environment. For example, if the request included identification of an area of the environment, such as a corner, or it included a contextual command to “open the window,” the system may limit the environment and the saliency map to only objects on walls. In this manner, the system can generate a focus saliency map with a more limited universe of potential objects or areas of interest ."). As per claim 2, Stent teaches the system as claimed in claim I, wherein the robotic system is selected from one or more of: a robotic arm; a humanoid robot: and a service robot with a set of mobility features (para [0022]: " computing machine, or robot 105, may include a 360-degree, monocular camera 111 as the image capturing device 110, or a component thereof. The use of a 360-degree, monocular camera 111 provides a substantially complete field-of view of the unconstrained environment 100. Therefore, regardless of a location of the individuals 112, 113, 114 or the multiple objects 115, the individuals 112, 113, 114 or the objects 115 may be captured by the camera 111. As such, the individuals 112, 113, 114 and the robot 105 may move within the environment 100 and the camera 111 may still capture the individuals 112, 113, 114 and multiple objects 115 within the unconstrained environment 100."). As per claim 3, Stent teaches the system as claimed in claim I. wherein the camera is mounted on at least one of: a wrist of the robotic system. a head of a humanoid robot. and an external stationary position within an operating environment (Figure 1, robot 105 with image capturing device 110, paras [0022]-[0023]:” robot 105 and the image capturing device 110 may include one or more infrared (“IR”) sensors 120 to scan or image the environment 100. Furthermore, the robot 105 and the image capturing device 110 may use the displacement of the IR signals to generate a depth image, which may be used by the system to estimate the depth of objects or individuals in the environment from the image capturing device 110.”). As per claim 4, Stent teaches the system as claimed in claim 3, wherein the camera mounted on the wrist provides a close-range view· of a set of objects held by the robotic system (para [0022]:” the individuals 112, 113, 114 and the robot 105 may move within the environment 100 and the camera 111 may still capture the individuals 112, 113, 114 and multiple objects 115 within the unconstrained environment 100.”). As per claim 5, Stent teaches the system as claimed in claim 3, wherein the camera mounted on head of the humanoid robot enables wide-angle human interaction and environment awareness (para [ 0022]). As per claim 6, Stent teaches the system as claimed in claim 1, ,wherein the robotic system's hand is configured to be compatible with human-like gloves while maintaining effective gesture recognition and object manipulation (para [0063]). As per claim 7, Stent teaches the system as claimed in claim 1, wherein the robotic system is configured to integrate hand gesture control and facial recognition to facilitate a c1uick order and confirmation process (paras [0063]. [0068]). As per claim 8, Stent teaches the system as claimed in claim 1, wherein the robotic system is trained through a three phase training process comprising: a data collection phase in which training data is captured using a human demonstration method and a teleoperation method (para [0048]:" association of objects with their properties and characteristics may be accomplished during training, semantic scene parsing or other machine learning stages as described herein. "); a model training phase where a set of AI models learn object interaction and behavior recognition (para [0069]:” system may, upon receiving a request, verbal or otherwise, parse the request using natural language processing, learned behavior, or the like to further refine the location and identification process. As detailed above, the system may be trained or programmed with a library of commands, objects and affordances that can be used refine saliency maps and object location and detection.”); and a deployment phase where the trained AI model autonomously executes tasks and adapts to environmental variables (para [0048] and Para. [0077] disclosing:” system may initiate a learning process by which the system may be trained, with interaction from the individual, by recording the object's image data in the knowledge base and an associated identity provided by the individual. The image data and associated identity may be recalled the next time the system is called upon to identify an object with similar image data.”). As per claim 9, Stent teaches the system as claimed in claim 1, wherein the robotic system is configured to receive a set of instructions from the user through one or more of: a set of voice commands; a set of text commands: and a set of gesture-based inputs to repeat learned actions (paras [0063], [0068], and [0069]). As per claim 10, Stent teaches the system as claimed in claim 1, wherein the robotic system is configured to navigate the environment autonomously to avoid one or more obstacles while executing a task. by utilizing one or more of a visual sensor, a positional sensor. and a depth sensor (See at least para [0047] and para [0070]:" depth image may be obtained using IR cameras, sensors or the like. Environment depth information may also be obtained using SLAM (e.g., for a moving robot system) or a deep neural network trained for depth estimation in unconstrained environments, such as indoor scenes. The depth information may be used to measure distances between the system, individuals, objects, and/or other definable contours of the environment, as detected by the system. "). As per claim 11, Stent teaches the system as claimed in claim 1, comprising a plurality of sensors integrated with the robotic system. wherein the sensors arc configured to provide an accurate hand gesture recognition and facial profiling (paras [0063], [0068]. [0069]). As per claim 12. Stent teaches the system as claimed in claim 1. wherein the gesture recognition module is configured to recognize a plurality of hand gestures, wherein each hand gesture is associated with a specific command for the robotic arm (paras [0063], [0068], [0069]). As per claim 13, Stent teaches the system as claimed in claim l, wherein the facial recognition module is configured to access a database of customer profiles (paras [0063 J. [0068], [0069]). As per claim 14, Stent teaches the system as claimed in claim l, wherein the control module is configured to integrate a hand gesture control and facial recognition to facilitate a quick order and confirmation process (para [0069]). As per claim 15, Stent teaches the system as claimed in claim 2,wherein the robotic arm is configured to mimic human arm movements (para [0022]). As per claim 16, Stent teaches the system as claimed in claim l, wherein the robotic system is trained to differentiate between a set of object categories based on a plurality of parameters comprising a location, appearance, and characteristics learned during the human demonstration (para [0047]). As per claim 17, Stent teaches the system as claimed in claim 1,wherein the automated service robot is configured to navigate the environment autonomously to avoid one or more obstacles while executing a task, by utilizing one or more of a visual sensor. a positional sensor. and a depth sensor (para. [0047]). As per claim 18. Stent teaches a method for multimodal robot-human interaction. comprising: capturing. by a computer processor. visual data using a camera mounted on a robotic system (para [0020]); processing, by the computer processor, the visual data to identify hand gestures and facial profiles (para [0063]); interpreting. by the computer processor. hand gestures to control a set of operations of the robotic system (para [0068]); identifying, by the computer processor, the user based on the facial profiles to facilitate personalized interactions (para [0068]); executing, by the computer processor, a set of predefined actions in response to recognized hand gestures and facial profiles (para [0069]); and enabling, by the computer processor, interaction via one or more of text, voice, and video through a communication interface (para [0083]). As per claim 19, Stent teaches the method as claimed in claim 18, further comprising accessing, by the computer processor, a database of the customer profiles (paras [0063], [0068], [0069]). As per claim 20, Stent teaches the method as claimed in claim 18, wherein the step of interpreting hand gestures comprising: recognizing, by the computer processor, a plurality of gestures, each associated with a specific command for the robotic system (paras [0063]. [0068], [0069]). As per claim 21, Stent teaches the method as claimed in claim 18, further comprising integrating, by the computer processor. a hand gesture control. and a facial recognition to facilitate a quick order and confirmation process (paras [0063]. [0068], [0069]). Claim Rejections --35 U.S.C. § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-21 are rejected under 35 U.S.C. 103 as being unpatentable over Hausman et al (US-20230311335-A1)(“Hausman”) and Florencio et al (US-20190248015-A1)(“Florencio”). Hausman discloses a robot system that can be commanded by a user to perform certain tasks for the user like cleaning, getting, and the like. To this end, Hausman discloses an interface where the user can command a robot to perform desired task through the use of a large language model that can reduced an utterance by the user to one or a series of commands in the form of task. Hausman discloses the invention substantially as claimed with following exception: Hausman fails to disclose an interface that can receive hand gestures and facial profiles of the user as a form of input. Florencio in the same field of endeavor discloses a system for human-robot interaction that includes a camera to acquire “facial expressions of the human 104” (Para. [0033]) and “air gestures, head and eye tracking, voice and speech, vision, touch, gestures”(Para. [0069]). A user in Florencio can used the enumerated forms or types of input to cause a robot to perform the task desired by the user. Florencio is considered to be analogous to the claimed invention because it is in the same field of systems which causes a robot to perform a task as commanded by a user. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hausman further in view of Florencio to allow for receiving a command from a user using verbal and non-verbal communications such as gestures, voice, and text. Motivation to do so would allow for reducing negative effects of service waiting time in a human-machine interaction by providing a more natural interaction during the service waiting time and increasing the likelihood that the robot will complete the task successfully with respect to a human that subsequently enters the environment of the robot (Florencio at Paras. [0005]-[0006]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Huang; Xiao-Qing William et al. (US-12397416-B2) Robot control method, storage medium and electronic device; Wang; Chao et al. (US-20250196363-A1) LLM driven multimodal human-robot interaction planning; Zass; Ron et al. (US-20250010458-A1) PERSONALIZING ROBOTIC INTERACTIONS; MARCHI; Eric et al. (US-20230368812-A1) DETERMINING WHETHER SPEECH INPUT IS INTENDED FOR A DIGITAL ASSISTANT; Hausman; Karol et al. (US-20230311335-A1) NATURAL LANGUAGE CONTROL OF A ROBOT; Wang; Jenny Z. (US-20230245651-A1) ENABLING USER-CENTERED AND CONTEXTUALLY RELEVANT INTERACTION; Steinberg; Erez et al. (US-20200142495-A1) GESTURE RECOGNITION CONTROL DEVICE; AL-GABRI; MOHAMMED MAHDI AHMED et al. (US-20190279529-A1) PORTABLE ROBOT FOR TWO-WAY COMMUNICATION WITH THE HEARING-IMPAIRED; MONCEAUX; Jérôme et al. (US-20170148434-A1) METHOD OF PERFORMING MULTI-MODAL DIALOGUE BETWEEN A HUMANOID ROBOT AND USER, COMPUTER PROGRAM PRODUCT AND HUMANOID ROBOT FOR IMPLEMENTING SAID METHOD; BRENCKLE; Wayne et al. (US-20140368423-A1) METHOD AND SYSTEM FOR LOW POWER GESTURE RECOGNITION FOR WAKING UP MOBILE DEVICES; Buehler; Christopher J. et al. (US-20130345870-A1) VISION-GUIDED ROBOTS AND METHODS OF TRAINING THEM. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ELLIS B. RAMIREZ whose telephone number is (571)272-8920. The examiner can normally be reached 7:30 am to 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ramon Mercado can be reached at 571-270-5744. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ELLIS B. RAMIREZ/ Examiner, Art Unit 3658
Read full office action

Prosecution Timeline

Mar 20, 2025
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699364
System and Method for Controlling an Entity
3y 11m to grant Granted Aug 04, 2026
Patent 12700489
CLOUD COACHING ARTIFICIAL INTELLIGENCE
3y 11m to grant Granted Aug 04, 2026
Patent 12690929
COMPENSATION OF GRAVITY-RELATED DISPLACEMENTS OF MEDICAL CARRIER STRUCTURES
5y 4m to grant Granted Jul 28, 2026
Patent 12691584
ROBOT SYSTEM AND ROBOT CELL
4y 2m to grant Granted Jul 28, 2026
Patent 12690931
METHOD FOR DETECTING, BASED ON THE MEASUREMENT OR DETECTION OF VELOCITIES, OPERATING ANOMALIES OF AN UNCONSTRAINED MASTER DEVICE OF A MASTER-SLAVE ROBOTIC SYSTEM FOR MEDICAL OR SURGICAL TELEOPERATION AND RELATED ROBOTIC SYSTEM
2y 11m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+18.0%)
3y 0m (~1y 8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 220 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month