DETAILED ACTION
This communication is a Non-Final Office Action on the Merits. Claims 1-20 as originally filed are currently pending and have been considered as follows:
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 16 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 16 recites “wherein controlling the robotic counterpart device with the trained neural network model comprises,” but claim 15 does not positively recite a step of controlling the robotic counterpart device. Claim 15 recites training a neural network model that is “configured to control” the robotic counterpart device. It is unclear whether claim 16 adds actual controlling steps to the method of claim 15 or merely describes the intended use or capability of the trained neural network model. Therefore, Claim 16 is rejected under 35 U.S.C. 112(b).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-5, 8-11, 13, 15, 16, 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang (NPL Title: Dexcap: Scalable and portable mocap data collection system for dexterous manipulation, Year: 2024) in view of Valentin (US Pub. No. 20210004979).
As per Claim 1, Wang discloses of scalable and portable mocap data collection system for dexterous manipulation, comprising:
a hand element configured to receive a hand of a user; (as per “DEXCAP (Fig. 1) is a portable hand mocap system that tracks the 6-DoF poses of the wrist and the finger motions in real-time (60Hz). The system includes a mocap glove to track finger joints, a camera mounted on top of each glove to track the 6-DoF poses of the wrists with SLAM, and an RGB-D LiDAR camera on the chest to observe the 3D environments.” in P2, INTRODUCTION, as per “In our system, finger motions are tracked using Rokoko motion capture gloves as illustrated in Figure 2. Each glove’s fingertip is embedded with a tiny magnetic sensor, while a signal receiver hub is placed on the glove’s dorsal side.” in P3, HARDWARE SYSTEM: DEXCAP)
a plurality of finger elements extending from the hand element; (as per “In our system, finger motions are tracked using Rokoko motion capture gloves as illustrated in Figure 2. Each glove’s fingertip is embedded with a tiny magnetic sensor, while a signal receiver hub is placed on the glove’s dorsal side.” in P3, HARDWARE SYSTEM: DEXCAP, as per “The data collection encompasses four data types, recorded at 60 frames per second: (1) the 6-DoF pose of the chest mounted LiDAR camera, as tracked by the top T265 camera; (2) the 6-DoF wrist poses, as captured by the two lower T265 cameras attached to the gloves; (3) the positions of finger joints within each glove’s reference frame, detected by the motion capture gloves; and (4) RGB-D image frames from the LiDAR camera.” in P15, APPENDIX A)
a mobile device mount coupled to the hand element configured to secure a mobile device to the wearable data collection device in a backward-facing orientation, (as per “The T265 cameras, initially in a known pose for calibration, are relocated to hand mounts during data collection to monitor palm positions, ensuring consistency through a click-in design. Finger motions are captured by Rokoko gloves, accurately tracking the finger joint positions.” in P3, RELATED WORK, as per “Then, we take off the tracking cameras from the rack and insert them into the camera slot attached to each glove. In this way, we can easily transform the hand pose tracking results into the observation frame of the chest camera with the constant initial transformation.” in P4, HARDWARE SYSTEM: DEXCAP, as per “The design features of the camera mounts on both the chest and gloves include a locking mechanism to prevent the cameras from accidentally slipping out. On the glove, the camera mount is positioned over the magnetic hub on its dorsal side, ensuring a firm attachment between the hub and the mount.” in P15, APPENDIX A)
wherein the backward-facing orientation positions a camera of the mobile device to face away from the plurality of finger elements; (as per “This system uses an Intel Realsense T265 camera, mounted on each glove’s dorsal side. It combines images from two fisheye cameras and IMU sensor signals to construct an environment map using the SLAM algorithm, enabling consistent tracking of the wrist’s 6-DoF pose.” in P4, HARDWARE SYSTEM: DEXCAP, as per “The glove mount follows the contour of the hump on the top of the Rokoko glove, and an opening is added to route the USB-C cable to the glove. The angle of the camera is set to 45 degrees facing upwards so that the camera view is less obstructed from the back of the hand.” in P15, APPENDIX A, as per “The T265 cameras, initially in a known pose for calibration, are relocated to hand mounts during data collection to monitor palm positions, ensuring consistency through a click-in design.” In P3, RELATED WORK)
a plurality of sensors mounted on the wearable data collection device configured to capture sensor data during a recording session; (as per “DEXCAP offers precise, occlusion-resistant tracking of wrist and finger motions based on SLAM and electromagnetic field together with 3D observations of the environment” in P1, ABSTRACT, as per “This system uses an Intel Realsense T265 camera, mounted on each glove’s dorsal side. It combines images from two fisheye cameras and IMU sensor signals to construct an environment map using the SLAM algorithm, enabling consistent tracking of the wrist’s 6-DoF pose.” in P4, HARDWARE SYSTEM: DEXCAP)
a processing circuit operatively coupled to the plurality of sensors configured to collect the sensor data; (as per “Central to the portability of DEXCAP is a compact mini-PC (Intel NUC 13 Pro), carried in a backpack, which serves as the primary computation unit for data recording.” in P4, HARDWARE SYSTEM: DEXCAP)
Wang fails to expressly disclose:
a mobile device;
transmit the sensor data to the mobile device.
Valentin discloses of depth from motion for augmented reality for handheld user devices, comprising:
a mobile device; (as per “FIG. 1 illustrates an example provisioning of an AR experience on a handheld user device 100 deployed in a real-world scene 102 using a depth-from-motion pipeline as described herein” in ¶11, as per “The handheld user device can be, for example, one of a compute-enabled cellular phone, a tablet computer, and a portable gaming device.” in ¶72)
transmit the sensor data to the mobile device. (as per “As a general operational overview, the monocular camera 214 operates to capture one or more sequences of real-world images 106 (FIG. 1) for inclusion in the camera feed 104, while the various sensors of the IMU 212 capture motion-related data representative of the pose, position, and movement of the handheld user device 100 for inclusion in a pose/position sensor feed 228.” in ¶15, as per “The sensor hub 210 operates to format, synchronize, and otherwise process the camera feed 104 and pose/position feed 228 and provide the resulting processed sensor streams for access by the application processor 202 (e.g., either through a direct input or via temporary storage in the memory 208, the mass storage device 218, or other storage component).” in ¶15)
In this way, Valentin operates to use a handheld user device, such as a smartphone or tablet, having a camera and IMU to capture real-world images, capture motion-related position data, and synchronize the resulting camera and position sensor streams (¶11, ¶15, ¶72). Like Wang, Valentin is concerned with IMU-based sensing and 6DoF tracking in a real-world environment.
It would have been obvious for one of ordinary skill in the art before the effective filing date to have modified the wearable data collection system of Wang with the handheld mobile-device sensing arrangement of Valentin to enable another standard means of using a portable mobile device camera and IMU to capture environmental image data and position data. Such modification also allows the system to use commonly available mobile-device components for camera-based and IMU-based tracking, thereby improving portability and reducing the need for specialized tracking hardware (¶1, ¶11, ¶15, ¶72).
As per Claim 2, the combination of Wang and Valentin teaches or suggests all limitations of Claim 1. Wang further discloses wherein the mobile device mount positions the mobile device such that the camera of the mobile device captures image data of an environment behind the wearable data collection device, and wherein the environment behind the wearable data collection device includes a wearer of the wearable data collection device. (as per “This system uses an Intel Realsense T265 camera, mounted on each glove’s dorsal side. It combines images from two fisheye cameras and IMU sensor signals to construct an environment map using the SLAM algorithm, enabling consistent tracking of the wrist’s 6-DoF pose.” in P4, HARDWARE SYSTEM: DEXCAP, as per “Upon initiating the program, the participant moves within the environment for several seconds, allowing the SLAM algorithm to build the map of the surroundings.” in P15, APPENDIX A, as per "The angle of the camera is set to 45 degrees facing upwards so that the camera view is less obstructed from the back of the hand." in P15, APPENDIX A)
As per Claim 3, the combination of Wang and Valentin teaches or suggests all limitations of Claim 1. Wang further discloses wherein the mobile device is configured to:
track a position and an orientation of the wearable data collection device in space during the recording session using at least one of an inertial measurement unit of the mobile device and the camera of the mobile device; (as per “To address these challenges, we develop a 6-DoF wrist tracking system based on the SLAM algorithm, as shown in Figure 2(c). This system uses an Intel Realsense T265 camera, mounted on each glove’s dorsal side. It combines images from two fisheye cameras and IMU sensor signals to construct an environment map using the SLAM algorithm, enabling consistent tracking of the wrist’s 6-DoF pose.” in P4, HARDWARE SYSTEM: DEXCAP, as per “This design has three key advantages: it is portable, allowing for wrist pose tracking without the need for hands to be visible in third-person camera frames; SLAM can autonomously correct position drift with the built map for long-time use; and the IMU sensor provides crucial wrist orientation information to train the robot policy in the subsequent pipeline.” in P4, HARDWARE SYSTEM: DEXCAP)
record the position and the orientation. (as per “The data collection encompasses four data types, recorded at 60 frames per second: (1) the 6-DoF pose of the chest mounted LiDAR camera, as tracked by the top T265 camera; (2) the 6-DoF wrist poses, as captured by the two lower T265 cameras attached to the gloves; (3) the positions of finger joints within each glove’s reference frame, detected by the motion capture gloves; and (4) RGB-D image frames from the LiDAR camera.” in P15, APPENDIX A, as per “This necessitates DEXCAP to estimate and record the 6-DoF pose trajectories of human hands during data collection.” in P3, HARDWARE SYSTEM: DEXCAP)
Wang fails to expressly disclose:
mobile device configured to track using an inertial measurement unit and camera
See Claim 1 for teachings of Valentin. Valentin further discloses:
mobile device configured to track using an inertial measurement unit and camera (as per “These portable devices support such AR capability using some form of six degree of freedom (6DoF) tracking capability using just typical sensors found inside these device, such as a color camera and an inertial measurement unit (IMU), leading to many developments in visual inertial odometry (VIO) and simultaneous localization and mapping (SLAM).” in ¶1, as per “A software-implemented depth-from-motion pipeline at the handheld user device 100 uses position/pose tracking data from an IMU (not shown in FIG. 1) to determine the current position/pose (6DoF) of the handheld user device 100 relative to the real-world scene 102, and from this 6DoF data and at least some of the captured real-world images 106 of the color camera feed 104, compute depth information 108 for the real-world scene in real time in the form of a sequence of dense depth maps 110.” in ¶11, as per ¶15)
In this way, Valentin operates to use a handheld user device, such as a smartphone or tablet, having a camera and IMU to capture real-world images, capture motion-related position data, and synchronize the resulting camera and position sensor streams (¶11, ¶15, ¶72). Like Wang, Valentin is concerned with IMU-based sensing and 6DoF tracking in a real-world environment.
It would have been obvious for one of ordinary skill in the art before the effective filing date to have modified the wearable data collection system of Wang with the handheld mobile-device sensing arrangement of Valentin to enable another standard means of using a portable mobile device camera and IMU to capture environmental image data and position data. Such modification also allows the system to use commonly available mobile-device components for camera-based and IMU-based tracking, thereby improving portability and reducing the need for specialized tracking hardware (¶1, ¶11, ¶15, ¶72).
As per Claim 4, the combination of Wang and Valentin teaches or suggests all limitations of Claim 1. Wang further discloses:
a connection interface configured to transmit the sensor data from the processing circuit to the mobile device, (as per “In the human-in-the-loop process, we employ the mini-PC to live stream data from all T265 tracking cameras. This tracking information is then transmitted to a Redis server configured on the local network.” in P18, APPENDIX A)
wherein the connection interface comprises at least one of a wired connection and a wireless connection. (as per “The glove mount follows the contour of the hump on the top of the Rokoko glove, and an opening is added to route the USB-C cable to the glove.” in P15, APPENDIX A)
As per Claim 5, the combination of Wang and Valentin teaches or suggests all limitations of Claim 4. Wang further discloses wherein the wired connection comprises a Universal Serial Bus (USB) connection. (as per “The glove mount follows the contour of the hump on the top of the Rokoko glove, and an opening is added to route the USB-C cable to the glove.” in P15, APPENDIX A)
Claim 8 is rejected using the same rationale, mutatis mutandis, applied to Claim(s) 1 & 3 above, respectively. Wang and Valentin teach or suggest receiving/securing the mobile device in the mobile device mount in the claimed backward-facing orientation, capturing sensor data during a recording session, capturing position and orientation data using the mobile device, and transmitting sensor data to the mobile device via a processing circuit, as set forth above with respect to claims 1 and 3.
Claim 9 is rejected using the same rationale, mutatis mutandis, applied to Claim 3 above, respectively.
Claim 10 is rejected using the same rationale, mutatis mutandis, applied to Claim 4 above, respectively.
Claim 11 is rejected using the same rationale, mutatis mutandis, applied to Claim 5 above, respectively.
As per Claim 13, the combination of Wang and Valentin teaches or suggests all limitations of Claim 8. Wang further discloses:
capturing environmental image data, wherein the environmental image data provides contextual information about an environment behind the wearable data collection device during the recording session; (as per “It combines images from two fisheye cameras and IMU sensor signals to construct an environment map using the SLAM algorithm, enabling consistent tracking of the wrist’s 6-DoF pose” in P4, HARDWARE SYSTEM: DEXCAP, as per “Capturing the data necessary for training robot policies requires not only the tracking of hand movement but also recording observations of the 3D environment as the policy input” in P4, HARDWARE SYSTEM: DEXCAP, as per “The angle of the camera is set to 45 degrees facing upwards so that the camera view is less obstructed from the back of the hand” in P15, APPENDIX A); and
associating the environmental image data with the sensor data and the position and orientation data. (as per “The data collection encompasses four data types, recorded at 60 frames per second: (1) the 6-DoF pose of the chest mounted LiDAR camera, as tracked by the top T265 camera; (2) the 6-DoF wrist poses, as captured by the two lower T265 cameras attached to the gloves; (3) the positions of finger joints within each glove’s reference frame, detected by the motion capture gloves; and (4) RGB-D image frames from the LiDAR camera. The initial pose of the top T265 camera establishes the world frame for all data, allowing for the integration of all streamed data—RGB-D point clouds, hand 6-DoF poses, and finger joint locations—into a unified world frame.” in P15, APPENDIX A)
Wang fails to expressly disclose:
capturing environmental image data using the camera of the mobile device,
See Claim 8 for teachings of Valentin. Valentin further discloses capturing environmental image data using the camera of the mobile device. (as per “the handheld user device 100 (e.g., a smartphone, tablet computer, personal digital assistant, portable video game device) employs a monocular camera sensor (not shown in FIG. 1) to capture a camera feed 104 composed of a sequence of captured real-world images 106 of the real-world scene 102” in ¶11, as per “the monocular camera 214 operates to capture one or more sequences of real-world images 106 (FIG. 1) for inclusion in the camera feed 104” in ¶15)
In this way, Valentin operates to use a handheld user device, such as a smartphone or tablet, having a camera and IMU to capture real-world images, capture motion-related position data, and synchronize the resulting camera and position sensor streams (¶11, ¶15, ¶72). Like Wang, Valentin is concerned with IMU-based sensing and 6DoF tracking in a real-world environment.
It would have been obvious for one of ordinary skill in the art before the effective filing date to have modified the wearable data collection system of Wang with the handheld mobile-device sensing arrangement of Valentin to enable another standard means of using a portable mobile device camera and IMU to capture environmental image data and position data. Such modification also allows the system to use commonly available mobile-device components for camera-based and IMU-based tracking, thereby improving portability and reducing the need for specialized tracking hardware (¶1, ¶11, ¶15, ¶72).
As per Claim 15, Wang discloses of scalable and portable mocap data collection system for dexterous manipulation, comprising:
receiving sensor data captured during a recording session by a plurality of sensors on a wearable data collection device, wherein: (as per “DEXCAP offers precise, occlusion-resistant tracking of wrist and finger motions based on SLAM and electromagnetic field together with 3D observations of the environment.” in P1, ABSTRACT, as per “The data collection encompasses four data types, recorded at 60 frames per second: (1) the 6-DoF pose of the chest-mounted LiDAR camera, as tracked by the top T265 camera; (2) the 6-DoF wrist poses, as captured by the two lower T265 cameras attached to the gloves; (3) the positions of finger joints within each glove’s reference frame, detected by the motion capture gloves; and (4) RGB-D image frames from the LiDAR camera.” in P15, APPENDIX A)
the wearable data collection device comprises a hand element configured to receive a hand of a user and a plurality of finger elements extending from the hand element; (as per “DEXCAP (Fig. 1) is a portable hand mocap system that tracks the 6-DoF poses of the wrist and the finger motions in real-time (60Hz). The system includes a mocap glove to track finger joints, a camera mounted on top of each glove to track the 6-DoF poses of the wrists with SLAM, and an RGB-D LiDAR camera on the chest to observe the 3D environments.” in P2, INTRODUCTION, as per “In our system, finger motions are tracked using Rokoko motion capture gloves as illustrated in Figure 2. Each glove’s fingertip is embedded with a tiny magnetic sensor, while a signal receiver hub is placed on the glove’s dorsal side.” in P3, HARDWARE SYSTEM: DEXCAP)
a mobile device is secured to the wearable data collection device in a backward-facing orientation, wherein the backward-facing orientation positions a camera of the mobile device to face away from the plurality of finger elements; (as per “The T265 cameras, initially in a known pose for calibration, are relocated to hand mounts during data collection to monitor palm positions, ensuring consistency through a click-in design.” in P3, RELATED WORK, as per “Then, we take off the tracking cameras from the rack and insert them into the camera slot attached to each glove.” in P4, HARDWARE SYSTEM: DEXCAP, as per “The design features of the camera mounts on both the chest and gloves include a locking mechanism to prevent the cameras from accidentally slipping out. On the glove, the camera mount is positioned over the magnetic hub on its dorsal side, ensuring a firm attachment between the hub and the mount.” in P15, APPENDIX A, as per “The angle of the camera is set to 45 degrees facing upwards so that the camera view is less obstructed from the back of the hand.” in P15, APPENDIX A)
receiving position and orientation data of the wearable data collection device captured by the mobile device during the recording session; (as per “This necessitates DEXCAP to estimate and record the 6-DoF pose trajectories of human hands during data collection.” in P3, HARDWARE SYSTEM: DEXCAP, as per “This system uses an Intel Realsense T265 camera, mounted on each glove’s dorsal side. It combines images from two fisheye cameras and IMU sensor signals to construct an environment map using the SLAM algorithm, enabling consistent tracking of the wrist’s 6-DoF pose.” in P4, HARDWARE SYSTEM: DEXCAP, as per “the 6-DoF wrist poses, as captured by the two lower T265 cameras attached to the gloves” in P15, APPENDIX A)
processing the sensor data and the position and orientation data to generate training data for a neural network; (as per “To tackle these challenges, we introduce DEXIL, a three-step framework to train dexterous robots using human hand motion capture data. The first step is to re-target the DEXCAP data into the action and observation spaces of the robot embodiment” in P5, LEARNING ALGORITHM: DEXIL, as per “The 6-DoF pose of the wrist pt = [Rt|Tt] and the finger joint positions Jt of the LEAP hands are then used as the robot’s proprioception state st = (pt, Jt).” in P5, DATA RE-TARGETING, as per “All of the point cloud observations are downsampled uniformly to 5000 points and stored together with robot proprioception states and actions into an hdf5 file.” in P16, APPENDIX A)
training the neural network using the training data to generate a trained neural network model, wherein the trained neural network model is configured to control a robotic counterpart device having a joint and sensor configuration that matches the wearable data collection device. (as per “Our goal is to use the human hand motion capture data recorded by DEXCAP to train dexterous robot policies.” in P5, LEARNING ALGORITHM: DEXIL, as per “Second step trains a point-cloud-based Diffusion Policy using the re-targeted data” in P5, LEARNING ALGORITHM: DEXIL, as per “More specifically, an policy model π, processes the point cloud observations ot and the robot’s current proprioception state st into an action trajectory (at, at+1, . . . , at+d)” in P6, POINT CLOUD-BASED DIFFUSION POLICY, as per “To validate the robot policy trained by the data from DEXCAP, we establish a bimanual dexterous robot setup. This setup comprises two Franka Emika robot arms, each equipped with a LEAP dexterous robotic hand (a four-fingered hand with 16 joints)” in P4, HARDWARE SYSTEM: DEXCAP, as per “Mirroring the human system, the robot system reuses the same chest cameras and mount.” in P4, HARDWARE SYSTEM: DEXCAP)
Wang fails to expressly disclose:
a mobile device;
the position and orientation data being captured by the mobile device.
Valentin discloses of depth from motion for augmented reality for handheld user devices, comprising:
a mobile device; (as per “FIG. 1 illustrates an example provisioning of an AR experience on a handheld user device 100 deployed in a real-world scene 102 using a depth-from-motion pipeline as described herein” in ¶11, as per “The handheld user device can be, for example, one of a compute-enabled cellular phone, a tablet computer, and a portable gaming device.” in ¶72)
the position and orientation data being captured by the mobile device. (as per “These portable devices support such AR capability using some form of six degree of freedom (6DoF) tracking capability using just typical sensors found inside these device, such as a color camera and an inertial measurement unit (IMU), leading to many developments in visual inertial odometry (VIO) and simultaneous localization and mapping (SLAM).” in ¶1, as per “A software-implemented depth-from-motion pipeline at the handheld user device 100 uses position/pose tracking data from an IMU (not shown in FIG. 1) to determine the current position/pose (6DoF) of the handheld user device 100 relative to the real-world scene 102, and from this 6DoF data and at least some of the captured real-world images 106 of the color camera feed 104, compute depth information 108 for the real-world scene in real time in the form of a sequence of dense depth maps 110.” in ¶11, as per ¶15)
In this way, Valentin operates to use a handheld user device, such as a smartphone or tablet, having a camera and IMU to capture real-world images, capture motion-related position data, and synchronize the resulting camera and position sensor streams (¶11, ¶15, ¶72). Like Wang, Valentin is concerned with IMU-based sensing and 6DoF tracking in a real-world environment.
It would have been obvious for one of ordinary skill in the art before the effective filing date to have modified the wearable data collection system of Wang with the handheld mobile-device sensing arrangement of Valentin to enable another standard means of using a portable mobile device camera and IMU to capture environmental image data and position data. Such modification also allows the system to use commonly available mobile-device components for camera-based and IMU-based tracking, thereby improving portability and reducing the need for specialized tracking hardware (¶1, ¶11, ¶15, ¶72).
As per Claim 16, the combination of Wang and Valentin teaches or suggests all limitations of Claim 15. Wang further discloses wherein controlling the robotic counterpart device with the trained neural network model comprises:
receiving real-time sensor data from multiple sensors on the robotic counterpart device; (as per “More specifically, an policy model π, processes the point cloud observations ot and the robot’s current proprioception state st into an action trajectory (at, at+1, . . . , at+d)” in P6, POINT CLOUD-BASED DIFFUSION POLICY, as per “To bridge the visual gap between human hands and the robot’s hand, we use forward kinematics to transform the links of the robot model with the proprioception state st and merge the point clouds of the transformed links into the observation ot.” in P6, POINT CLOUD-BASED DIFFUSION POLICY, as per “The RGB-D LiDAR camera, positioned on the central bar between the robot arms, connects to the workstation to capture observation data.” in P18, APPENDIX A, as per “Following each robot action, we calculate the distance between the robot’s current proprioception and the target pose.” in P16, APPENDIX A)
processing the real-time sensor data using the trained neural network model to determine control signals; (as per “More specifically, an policy model π, processes the point cloud observations ot and the robot’s current proprioception state st into an action trajectory (at, at+1, . . . , at+d)” in P6, POINT CLOUD-BASED DIFFUSION POLICY, as per “At the high level, the learned policy generates the goal position for the next step, which encompasses the 6-DoF pose of the end-effector for both robot arms and a 16-dimensional finger joint position for both hands.” in P16, APPENDIX A)
transmitting the control signals to the robotic counterpart device to control movement of the robotic counterpart device. (as per “At the low level, an Operational Space Controller (OSC), continuously interpolates the arm’s trajectory towards the high-level specified goal position and relays interpolated OSC actions to the robot for execution. Meanwhile, finger movements are directly managed by a joint impedance controller.” in P16, APPENDIX A, as per “Both the robot arms and the LEAP hands operate at a control frequency of 20Hz. We use end-effector position control for both robot arms and joint position control for both LEAP hands.” in P4, HARDWARE SYSTEM: DEXCAP)
Claim 18 is rejected using the same rationale, mutatis mutandis, applied to Claim 3 above, respectively.
As per Claim 19, the combination of Wang and Valentin teaches or suggests all limitations of Claim 15. Wang further discloses:
receiving environmental image data captured during the recording session, wherein the environmental image data provides contextual information about an environment behind the wearable data collection device; (as per “Capturing the data necessary for training robot policies requires not only the tracking of hand movement but also recording observations of the 3D environment as the policy input.” in P4, HARDWARE SYSTEM: DEXCAP, as per “It incorporates an Intel Realsense L515 RGB-D LiDAR camera, mounted on the top of the chest, to capture the observations during human data collection.” in P4, HARDWARE SYSTEM: DEXCAP, as per “This system uses an Intel Realsense T265 camera, mounted on each glove’s dorsal side. It combines images from two fisheye cameras and IMU sensor signals to construct an environment map using the SLAM algorithm, enabling consistent tracking of the wrist’s 6-DoF pose.” in P4, HARDWARE SYSTEM: DEXCAP, as per “The angle of the camera is set to 45 degrees facing upwards so that the camera view is less obstructed from the back of the hand.” in P15, APPENDIX A)
incorporating the environmental image data into the training data for the neural network. (as per “We convert the RGB-D images captured by the LiDAR camera in the DEXCAP data into point clouds using the camera parameters.” in P5, DATA RE-TARGETING, as per “Based on these findings, all RGB-D frames from the mocap data are processed into point clouds aligned with the robot’s space, and the task-irrelevant elements, such as the table surface points, are excluded. This refined point cloud data thus becomes the observation inputs ot fed into the robot policy π.” in P6, DATA RE-TARGETING, as per “With the transformed robot’s state st, action at and corresponding 3D point cloud observation ot, we formalize the robot policy learning process as a trajectory generation task.” in P6, POINT CLOUD-BASED DIFFUSION POLICY)
Wang fails to expressly disclose:
receiving environmental image data captured by the camera of the mobile device;
See Claim 15 for teachings of Valentin. Valentin further discloses:
receiving environmental image data captured by the camera of the mobile device; (as per “the handheld user device 100 (e.g., a smartphone, tablet computer, personal digital assistant, portable video game device) employs a monocular camera sensor (not shown in FIG. 1) to capture a camera feed 104 composed of a sequence of captured real-world images 106 of the real-world scene 102” in ¶11, as per “the monocular camera 214 operates to capture one or more sequences of real-world images 106 (FIG. 1) for inclusion in the camera feed 104” in ¶15)
In this way, Valentin operates to use a handheld user device, such as a smartphone or tablet, having a camera and IMU to capture real-world images, capture motion-related position data, and synchronize the resulting camera and position sensor streams (¶11, ¶15, ¶72). Like Wang, Valentin is concerned with IMU-based sensing and 6DoF tracking in a real-world environment.
It would have been obvious for one of ordinary skill in the art before the effective filing date to have modified the wearable data collection system of Wang with the handheld mobile-device sensing arrangement of Valentin to enable another standard means of using a portable mobile device camera and IMU to capture environmental image data and position data. Such modification also allows the system to use commonly available mobile-device components for camera-based and IMU-based tracking, thereby improving portability and reducing the need for specialized tracking hardware (¶1, ¶11, ¶15, ¶72).
As per Claim 20, the combination of Wang and Valentin teaches or suggests all limitations of Claim 15. Wang further discloses:
receiving additional sensor data and position and orientation data from multiple recording sessions from the wearable data collection device, wherein the multiple recording sessions comprise recordings of different tasks performed with the wearable data collection device; (as per “we evaluate DEXIL using six tasks of varying difficulty to assess its performance with DEXCAP data.” in P7, EXPERIMENT SETUPS, as per “We utilize three data types: (1) DEXCAP data capturing human hand motion within the robot’s operational space, (2) in-the-wild DEXCAP data from outside lab environments, and (3) human-in-the-loop correction data for adjusting robot actions or enabling teleoperation to correct errors, collected using a foot pedal.” in P7, EXPERIMENT SETUPS, as per “For data collection, we gathered 30 minutes of DEXCAP data across the first three tasks, resulting in 201, 129, and 82 demos respectively. An hour of in-the-wild DEXCAP data provided 96 demos for Packaging. Scissor Cutting and Tea Preparing tasks each received an hour of DEXCAP data, yielding 104 and 55 demos respectively.” in P7, EXPERIMENT SETUPS)
analyzing the additional sensor data and position and orientation data to identify one or more patterns; and (as per “With the transformed robot’s state st, action at and corresponding 3D point cloud observation ot, we formalize the robot policy learning process as a trajectory generation task.” in P6, POINT CLOUD-BASED DIFFUSION POLICY, as per “More specifically, an policy model π, processes the point cloud observations ot and the robot’s current proprioception state st into an action trajectory (at, at+1, . . . , at+d)” in P6, POINT CLOUD-BASED DIFFUSION POLICY, as per “During training, we also use data augmentation over the inputs by applying random 2D translations to the point clouds and motion trajectories within the robot’s operational space.” in P6, POINT CLOUD-BASED DIFFUSION POLICY)
refining the trained neural network model based on the one or more patterns to improve performance of the robotic counterpart device. (as per “DEXCAP also offers an optional human-in-the-loop correction mechanism to refine and further improve robot performance.” in P1, ABSTRACT, as per “The corrected actions and observations are stored in a new dataset D′. Training data is sampled with equal probability from D′ and the original dataset D to fine-tune the policy model, similar to IWR [46].” in P7, HUMAN-IN-THE-LOOP CORRECTION, as per “The last three columns of Table II showcase the effectiveness of using human-in-the-loop correction together with policy fine-tuning to improve the model performance.” in P9, EXPERIMENTS)
Claim(s) 6, 12, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wang (NPL Title: Dexcap: Scalable and portable mocap data collection system for dexterous manipulation, Year: 2024) in view of Valentin (US Pub. No. 20210004979) in further view of Wang (CN Pub. No. 103251419).
As per Claim 6, the combination of Wang and Valentin teaches or suggests all limitations of Claim 1. Wang further discloses wherein the plurality of sensors comprises:
at least one position sensor at each joint of a plurality of joints that couple the plurality of finger elements to the hand element; (as per “The data collection encompasses four data types, recorded at 60 frames per second: (1) the 6-DoF pose of the chest mounted LiDAR camera, as tracked by the top T265 camera; (2) the 6-DoF wrist poses, as captured by the two lower T265 cameras attached to the gloves; (3) the positions of finger joints within each glove’s reference frame, detected by the motion capture gloves; and (4) RGB-D image frames from the LiDAR camera.” in P15, APPENDIX A, as per “The T265 cameras, initially in a known pose for calibration, are relocated to hand mounts during data collection to monitor palm positions, ensuring consistency through a click-in design. Finger motions are captured by Rokoko gloves, accurately tracking the finger joint positions.” in P3, RELATED WORK)
at least one camera mounted on the wearable data collection device and oriented to face in a direction different from the camera of the mobile device. (as per “The system includes a mocap glove to track finger joints, a camera mounted on top of each glove to track the 6-DoF poses of the wrists with SLAM, and an RGB-D LiDAR camera on the chest to observe the 3D environments” in P2, INTRODUCTION, as per “Mirroring the human system, the robot system reuses the same chest cameras and mount.” in P4, HARDWARE SYSTEM: DEXCAP, as per “As depicted in Figure 2(a), we design a wearable camera vest for this purpose. It incorporates an Intel Realsense L515 RGB-D LiDAR camera, mounted on the top of the chest, to capture the observations during human data collection” in P4, HARDWARE SYSTEM: DEXCAP)
Wang and Valentin fail to expressly disclose:
at least one pressure sensor positioned on each finger element of the plurality of finger elements;
Wang discloses of hand function rehabilitation training and evaluation of the glove, comprising:
at least one pressure sensor positioned on each finger element of the plurality of finger elements; (as per “the data glove further comprises packaging in the glove hand bending sensor set on one side of package in the glove palm flexible pressure sensor set on one side. a bending sensor group is located at each finger joint, the pressure sensor assembly is the grabbed object several key node in contact with the object is curved sensor set, a pressure sensor group of the lead package forearm outer dorsal side in the glove; and connected with the subsequent hardware processing circuit.” in Abstract, as per “said pressure sensor group includes a proximal joint 12 to 15 packaged in a palm side of the glove, the forefinger, middle finger, ring finger and small finger phalanx, a distal phalanx and the base joint and the thumb distal phalanx bone, thumb flexor and small finger bending flexible pressure sensor of muscle and so on” in ¶9)
In this way, Wang operates to provide a data glove having bending sensors and flexible pressure sensors positioned on the glove, including pressure sensors at finger phalanx/joint regions, with the sensors connected to a subsequent hardware processing circuit (Abstract, ¶9). Like Wang and Valentin, Wang is concerned with wearable glove-based sensing of hand movement and hand interaction data.
It would have been obvious for one of ordinary skill in the art before the effective filing date to have modified the system(s) of Wang and Valentin with the glove pressure-sensor arrangement of Wang to enable another standard means of collecting pressure data from the user’s fingers during hand manipulation. Such modification also allows the system to capture grip or contact information in addition to finger position and wrist pose information, thereby providing more complete hand-interaction data for training or evaluating manipulation tasks (Abstract, ¶9).
Claim 12 is rejected using the same rationale, mutatis mutandis, applied to Claim 6 above, respectively.
Claim 17 is rejected using the same rationale, mutatis mutandis, applied to Claim 6 above, respectively.
Claim(s) 7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wang (NPL Title: Dexcap: Scalable and portable mocap data collection system for dexterous manipulation, Year: 2024) in view of Valentin (US Pub. No. 20210004979) in further view of Jarvis (WO Pub. No. 2023039088).
As per Claim 7, the combination of Wang and Valentin teaches or suggests all limitations of Claim 1. Wang further discloses wherein the sensor data captured during the recording session is used to train a neural network that controls a robotic counterpart device having a joint and sensor configuration that matches the wearable data collection device.
Wang and Valentin fail to expressly disclose an activation mechanism configured to:
initiate the recording session in response to a first user input; and terminate the recording session in response to a second user input,
Jarvis discloses a wearable robot data collection system with human-machine operation interface, comprising:
initiate the recording session in response to a first user input; and terminate the recording session in response to a second user input, (as per “in one embodiment the data collector 105 may use the voice user interface 122 to provide audio commands such as “begin recording” at the start of a data collection process or at the start of an instructed action or “stop recording” at the end of the data collection process or the end of an instructed action” in ¶34, as per “the data collector can make a hand gesture that is tracked by the VR/AR glasses 121 to start and stop recording at the start and completion of the data collection process and/or the start and completion of an instructed action. In still other embodiments, a QR code that is located in the data collection location and that can be scanned by the VR/AR glasses 121 can be provided to start and stop recording at the start and completion of the data collection process and/or the start and completion of an instructed action” in ¶34)
In this way, Jarvis operates to start and stop a data collection process or instructed action using user inputs, including audio commands such as “begin recording” and “stop recording,” tracked hand gestures, or a scanned QR code (¶34). Like Wang and Valentin, Jarvis is concerned with collecting data from human-performed actions for use in robotic operation or training.
It would have been obvious for one of ordinary skill in the art before the effective filing date to have modified the system(s) of Wang and Valentin with the start/stop activation arrangement of Jarvis to enable another standard means of initiating and terminating a recording session in response to user input. Such modification also allows the system to define the beginning and end of a data collection process or instructed action, thereby reducing unnecessary recorded data and improving organization of recorded training sessions (¶34).
Claim 14 is rejected using the same rationale, mutatis mutandis, applied to Claim 7 above, respectively.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zhang (US Pub. No. 20230260155) discloses deep continuous 3d hand pose tracking.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER R ROBARGE whose telephone number is (703)756-5872. The examiner can normally be reached Monday - Friday, 8:00 am - 5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ramon Mercado can be reached on (571) 270-5744. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.R.R./Examiner, Art Unit 3658
/Ramon A. Mercado/Supervisory Patent Examiner, Art Unit 3658