Prosecution Insights
Last updated: October 04, 2026
Application No. 18/812,487

SYSTEM AND METHOD FOR HARVESTING FRUIT

Non-Final OA §103§112
Filed
Aug 22, 2024
Priority
Dec 21, 2023 — provisional 63/613,377 +1 more
Examiner
XU, PETER
Art Unit
Tech Center
Assignee
Oishii Farm Corporation
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
8m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
29 currently pending
Career history
26
Total Applications
across all art units

Statute-Specific Performance

§101
4.5%
-35.5% vs TC avg
§103
71.3%
+31.3% vs TC avg
§102
3.8%
-36.2% vs TC avg
§112
16.6%
-23.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the applicant’s communication filed on 8/22/2024 Claims 1-12 are pending Claim Objections Claim 1 objected to because of the following informalities: “(D” in line 12 is missing a “)” and should be corrected to “(D)”, and “at at least one point in time” in line 23 is suggested to be changed to “at one or more points in time”. Appropriate correction is required. Claim 4 objected to because of the following informalities: “where in” in line 1 should be corrected to “wherein”. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 10-11 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 10 recites the limitation "the fruit location" in line 2. There is insufficient antecedent basis for this limitation in the claim. Claim 11 recites the limitation "the fruit location" in line 2. There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-6, 8-9, and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Faulring et al. USPGPUB 2022/0078972 A1 (hereinafter Faulring) in view of Russel et al. WO 2017/152224 A1 (hereinafter Russel) and Cooper et al. US 10,296,602 B1 (hereinafter Cooper), and further in view of Parsa et al. (Autonomous Strawberry Picking Robotic System, 1/10/2023) (hereinafter Parsa). Regarding claim 1, Faulring teaches a system for automatically harvesting fruit from plants (Par. [0022], “the systems and methods described here include an automated and/or semi-automated system with machine(s) that is/are capable of harvesting agricultural targets such as berries from their planter beds without human hands touching the plants or targets themselves”), comprising: (A) one or more robots (Par. [0036], “multiple robotic arms 160 may be fit onto one overall traversing vehicle 152”), each of the one or more robots comprising: (i) a camera (Par. [0042], “at least one picker head may be mounted on or partially mounted on a robotic harvesting arm 160, alone or in combination with sensors such as cameras and/or lighting system(s)”); and (ii) an end effector (Par. [0042], “the harvesting subassembly may include at least one picker head at the end of the robotic arm 160 that first interacts with the target in the field to remove or detach the target from the plant 180 it grows on”; Par. [0043], “the picker head assembly is designed to grasp and remove targets from the plants” – the picker head corresponds to the end effector because it is located at the end of the robotic arm and grasps and removes the fruit); (B) one or more edge devices (Par. [0107], “The computer(s) 802 could be any number of kinds of computers such as those included in the sensors themselves, in the robotic assemblies, image processing and/or another computer arrangement in communication with the camera computer components”; Par. [0123], “Each picking segment 940, 950, etc. may also include a target acquisition or an identification processor 944, 954”), each of the one or more edge devices operatively connected to a corresponding camera of a corresponding one of the one or more robots (Par. [0123], “Each identification processor 944, 954 may also be in communication with an identification camera 945, 955 or two cameras 946, 956” - identification processors 944 and 954 correspond to the edge devices because each processor is locally associated with a respective robotic picking segment and is connected to the camera of that segment.) and configured to receive first image data associated with at least one two-dimensional image captured by the corresponding camera (Par. [0062], “digital, pixelated images taken from the multiple stereo cameras may be processed by a computing system to create three-dimensional (3-D) images using machine vision. In some examples, these images are made of pixels”; Par. [0123], “Such cameras 945, 955, 946, 956 may provide the pixelated image data taken of the targets in order to process for target selection, target coordination, and/or classification of targets” – the individual digital, pixelated image taken by each camera corresponds to the two-dimensional image, while the pixelated image data provided by the corresponding camera to identification processor 944 or 954 correspond to the first image data associated with that two-dimensional image.) and output second image data comprising information associated with the at least one two-dimensional image (Par. [0124], “The identification processor 944, 954 may include artificial intelligence subcomponents and/or neural network programming used to make determinations of target selection and/or coordinate mapping of targets using the image data from the cameras as described below … Such a processing center may perform neural network processing to identify all candidate harvest targets in the acquired imagery and transmit all processed results to the system controller 902” – the pixelated image data received from the cameras, discussed above, correspond to the first image data, while the processed target-identification and coordinate-mapping results generated from that image data correspond to the second image data because they comprise information associated with the captured two-dimensional image. Transmitting the processed results to system controller 902 corresponds to outputting the second image data.) (C) a programmatic logic controller operatively connected to the one or more robots (Par. [0115], “The system controller 902 may coordinate the subfunctions and be in communication with multiple computer components”; Par. [0125], “The motion processing center 947, 957 can receive targets to harvest from the system controller 902 or identification processor 944, 954, in fully autonomous examples. The motion processor 947, 957 may use vehicle position information from the system controller 902 to resolve the relative position of the target to be picked and computes a path for the robotic arm 948, 958 to reach the target” – system controller 902 corresponds to the programmatic logic controller because it coordinates the harvesting subfunctions and supplies target and position information used by the motion processors to control the corresponding robotic arms.); and (D a server comprising a non-transitory computer-readable memory (Par. [0127], “Fig. 10 shows a back-end architecture with a back-end server 1030 and associated data storage 1032”; Par. [0129], “The backend system 1030 may employ the use of a database 1032 to organize, store and retrieve results during and after harvesting operations” - back-end server 1030 corresponds to the server, and associated data storage/database 1032 corresponds to the non-transitory computer-readable memory.) and operatively connected to the programmatic logic controller (Par. [0119], “The controller 902 may also be in communication with a field network 920”; Par. [0128], “images or other sensor data taken from the harvesters 1090 in the field may be communicated through the field network 1020 and data radios 1022, 1024 and an internet router 1028 and through the network 1060 to the back-end server 1030 to be analyzed”; Par. [0129], “The backend system 1030 may route coordinates for harvest targets from remote users 1040 to the appropriate harvester 1090 in the field when the remote user evaluation has been completed.” – the transmission of harvester data to back-end server 1030 and target coordinates back to the harvester through the field network operatively connects back-end server 1030 to system controller 902.), the server comprising: (i) a programmatic logic controller module configured to receive operating state data of the one or more robots from the programmatic logic controller (Par. [0116], “The overall or master system controller 902 may be in communication with a navigation system 910 to receive and analyze positional data of the harvester , such as geographical and/or relative positional data within a field”; Par. [0125], “continuous feedback from the servo camera may monitor the progress of the motion of the arm 948, 958 towards the picking target”; Par. [0119], “The controller 902 may also be in communication with a field network 920”; Par. [0128], “images or other sensor data taken from the harvesters 1090 in the field may be communicated through the field network 1020 and data radios 1022 , 1024 and an internet router 1028 and through the network 1060 to the back-end server 1030 to be analyzed” – The positional data and monitored progress of the robotic arm correspond to operating-state data because they indicate the current position and movement of the harvester and robot; system controller 902 receives and analyzes the positional data and communicates with the field network through which harvester sensor data are communicated to the software of back-end server 1030, which corresponds to the programmatic logic controller module.) input the operating state data to the memory (Par. [0129], “The backend system 1030 may or may not store 1032 acquired imagery and operational data for further analysis to improve situational knowledge of harvester operations regarding business efficiency. The backend system 1030 may employ the use of a database 1032 to organize, store and retrieve results during and after harvesting operations” – the positional, movement, and other harvester-state information are stored as operational data in database 1032, thereby inputting the operating-state data to the memory.) and send robot operating instructions to the programmatic logic controller (Par. [0119], “remote harvesting target selections may be made using the image and/or coordinate data determined by the harvester, sent to a remote user for target selection/classification, and the coordinate data and harvesting instruction sent back to the harvester for harvesting.”; Par. [0129], “The backend system 1030 may route coordinates for harvest targets from remote users 1040 to the appropriate harvester 1090 in the field when the remote user evaluation has been completed.”; Par. [0115], “The system controller 902 may coordinate the subfunctions” – the target coordinates and harvesting instructions correspond to robot operating instructions and are sent by back-end server 1030 to the appropriate harvester, where system controller 902 coordinates the harvesting subfunctions based on the instructions); (ii) one or more communication bridges each associated with a corresponding one of the one or more robots (Par. [0121], “Such a switch may be in communication with one or more picking segments 940, 950 and their associated own network switches 942, 952 respectively.”; Par. [0122], “Each picking segment 940, 950, etc. may include many multiple component parts including a network switch 942, 952 for communication with the system controller 902 by way of the main network switch 930” – network switches 942 and 952 and their associated communication software correspond to the communication bridges because each is respectively assigned to a picking segment having a corresponding robotic arm.), each of the one or more communication bridges configured to receive the second image data (Par. [0123], “Each picking segment 940, 950, etc. may also include a target acquisition or an identification processor 944, 954 and/or a motion processor 947, 957 in communication with the respective network switches 942, 952.”; Par. [0124], “Such a processing center may perform neural network processing to identify all candidate harvest targets in the acquired imagery and transmit all processed results to the system controller 902” – the processed results transmitted by identification processors 944 and 954 correspond to the second image data. Because each identification processor communicates with system controller 902 through its respective network switch 942 or 952, the respective network switch receives the second image data transmitted from the corresponding identification processor.) and store the second image data in the memory (Par. [0109], “In such examples, the image and target data may be stored, analyzed, used to train models, or any other kind of image data analysis.”; Par. [0129], “The backend system 1030 may or may not store 1032 acquired imagery and operational data … The backend system 1030 may employ the use of a database 1032 to organize, store and retrieve results during and after harvesting operations” – the processed image and target information transmitted through the communication bridges is stored in data storage/database 1032 of back-end server 1030.); and 3. make available the pick data to the programmatic logic controller module (Par. [0129], “The backend system 1030 may route coordinates for harvest targets from remote users 1040 to the appropriate harvester 1090 in the field when the remote user evaluation has been completed.”) so that the programmatic logic controller can control the corresponding one of the one or more robots (Par. [0115], “The system controller 902 may coordinate the subfunctions and be in communication with multiple computer components”; Par. [0125], “The motion processing center 947, 957 can receive targets to harvest from the system controller 902 or identification processor 944, 954, in fully autonomous examples” - system controller 902 controls the corresponding robot by coordinating the picking segment and providing the harvesting target to its motion processor.) to move the corresponding end effector in accordance with the pick data to pick the fruit (Par. [0125], “The motion processor 947, 957 may use vehicle position information from the system controller 902 to resolve the relative position of the target to be picked and computes a path for the robotic arm 948, 958 to reach the target”; Par. [0126], “Results from the real-time neural network processing can be used by the motion processor 947, 957 to correct the target path of the robotic arm 948, 958 motion to compensate for variable conditions … Upon reaching the harvesting target, the motion processor 947, 957 may command the actuation of the gripper on the robotic arm 948, 958 to acquire the target and deposit the target for harvesting”). Faulring does not explicitly teach the second image data comprising a corresponding time stamp; (iii) one or more frame synchronization modules each associated with a corresponding one of the one or more robots, each of the one or more frame synchronization modules configured to, at at least one point in time: 1. obtain first operating state data and the second image data from the memory for a corresponding one of the one or more robots; and 2. synchronize the second image data with the corresponding first operating state data; and 3. output, based on the synchronization, first synchronization data to the memory, the synchronization data comprising information associated with the captured at least one image and the corresponding first operating state of the corresponding robot; (iv) an inference module configured to process the first synchronization data output by each of the one or more frame synchronization modules using a neural network, the neural network having been configured through training to receive the synchronization data and to process the synchronization data to generate corresponding output that comprises depth of a fruit image of a fruit within the at least one images captured by the one or more cameras, at least one mask associated with the fruit image within the at least one images, and at least one keypoint associated with the fruit image within the at least one images; (v) a 3D module configured to generate, based on the processed synchronization data and the robot operating state data of each of the one or more robots, three-dimensional image information containing a set of points within three dimensions representing location of the fruit within a three-dimensional world frame; and (vi) an aggregator module configured to: 1. generate, based on the three-dimensional image information, a world map comprising the location of the fruit within the world frame and location of the end effectors within the world frame; 2. determine, based on the world map, pick data comprising information associated with an ideal approach angle for the end effector of a corresponding one of the one or more robots to the fruit to pick the fruit. However, Russel teaches (iii) one or more frame synchronization modules each associated with a corresponding one of the one or more robots (Page 19, Par. 2, “accurate time synchronisation is required between the joint states and the camera data.”; Page 18, Par. 4, “The Kinect Fusion subsystem 202 is configured to receive raw point clouds from an RGB-D sensor and to register consecutive frames into a smoothed point cloud for further processing”; Page 18, Par. 6, “A pose detection state 224 consists of combining the point clouds into a coherent point cloud from multiple viewpoints as the robot arm 110 moves through the scanning motion” - Kinect Fusion subsystem 202 is a frame-processing software module because it registers consecutive RGB-D frames, and the module is associated with robot arm 110 because it processes and combines point clouds obtained as robot arm 110 performs the scanning motion.), each of the one or more frame synchronization modules configured to, at at least one point in time (Page 18, Par. 6, “A pose detection state 224 consists of combining the point clouds into a coherent point cloud from multiple viewpoints as the robot arm 110 moves through the scanning motion” – during the pose detection state, Kinect Fusion performs the frame-registration and point-cloud-combination operations as robot arm 110 moves through the scanning motion, thereby configuring the module to perform the following operations at one or more points in time.): 1. obtain first operating state data and the second image data from the memory for a corresponding one of the one or more robots (Page 18, Par. 6, “The Kinect Fusion subsystem 202 receives raw point clouds from the RGB-D camera 72 and registers consecutive frames into a smoothed point cloud for further processing.”; Page 19, Par. 2, “the robot arm joint states provide a high bandwidth update about the camera's pose” – the raw RGB-D point clouds correspond to image data for robot arm 110, and the joint states correspond to operating-state data because they identify the current robot-arm configuration and resulting camera pose. In the modified system, Russel’s subsystem obtains the corresponding image and operating-state data from Faulring’s database 1032, where acquired imagery and operational data are stored and retrieved.); and 2. synchronize the second image data with the corresponding first operating state data (Page 19, Par. 2, “the robot arm joint states provide a high bandwidth update about the camera's pose. However, an accurate rigid calibration between the camera and the end effector of the robot arm is required. Also, accurate time synchronisation is required between the joint states and the camera data.” - the camera data correspond to the second image data, the robot-arm joint states correspond to the operating-state data, and Russel expressly requires the two corresponding data sets to be synchronized); and 3. output, based on the synchronization, first synchronization data to the memory, the synchronization data comprising information associated with the captured at least one image and the corresponding first operating state of the corresponding robot (Page 19, Par. 1, “The registration method produces two key outputs: an estimate of the current camera pose and a merged point cloud”; Page 19, Par. 2, “the robot arm joint states provide a high bandwidth update about the camera's pose … Also, accurate time synchronisation is required between the joint states and the camera data.” – the merged point cloud is generated from the captured camera data and therefore comprises information associated with the captured image, while the current camera-pose estimate is determined using the corresponding robot-arm joint states and therefore comprises information associated with the corresponding robot operating state. Because Russel requires the camera data and joint states to be time synchronized for the registration process, these outputs constitute synchronization data based on the synchronized image and robot-state information. In the modified system, the synchronization data are output to Faulring’s database 1032 for storage.); (v) a 3D module configured to generate (Page 17, Par. 6, “The software is broken into five different subsystems, which include the Kinect Fusion subsystem 202, a detection and segmentation subsystem 204, a superquadric fitter subsystem 206, a state machine 208 and a path planner subsystem 210.”; Page 17, Par. 7, “The raw information from the RGB-D camera 72 is used within the Kinect Fusion subsystem 202 to reconstruct the 3D scene”; Page 11, Par. 5, “In the first stage, a scanning motion is used to build up a 3D scene of the world using the RGB-D camera 72 … The information from the RGB-D camera 72 is registered using a Kinect Fusion (trademark) (Kinfu) algorithm to produce a high-fidelity 3D scene.” – Russel’s Kinect Fusion subsystem 202 corresponds to the 3D module because the subsystem registers the RGB-D camera information to generate the high-fidelity 3D scene of the world.), based on the processed synchronization data and the robot operating state data of each of the one or more robots (Page 20, Par. 1, “To obtain the point cloud data, and pose ground truth, an RGB-D camera was mounted on a robotic arm and moved over a known trajectory (forwards and backwards only). The pose ground truth was obtained from the odometry of the robotic arm, which was accurate up to 0.1 mm.”; Page 19, Par. 3, “Kinect Fusion provides accurate tracking of the camera pose whilst producing rich scene reconstruction from multiple viewpoints of the camera” – Russel’s three-dimensional reconstruction uses the image-derived point-cloud data together with the robot-arm trajectory and odometry that establish the camera pose, thereby generating the reconstruction based on the processed image information and robot operating-state information.), three-dimensional image information containing a set of points within three dimensions representing location of the fruit within a three-dimensional world frame (Page 13, Par. 6, “Segmented 3D information about a target capsicum is isolated to estimate the pose or orientation of the crop”; Page 14, Par. 1, “The optimisation returns the parameters of the model which describe the shape, size and pose of the capsicum in the world” – Russel processes the segmented three-dimensional point information to determine the fruit’s shape, size, and pose within the world frame.); and (vi) an aggregator module configured to (Page 18, Par. 1, “The state machine 208 is the central node of the software system 200 which interfaces with each other process to perform the harvesting operation.”; Page 17, Par. 7, “The state machine 208 uses the registered scene and the detection and superquadric fitter subsystem 206 is used to estimate the pose of a capsicum. The pose of the capsicum is then used to perform the harvesting actions using a path planner subsystem 210, a robot arm controller 212, and an end effector controller 214.” – state machine 208 corresponds to the aggregator module because it is the central software node that interfaces with the other processes to perform the harvesting operation and uses the registered scene and capsicum-pose information in connection with the path planner, robot-arm controller, and end-effector controller.): 1. generate, based on the three-dimensional image information, a world map comprising the location of the fruit within the world frame and location of the end effectors within the world frame (Page 11, Par. 5, “a scanning motion is used to build up a 3 D scene of the world using the RGB-D camera 72. The camera 72 is part of the end effector 10, which is moved to build up the 3D scene in an eye-in-hand configuration”; Page 14, Par. 1, “The optimisation returns the parameters of the model which describe the shape, size and pose of the capsicum in the world”; Page 19, Par. 1, “The registration method produces two key outputs: an estimate of the current camera pose and a merged point cloud.” - the 3D scene corresponds to the world map, the capsicum pose in the world provides the fruit location in the world map, and the current camera pose provides the location of the end effector because Russel expressly states that camera 72 is part of end effector 10.); and 2. determine, based on the world map, pick data comprising information associated with an ideal approach angle for the end effector of a corresponding one of the one or more robots to the fruit to pick the fruit (Page 15, Par. 4, “The last step in this approach is to estimate the grasp pose using the pose of the capsicum. The rotation of the grasp is first determined in world coordinates”; Page 15, Par. 5, “the x-axis of the world represents the front of the robot and is the axis the grasp pose is to be aligned with”; Page 16, Par. 1-2, “The suction cup 40 is aligned with a selected face of the capsicum using the pose and shape information. The attachment stage includes moving the arm 110 so that the suction cup 40 can suction grip or latch onto the capsicum.” – the grasp pose corresponds to the pick data, and the world-coordinate rotation and alignment of the suction cup with a selected fruit face define the approach angle at which the end effector approaches the fruit for picking.). Faulring and Russel are analogous art because they are from the same field of endeavor and contain functional similarities. They both relate to robotic fruit-harvesting systems that use cameras to identify fruit, determine fruit location, control a robotic arm, and operate an end effector to harvest the fruit. Therefore, at the time of effective filing date, it would have been obvious to a person of ordinary skill in the art to modify the above robotic fruit-harvesting system, as taught by Faulring, and incorporate synchronized camera-data and robot-joint state registration, three-dimensional scene reconstruction, world-frame fruit-pose determination, and grasp-pose determination techniques, as taught by Russel. One of ordinary skill in the art would have been motivated to improve the accuracy of camera-pose tracking and three-dimensional scene reconstruction, as suggested by Russel (Page 19, Par. 3). Faulring and Russel do not explicitly teach the second image data comprising a corresponding time stamp; and (iv) an inference module configured to process the first synchronization data output by each of the one or more frame synchronization modules using a neural network, the neural network having been configured through training to receive the synchronization data and to process the synchronization data to generate corresponding output that comprises depth of a fruit image of a fruit within the at least one images captured by the one or more cameras, at least one mask associated with the fruit image within the at least one images, and at least one keypoint associated with the fruit image within the at least one images. However, Cooper teaches the second image data comprising a corresponding time stamp (Col. 4, lines 39-41, “receiving, from the vision component, the image frame and corresponding metadata generated by the vision component”; Col. 5, lines 10-13, “the corresponding metadata generated by the vision component comprises a vision component generated timestamp that is based on the vision component clock domain”). Faulring, Russel, and Cooper are analogous art because they contain functional similarities. They all relate to robotic systems that use camera data and robot operating-state data for robotic control. Therefore, at the time of effective filing date, it would have been obvious to a person of ordinary skill in the art to modify the above synchronized robotic fruit-harvesting system, as taught by Faulring and Russel, and incorporate a corresponding timestamp with the image data, as taught by Cooper. One of ordinary skill in the art would have been motivated to accurately correlate the image data with the corresponding robot operating-state data at the time the image was captured, as suggested by Cooper (Col. 1, lines 21-25). Faulring, Russel, and Cooper do not explicitly teach (iv) an inference module configured to process the first synchronization data output by each of the one or more frame synchronization modules using a neural network, the neural network having been configured through training to receive the synchronization data and to process the synchronization data to generate corresponding output that comprises depth of a fruit image of a fruit within the at least one images captured by the one or more cameras, at least one mask associated with the fruit image within the at least one images, and at least one keypoint associated with the fruit image within the at least one images. However, Parsa teaches (iv) an inference module configured to process the first synchronization data output by each of the one or more frame synchronization modules using a neural network (Page 7, Par. 1, “Our comprehensive perception system includes an RGB-D sensor and three RGB cameras, a novel dataset, and state-of-the-art algorithms to detect and localise the fruit and determine its suitability for picking … These sensors are coupled with a novel Mask-RCNN-based algorithm to form the perception system”; Page 9, Par. 1, “perception acquires the data, i.e. images, depths, point cloud, from multiple sensors. After pre-processing the received data, the perception algorithm detects all strawberries in the field of view and publishes their coordinates through the berry topic.”; Page 17, Par. 1, “Our proposed approach includes Detectron-2 (Wu et al., 2019) for segmentation and key-points estimation. The Detectron-2 model is based on MRCNN” – Parsa’s perception system corresponds to the inference module because it receives and processes image, depth, and point-cloud information using the MRCNN-based Detectron-2 neural network to detect and localize strawberries. In the modified system, the synchronized RGB-D information produced according to Russel is supplied to Parsa’s perception system.), the neural network having been configured through training to receive the synchronization data and to process the synchronization data (Page 16, Perception for selective harvesting robotic system, “our approach includes SOTA MRCNN models. We collected two datasets to train our models … Dataset-1 is a novel dataset that presents strawberry dimensions, weights, suitability for picking, instance segmentation, and key-points for grasping and picking action”; Page 17, Segmentation, Key-points and Pluckable Detection, “The datasets' key-points, segmentation masks, and strawberry categories (`pluckable' and ‘"unpluckable"') are converted to MSCOCO JSON format (Lin et al., 2015). This MSCOCO JSON is the default format for feeding data into Detectron-2” - Parsa trains the MRCNN-based Detectron-2 model using strawberry images annotated with segmentation masks, keypoints, and picking classifications. In the modified system, the image portion of the synchronized RGB-D information generated according to Russel is provided to the trained model for processing.) to generate corresponding output that comprises depth of a fruit image of a fruit within the at least one images captured by the one or more cameras (Page 17, Perception Setup, “The vision system consists of three cameras: An Intel Realsense d435i color and depth-sensing camera and three colors (RGB) cameras”; Pages 17–18, Perception pipeline and approach, “At the home position, all strawberries are detected in the top (RealSense) camera image frame and scheduled … the 2D segmented strawberry pixels are used as binary masks on the depth image to filter the depth pixels belonging to the strawberry … Then we take the average value of the remaining strawberry depth pixels which gives us a more reliable 3D coordinate” - Parsa’s inference module uses the detected strawberry segmentation to select and process the corresponding strawberry-depth pixels and generates depth-based three-dimensional coordinate information associated with the strawberry image.), at least one mask associated with the fruit image within the at least one images (Page 17, Segmentation, Key-Points and Pluckable Detection, “Our proposed approach includes Detectron-2 (Wu et al., 2019) for segmentation and key-points estimation. The Detectron-2 model is based on MRCNN (He et al., 2017) and has become the standard for instance segmentation”; Pages 17, Perception pipeline and approach, “the 2D segmented strawberry pixels are used as binary masks on the depth image to filter the depth pixels belonging to the strawberry” - Detectron-2 generates an instance segmentation of the detected strawberry, and the segmented strawberry pixels form a binary mask associated with that strawberry image.), and at least one keypoint associated with the fruit image within the at least one images (Page 16, Par. 4, “For each strawberry, the dataset presents five different key-points: the picking point (PP), the top, and bottom points of the fruit, the left grasping point (LGP), and the right grasping point”; Page 17, Segmentation, Key-points and Pluckable Detection, “We adapted this key-points detection method integrated within Detectron-2 to estimate the strawberry key-points” - the trained Detectron-2 model estimates multiple keypoints specifically associated with each detected strawberry, including its picking point and grasping points); Faulring, Russel, Cooper, and Parsa are analogous art because they are from the same field of endeavor and/or contain functional similarities. Faulring, Russle, and Parsa all relate to robotic fruit-harvesting systems that use camera-based perception to detect and localize fruit, determine picking information, and guide a robotic end effector to harvest the fruit. Cooper relates to robotic systems that correlate image data with corresponding robot sensor-state data. Therefore, at the time of effective filing date, it would have been obvious to a person of ordinary skill in the art to modify the above timestamped and synchronized robotic fruit-harvesting system, as taught by Faulring, Russel, and Cooper, and incorporate a trained Mask-RCNN-based Detectron-2 perception process and associated depth-filtering process so that the synchronized RGB-D image information is processed to generate strawberry segmentation masks, strawberry keypoints, and depth-based strawberry-location information, as taught by Parsa. One of ordinary skill in the art would have been motivated to improve the robustness of fruit picking-point localization, as suggested by Parsa (Page 6, Par. 4). Regarding claim 2, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Parsa further teaches wherein the output of the inference module further comprises fruit ripeness detection (Page 7, Par. 1, “This sensor provides an RGB image of the plant and also a three-dimensional point cloud to detect fruit, localise the picking point and predict the ripeness of the fruit”; Page 9, System work-flow and algorithm, “the perception classifies all detected berries as "pluckable" or "unpluckable" which is included in the berry topic.”; Page 16, Par. 4, “"unpluckable" strawberries include unripe, semi-, and over-ripe or rotten berries. The ‘pluckable’ category includes strawberries that are nearly ripe and perfectly ripe”), at least one bounding box (Page 11, Target Berry Selection Using Min/Max, “The min-max algorithm attempts to find the maximum of the minimum distances among all the bounding boxes of the detected berries”), and at least one object detection (Page 9, System work-flow and algorithm, “After pre-processing the received data, the perception algorithm detects all strawberries in the field of view and publishes their coordinates through the berry topic.”). Regarding claim 3, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Parsa further teaches wherein the server further comprises a safety module configured to filter the pick data based on safety parameters (Page 9, Control and Motion Planning, “As the space between the strawberry plants and the robot arm is very limited and the environment contains a high level of uncertainty, there is a high possibility of collision of the robot arm or the end-effector with different objects … Different motion planning algorithms and trajectory/velocity/acceleration profiles were employed for each segment of movement to ensure collision-free and efficient manipulation” – Parsa’s motion-planning software filters and modifies the end-effector path, velocity, and acceleration before execution based on collision-avoidance and safe-manipulation requirements. In the modified system, the safety module is implemented in Faulring’s server.). Regarding claim 4, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Faulring further teaches where in the end effector is a gripper (Par. [0126], “the motion processor 947, 957 may command the actuation of the gripper on the robotic arm 948, 958 to acquire the target and deposit the target for harvesting”). Regarding claim 5, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Parsa further teaches wherein the gripper comprises a grip portion configured to hold a stem of a fruit and a cutting portion configured to cut the stem while the stem is held by the grip portion (Page 14, End-effector design, “the second pair of fingers for gripping a stem of the identified ripe fruit (Grippers); and a cutting mechanism for cutting the stem of the identified ripe fruit when the stem is gripped between the second pair of fingers (Cutter)”). Regarding claim 6, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Faulring further teaches wherein the one or more robots comprise a plurality of robots (Par. [0036], “multiple robotic arms 160 may be fit onto one overall traversing vehicle 152 … up to eight picker arms 160 may be employed”). Regarding claim 8, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Parsa further teaches wherein the one or more communication bridges are ethernet bridges (Page 7, Par. 2, “The second laptop is connected to the first laptop using an Ethernet connection and communicating through ROS”; Page 9, System work-flow and algorithm, “the entire algorithm was implemented in two laptops communicating through Ethernet protocol” – In the combined system, Faulring’s network switches 942 and 952 remain the communication bridges respectively associated with the robotic picking segments, while Parsa’s Ethernet protocol is used to communicate the processed perception data between the perception processing and robot-control portions of the system.). Regarding claim 9, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Parsa further teaches wherein the neural network is of a type selected from the group consisting of Mask R-CNN, YOLOACT, Keypoint R-CNN, GSNet, Detectron2 and PointRend (Page 17, Par. 1, “Our proposed approach includes Detectron-2 (Wu et al., 2019) for segmentation and key-points estimation. The Detectron-2 model is based on MRCNN”). Regarding claim 12, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Parsa further teaches wherein the plants are flowering crops (Page 16, Par. 5, “The strawberries not annotated for key-points are either severely occluded or are in an early flowering stage where a meaningful annotation is not possible”). Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Faulring et al. USPGPUB 2022/0078972 A1 (hereinafter Faulring) in view of Russel et al. WO 2017/152224 A1 (hereinafter Russel), Cooper et al. US 10,296,602 B1 (hereinafter Cooper), and Parsa et al. (Autonomous Strawberry Picking Robotic System, 1/10/2023) (hereinafter Parsa), and further in view of Avigad et al. USPGPUB 2020/0128744 A1 (hereinafter Avigad). Regarding claim 7, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Faulring, Russel, Cooper, and Parsa do not explicitly teach wherein the plurality of robots are arranged on scaffolding that has multiple levels. However, Avigad teaches wherein the plurality of robots are arranged on scaffolding that has multiple levels (Par. [0020], “A plurality of four-dimensional (4-Degrees-of-Freedom-D.O.F) linear robots are mounted in the frame and configured to harvest fruit from the sector. The robots are arranged in pairs, which are stacked vertically in the frame.”; Par. [0033], “three actuators 32 (also referred to as "linear stages" or simply "stages") are mounted in frame 28. Two actuators 36 (also referred to as "robots") are mounted on each actuator 32” – vertical frame 28 corresponds to the scaffolding because it is a supporting framework having three vertically arranged stages, with a pair of fruit-harvesting robots mounted on each stage, thereby arranging the plurality of robots at multiple levels.). Faulring, Russel, Cooper, Parsa, and Avigad are analogous art because they are from the same field of endeavor and/or contain functional similarities. Faulring, Russel, Parsa, and Avigad all relate to robotic fruit-harvesting systems that use one or more robotic manipulators, camera-based fruit perception, and automated control to locate and harvest fruit. Therefore, at the time of effective filing date, it would have been obvious to a person of ordinary skill in the art to modify the above synchronized, vision-guided, multi-robot fruit-harvesting system, as taught by Faulring, Russel, Cooper, and Parsa, to incorporate a vertically staged supporting frame so that the harvesting robots are arranged at multiple levels, as taught by Avigad. One of ordinary skill in the art would have been motivated to improve the vertical harvesting coverage of the multi-robot system and avoid collision between stages, as suggested by Avigad (Par. [0034]). Claim(s) 10-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Faulring et al. USPGPUB 2022/0078972 A1 (hereinafter Faulring) in view of Russel et al. WO 2017/152224 A1 (hereinafter Russel), Cooper et al. US 10,296,602 B1 (hereinafter Cooper), and Parsa et al. (Autonomous Strawberry Picking Robotic System, 1/10/2023) (hereinafter Parsa), and further in view of Robertson et al. USPGPUB 2019/0261566 A1 (hereinafter Robertson). Regarding claim 10, the combination of Faulring, Russel, Cooper, and Parsa teaches all the limitations of the base claims as outlined above. Faulring, Russel, Cooper, and Parsa do not explicitly teach wherein the 3D module is further configured to determine points within the three dimensions around the fruit location that are occluded. However, Robertson teaches wherein the 3D module is further configured to determine points within the three dimensions around the fruit location that are occluded (Par. [0128], “an alternative and innovative approach is to use an implicit 3D model of the scene formed by the range of viewpoints from which the target fruit can be observed without occlusion … By identifying one or more viewpoints from which the target fruit appears un-occluded, obstacle free region of space is found … Occlusion of the target fruit by an obstacle between the fruit and the camera when viewed from a particular viewpoint can be detected by several means including e.g. stereo matching” - Robertson forms a three-dimensional model of the scene and detects whether an obstacle blocks the spatial region extending between a viewpoint and the fruit. In the combined system, Russel’s point-based three-dimensional scene represents the obstacle-free and blocked spatial regions as points within three dimensions around the fruit location, and Robertson’s occlusion-detection process determines which of those points are occluded). Faulring, Russel, Cooper, Parsa, and Robertson are analogous art because they are from the same field of endeavor and/or contain functional similarities. They all relate to robotic fruit-harvesting systems that use camera-based three-dimensional perception to locate fruit, determine an approach to the fruit, and control a robotic end effector to pick the fruit. Therefore, at the time of the effective filing date, it would have been obvious to a person of ordinary skill in the art to modify the above synchronized, vision-guided robotic fruit-harvesting system, as taught by Faulring, Russel, Cooper, and Parsa, to incorporate a three-dimensional occlusion-detection and viewpoint-selection process, as taught by Robertson. One of ordinary skill in the art would have been motivated to improve the identification of an obstacle-free approach region so that the picking head and robotic arm can approach the fruit without colliding with surrounding obstacles, as suggested by Robertson (Par. [0128]). Regarding claim 11, the combination of Faulring, Russel, Cooper, Parsa, and Robertson teaches all the limitations of the base claims as outlined above. Robertson further teaches wherein the pick data is determined by selecting a least occluded three-dimensional image around the fruit location (Par. [0184], “the system can gain more information about the target (e.g. its shape and size , its suitability for picking , its pose ) by obtaining more views from new viewpoints … the best viewpoint might be selected”; Par. [0185], “If a target fruit (or its stem) is partially occluded (by foliage, other fruits, etc.) then it may be valuable to move in the direction required to reduce the amount of occlusion. Generally, it is desirable to find a viewpoint from which the whole fruit is visible without occlusion because such a viewpoint defines, via the back-projected silhouette, a volume of space in which the picking head (and picked fruit) can be moved towards the target fruit without colliding with any other obstacles.” - Robertson evaluates views around the fruit and selects a view having reduced or no occlusion. A view from which the entire fruit is visible without occlusion is the least-occluded view, and, in the combined three-dimensional scene, corresponds to the least-occluded three-dimensional image around the fruit location. The selected view defines an obstacle-free approach volume for moving the picking head toward the fruit and therefore is used to determine the corresponding pick data). Citation of Pertinent Prior Art The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Davidson et al. [USPGPUB 2016/0073584 A1] teaches an autonomous robotic fruit-harvesting system comprising a manipulator, a machine-vision system that determines the location of fruit, and an end effector movable in three-dimensional space, wherein the end effector approaches the fruit along an angle providing a direct approach, grasps the fruit and its stem, and picks the fruit. Zhang et al. [USPGPUB 2021/0212257 A1] teaches a machine-vision fruit and vegetable picking system that uses a pretrained Mask R-CNN neural-network model to identify fruit from captured images, performs instance segmentation, determines a cuttable stalk area and a cutting point, and controls an end-picking apparatus to clamp and cut the fruit stalk. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PETER XU whose telephone number is (571)272-0792. The examiner can normally be reached Monday-Friday 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mohammad Ali can be reached at (571) 272-4105. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PETER XU/ Examiner, Art Unit 2119 /MOHAMMAD ALI/ Supervisory Patent Examiner, Art Unit 2119
Read full office action

Prosecution Timeline

Aug 22, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
2y 10m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month