DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This correspondence is in response to amendments filed on May 15, 2026. Claims 1, 6, 9, and 14 are amended. Claims 2-4, 7-8, 10-13, and 15 are as originally filed or previously presented. Claim 5 is cancelled. Claims 16-19 are new. An Examiner’s response to arguments is included below. Amendments to claims 9 and 14 obviate the claim objections set forth in the previous rejection and as such those rejections have been withdrawn.
Response to Arguments
Applicant argues that Maehara does not teach a “circumscribed circle of the polygon GC” (Remarks Pages 14 and 15). Applicant’s arguments with respect to the “circumscribed circle of the polygon GC” have been considered but are moot because the new ground of rejection does not rely on the same combination of references applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-4 and 6-19 are rejected under 35 U.S.C. 103 as being unpatentable over Humayun et al. (US 2022/0016766 A1; hereinafter “Humayun”) in view of Maehara et al. (US 2011/00774171 A1; hereinafter “Maehara”) and further in view of Redmon et al. (US Patent No. 10,089,575; hereinafter “Redmon”).
NOTE: Wherein the claims have been rejected below, Examiner includes a “supplementary document” attached to the file to further Applicant’s understanding of the alleged combination of Maehara and Redmon. Examiner will hereinafter refer to Figs. 1-3 of “supplementary document” depending on which information is to be referred to.
Regarding claim 1, Humayun teaches an object manipulation apparatus (robotic arm 120 in Fig. 2A and 2B) comprising:
one or more hardware processors coupled to a memory (The “processing system” includes one or more processors connected to a non-transitory computer readable storage medium [0124].) and configured to function as
a feature calculation unit to calculate a feature map indicating a feature of a captured image of a grasping target object ("The graspability network can rapidly generate graspability scores for pixels and/or a graspability map 108 for an image of the scene, wherein the grasp(s) can be selected based on the graspability scores. In some variants, auxiliary scene information can also be generated in parallel (e.g., the object detector can be run on the image to extract object poses), wherein the grasps can be further selected based on the auxiliary data (e.g., the grasps identified from the heatmap can be prioritized based on the corresponding object poses)." [0024]. The features calculated for the associated map are the graspability scores and grasping object poses.);
a region calculation unit to …
calculate a position and a posture of a handling tool ("The grasp selector 146 is preferably configured to select grasp points from the output of the graspability network, but can additionally or alternatively be configured to select grasp points from the output of the object detector (e.g., an object detector can pre-process inputs to the grasp selector)." [0041]. "Planning the grasp can include determining a grasp pose, where the grasp is planned based on the grasp point and the grasp pose. In a first variant, the grasp pose can be determined from the object parameters output by an object detector (e.g., running in series and/or parallel with the graspability network/grasp selector, based on the same or a contemporaneously-captured image), and planning the grasp for the object parameters for the detected object that encompasses (e.g., includes, is associated with) the grasp point" [0099]. The system, as outlined, thus uses a grasp selector to select a grasp point and object to grasp before planning the grasp, inclusive of the position and posture of the robotic manipulator. The grasp selector selects grasp points on the basis of the graspability network, i.e. the feature map.) … , the position and the posture enabling the handling tool to grasp the grasping target object (“The computing system can include a motion planner 148, which functions to determine control instructions for the robotic arm to execute a grasp attempt for a selected grasp point” [0043]. Thus, the position and posture enable the handling tool to grasp the grasping target object by determining control instructions for a selected grasp point.);…
However, Humayun does not explicitly teach …detect a circular anchor on the feature map, and
calculate a position and a posture of a handling tool by a first parameter representing a polygon for which the circular anchor is a circumscribed circle, the position and the posture enabling the handling tool to grasp the grasping target object; and
a grasp configuration (GC) calculation unit to calculate a polygon GC of the handling tool by converting the position and the posture indicated by the first parameter into a second parameter indicating a position and a posture of the handling tool on the captured image,
wherein the circular anchor is a circumscribed circle of the polygon GC,
the first parameter includes a first sub-parameter indicating coordinates of a center of the circumscribed circle and a second sub-parameter indicating a radius of the circumscribed circle, and
the GC calculation unit calculates the polygon GC by acquiring the second parameter from the first sub-parameter, the second sub-parameter, and a third sub-parameter indicating a midpoint of a side of the polygon GC.
Maehara, pertinent to the problem at hand, teaches …detect a circular anchor on the feature map (Maehara sets a “possibility of interference J” as shown in Fig. 6 “anchored” at the grasping center which the fingers symmetrically open and close around as designated in [0086].), and
calculate a position and posture of a handling tool by a first parameter representing an estimation of the grasping attitude for which the circular anchor is a circumscribed circle, the position and posture enabling the handling tool to grasp the grasping target object (See Fig. 3 S130 “set grasping target” which determines the grasping target based on grasping candidates and grasping attitude (see Paragraphs [0074-0079]). Viable grasps are selected based on the center of the graspable member, i.e., center of the circular anchor, as well as the non-interference attitude ranges, i.e., radius of the circular anchor (see Fig. 15-16). The center and radius are determined as the first parameter of the circular anchor which estimates the grasping attitude. Furthermore, we will consider the center of the estimated midpoint of the finger on one side of the grip configuration. Further, S150 performs grasping based on the position and posture of the fingers determined when selecting the grasping target.); …
a grasp configuration (GC) calculation unit to calculate a …GC of the handling tool (The grasp configuration is determined with respect to the position of the approximated finger positions within the circular anchor in [0055].)…
wherein the circular anchor is a circumscribed circle of the … GC (As the “possibility of interference J” is the circular anchor, it can be seen that said circle is circumscribed around the estimated grip configuration of fingers 5 which are equally separated.),
the first parameter includes a first sub-parameter indicating coordinates of a center of the circumscribed circle (Center O which indicates the center of the grasp configuration between fingers.) and a second sub-parameter indicating a radius of the circumscribed circle (The radius of “possibility of interference J” which is shown in Fig. 6 as r(i,j)+Rf which is the radius of the circle of the graspable member and associated diameter of the estimated finger.)…
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to have modified the estimation of a grasp pose as taught by Humayun by fixing a circular anchor to the image as taught by Maehara with a reasonable expectation of success. One of ordinary skill in the art would have been motivated to make such a modification because the inclusion of the circular anchor (“the possibility of interference”) in determining the grip configuration for the fingers in grasping the graspable member allows for steady grasps regardless of any attitude recognition errors which result from the visual sensing or positioning errors of the grasping attitude (Maehara, [0082]).
However, given that Maehara merely teaches an estimation for individual fingers, rather than the entire span of the grip configuration, Humayun as modified by Maehara still does not teach …calculate a position and a posture of a handling tool by a first parameter representing a polygon for which the circular anchor is a circumscribed circle…
…a grasp configuration (GC) calculation unit to calculate a polygon GC of the handling tool by converting the position and the posture indicated by the first parameter into a second parameter indicating a position and a posture of the handling tool on the captured image,
wherein the circular anchor is a circumscribed circle of the polygon GC, … and
the GC calculation unit calculates the polygon GC by acquiring the second parameter from the first sub-parameter, the second sub-parameter, and a third sub-parameter indicating a midpoint of a side of the polygon GC.
Redmon, pertinent to the problem at hand, teaches an estimation of a polygon GC as shown in Figs. 2-4. The specification describes, “The graphical representation 260 is a “grasping rectangle” that defines a grasping coordinate 262, a grasping width 264, a grasping height 266, and an orientation parameter 268 of a robot end effector for the grasp. The grasping coordinate 262 defines a two-dimensional (e.g., “X” and “y”) coordinate that defines a “center” or other reference point of a robot end effector. The grasping width 264 defines a distance between two or more opposing actuable members of the robot end effector, such as a distance between opposing plates of a parallel plate gripper end effector. The grasping height 266 defines a span of each of one or more opposing actuable members of a robot end effector, such as the span of each plate of a parallel plate gripper. It is noted that some robot end effectors may have fixed heights and other end effectors may have adjustable heights (e.g., plates of adjustable sizes). The orientation parameter 268 defines an orientation angle of the end effector for the grasp of the object relative to a reference orientation such as an “X” axis (left to right in FIG. 2) of the image 250. In some implementations, the orientation parameter 268 may comprise two separate parameters to account for the two-fold rotationally symmetric nature of grasp angles” (C5,L65 through C6,L20). Thus, provided that the “grasping rectangle” details similar grasping parameters which describe the grasping configuration of Fig. 6 in Maehara, one of ordinary skill in the art may be adequately motivated to simply substitute the estimation of the finger positions as detailed by Maehara (one known element) with the estimation of the “grasping rectangle” as detailed by Redmon (another known element) to obtain predictable results. This substitution would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (see MPEP 2143.I(B)). Furthermore, one of ordinary skill in the art may be motivated to make this modification because the “grasping rectangle” of Redmon would not only indicate the position of the opposing gripping plates of the end effector, i.e., gripper, but additionally the potential occupancy of the opposing gripping plates as they close in around the object to be grasped.
Reliant on the substitution detailed above, Examiner concludes that the combination of references teaches …detect a circular anchor on the feature map (Newly depicted “possibility of interference J” as shown in Fig. 1 of supplementary document.), and
calculate a position and a posture of a handling tool by a first parameter representing a polygon for which the circular anchor is a circumscribed circle (The position and posture of the handling tool is represented as the “grasping rectangle” of Redmon, i.e., first parameter representing a polygon, which is circumscribed by the “possibility of interference J” of Maehara as shown in Fig. 1 in supplementary document.), the position and the posture enabling the handling tool to grasp the grasping target object (S150 of Maehara and additionally 812 of Fig. 8 of Redmon detail that the estimated grasp configuration which is the associated position and posture enable the handling tool to grasp the grasping target object.)…
wherein the circular anchor is a circumscribed circle of the polygon GC (Again, as shown in Fig. 1 of supplementary document, “the possibility of interference J”, i.e., circular anchor, is a circumscribed circle of the “grasping rectangle”, i.e., polygon GC.)…
Although not explicit, the following limitations are implied such that the modified references teach …a grasp configuration (GC) calculation unit to calculate a polygon GC of the handling tool by converting the position and the posture indicated by the first parameter into a second parameter indicating a position and a posture of the handling tool on the captured image (Figs. 2-3 of the supplementary document show how the variables indicated by Applicant as the first parameter (center of a circumscribed circle, radius of a circumscribed circle, and midpoint of a side of the polygon GC) are obvious variants of the second parameter (height, width, angle, and center) of the “grasping rectangle” disclosed by Redmon through the use of basic geometry.),
…the first parameter includes a first sub-parameter indicating coordinates of a center of the circumscribed circle (Center O as shown in Fig. 1 of supplementary document and additionally coordinate (x,y) as shown in Figs. 2-3 of supplementary document.) and a second sub-parameter indicating a radius of the circumscribed circle (Radius r shown in Fig. 2 of supplementary document.), and
the GC calculation unit calculates the polygon GC by acquiring the second parameter from the first sub-parameter, the second sub-parameter, and a third sub-parameter indicating a midpoint of a side of the polygon GC (As can be seen in Figs. 2-3 of supplementary document the second parameter (grasp height, grasp width, grasp orientation, and grasp coordinate) are obvious variants of the first sub-parameter, i.e., center (x,y), second sub-parameter, i.e., radius r, and third sub-parameter, i.e., midpoint dRx and dRy.).
Therefore, although the references do not explicitly determine that the polygon GC of the handling tool converts the position and posture indicated by the first parameter into a second parameter, both Maehara and Redmon provide the necessary information which would render one of ordinary skill in the art to arrive at the claimed invention. This conclusion is motivated as a combination of variables in that the set of variables as taught by Redmon may be included with the desired variables by known methods of basic geometry to yield predictable results (MPEP 2143.I(A)). Such a modification would be obvious to try as this change of variables is one of a finite number of identified, predictable solutions, with a reasonable expectation of success (MPEP 2143.I(E)). Additionally, one of ordinary skill in the art may be prompted to vary the variables which were used based on design incentives as such variations are predictable to one of ordinary skill in the art (MPEP 2143.I(F)).
Regarding claim 2, Humayun as modified by Maehara further modified by Redmon teaches the apparatus according to claim 1,
with Humayun further teaching wherein the feature calculation unit implements the calculation of the feature map by receiving input of a plurality of pieces of image sensor information ("The input to the graspability network can include: an RGB image, receptive field from image, optionally depth, optionally object detector output (e.g., object parameters, etc.), and/or any other suitable information. In a first variant, the input to the graspability network is a 2D image having 3 channels per pixel (i.e., red-green-blue; RGB). In a second variant, the input to the graspability network can be a 2.5D image having 4 channels per pixel (RGB-depth image). In a first example, the depth can be a sensed depth (e.g., from a lower-accuracy sensor or a higher-accuracy sensor such a Lidar). In a second example, the depth can be a ‘refined’ depth determined by a trained depth enhancement network (e.g., wherein the depth enhancement network can be a precursor neural network or form the initial layers of the graspability network; etc.). In a third variant, the input to the graspability network can include an object detection output as an input feature (e.g., an object parameter, such a characteristic axis of a detected object)" [0069]. Thus, the feature map receives a plurality of input from a plurality of image sensor information including color, object detection, and depth perception images.),
extracting a plurality of intermediate features by a plurality of feature extractors from the plurality of pieces of image sensor information, integrating the plurality of intermediate features, and fusing, by convolution calculation, features of the plurality of pieces of image sensor information including the plurality of intermediate features ("The graspability network is preferably trained based on the labelled images. The labelled images can include: the image (e.g., RGB, RGB-D, RGB and point cloud, etc.), grasp point (e.g., the image features depicting a 3D physical point to grasp in the scene), and grasp outcome; and optionally the object parameters (e.g., object pose, surface normal, etc.), effector parameters (e.g., end effector pose, grasp pose, etc.), and/or other information. In particular, the graspability network is trained to predict the outcome of a grasp attempt at the grasp point, given the respective image as the input. However, the network can additionally or alternatively be trained based on object parameters and/or robotic manipulator parameters, such as may be used to: train the graspability network to predict the object parameter values (or bins) and/or robotic manipulator parameter values (or bins), given the respective image as input" [0075]. "The graspability network can be a neural network (e.g., CNN, fully connected, etc.), such as a convolutional neural network (CNN), fully convolutional neural network (FCN), artificial neural network (ANN), a feed forward network, a clustering algorithm, and/or any other suitable neural network or ML model" [0039]. The image sensor information is labelled such that it may be fed into a training model in order to delineate features for each of the images. The features are integrated such that they are provided to a convolutional neural network model in order to fuse the image data and determine graspability outcomes.).
Regarding claim 3, Humayun as modified by Maehara further modified by Redmon teaches the apparatus according to claim 2,
with Humayun further teaching wherein the plurality of pieces of image sensor information include a color image indicating a color of the captured image and a depth image indicating a distance from camera to objects in the captured image (Input for the graspability network includes RGB color images with red-green-blue inputs and depth images indicating distance inputs from the sensor data [0069].), and
the plurality of feature extractors is implemented by a neural network having an encoder-decoder model structure ("The graspability network can include an encoder (e.g., VGG-16, ResNet, etc.), a decoder (e.g., CCN decoder, FCN decoder, RNN-based decoder, etc.), and/or any other suitable components" [0039].).
Regarding claim 4, Humayun as modified by Maehara further modified by Redmon teaches the apparatus according to claim 1,
with Humayun further teaching wherein the one or more hardware processors are further configured to function as a position heatmap calculation unit to calculate a position heatmap indicating success probability for grasping target object ("The computing system can include a graspability network 144 which functions to determine a grasp score (e.g., prediction of grasp success probability) for points and/or regions of an image... In one example, the graspability network 144 functions to generate a graspability map (e.g., grasp score mask, a heatmap) for an object scene" [0038]. The grasp scores determine the probability of success for each predicted grasp point and applies this information to each pixel/image region to produce the graspability map which may take the form of a heatmap.), and
the region calculation unit is implemented by a neural network … on the basis of the position heatmap ("The graspability network can be a neural network (e.g., CNN, fully connected, etc.), such as a convolutional neural network (CNN), fully convolutional neural network (FCN), artificial neural network (ANN), a feed forward network, a clustering algorithm, and/or any other suitable neural network or ML model. The graspability network can include an encoder (e.g., VGG-16, ResNet, etc.), a decoder (e.g., CCN decoder, FCN decoder, RNN-based decoder, etc.), and/or any other suitable components" [0039]. "The grasp selector 146 is preferably configured to select grasp points from the output of the graspability network.... Additionally or alternatively, the grasp selector can function to select a grasp point based on a plurality of object poses and/or based on a graspability heat map (e.g., grasp score mask; examples are shown in FIGS. 6, 7, and 8)" [0041]. Thus, with regard to the grasp selector, i.e., the region calculation unit, the graspability network and graspability heat map is the driving force behind the grasp selector. The graspability network is a trained neural network.)…
However, Humayun does not teach …detecting the circular anchor on the feature map… and calculating the first parameter on the detected circular anchor.
Provided that Humayun determines a “grasping point”, i.e., grasp center, based on the graspability heat map, Maehara further modified by Redmon additionally teaches …detecting the circular anchor on the feature map on the basis of the position heat map and calculating the first parameter on the detected circular anchor (The circlular anchor is centered on the “grasping point”, i.e., center O of Fig. 6 of Maehara and Fig. 1 of the supplementary document, such that the circular anchor is detected on the feature map on the basis of the position and thereby calculates the first parameter on the detected circular anchor as described above in the rejection of claim 1.).
Regarding claim 6, Humayun teaches a handling method implemented by a computer (The processing system, i.e., computer, is responsible for performing the methods [0124].), the method comprising:
calculating a feature map indicating a feature of an image on the basis of image sensor information including a grasping target object ("The graspability network can rapidly generate graspability scores for pixels and/or a graspability map 108 for an image of the scene, wherein the grasp(s) can be selected based on the graspability scores. In some variants, auxiliary scene information can also be generated in parallel (e.g., the object detector can be run on the image to extract object poses), wherein the grasps can be further selected based on the auxiliary data (e.g., the grasps identified from the heatmap can be prioritized based on the corresponding object poses)." [0024]. The features calculated for the associated map are the graspability scores and grasping object poses.); …
calculating a position and a posture of a handling tool ("The grasp selector 146 is preferably configured to select grasp points from the output of the graspability network, but can additionally or alternatively be configured to select grasp points from the output of the object detector (e.g., an object detector can pre-process inputs to the grasp selector)." [0041]. "Planning the grasp can include determining a grasp pose, where the grasp is planned based on the grasp point and the grasp pose. In a first variant, the grasp pose can be determined from the object parameters output by an object detector (e.g., running in series and/or parallel with the graspability network/grasp selector, based on the same or a contemporaneously-captured image), and planning the grasp for the object parameters for the detected object that encompasses (e.g., includes, is associated with) the grasp point" [0099]. The system, as outlined, thus uses a grasp selector to select a grasp point and object to grasp before planning the grasp, inclusive of the position and posture of the robotic manipulator. The grasp selector selects grasp points on the basis of the graspability network, i.e. the feature map.) …, the position and the posture enabling the handling tool to grasp the grasping target object (“The computing system can include a motion planner 148, which functions to determine control instructions for the robotic arm to execute a grasp attempt for a selected grasp point” [0043]. Thus, the position and posture enable the handling tool to grasp the grasping target object by determining control instructions for a selected grasp point.);…
However, Humayun does not explicitly teach …detecting a circular anchor on the feature map;
calculating a position and a posture of a handling tool by a first parameter representing a polygon for which the circular anchor is a circumscribed circle, the position and the posture enabling the handling tool to grasp the grasping target object; and
calculating a polygonal grasp configuration (GC) of the handling tool by converting the position and the posture indicated by the first parameter into a second parameter indicating a position and a posture of the handling tool on the captured image,
wherein the circular anchor is a circumscribed circle of the polygonal GC,
the first parameter includes a first sub-parameter indicating coordinates of a center of the circumscribed circle and a second sub-parameter indicating a radius of the circumscribed circle, and
the calculating of the polygonal GC is executed by acquiring the second parameter from the first sub-parameter, the second sub-parameter, and a third sub-parameter indicating a midpoint of a side of the polygonal GC.
Maehara, pertinent to the problem at hand, teaches …detecting a circular anchor on the feature map (Maehara sets a “possibility of interference J” as shown in Fig. 6 “anchored” at the grasping center which the fingers symmetrically open and close around as designated in [0086].), and
calculating a position and posture of a handling tool by a first parameter representing an estimation of the grasping attitude for which the circular anchor is a circumscribed circle, the position and posture enabling the handling tool to grasp the grasping target object (See Fig. 3 S130 “set grasping target” which determines the grasping target based on grasping candidates and grasping attitude (see Paragraphs [0074-0079]). Viable grasps are selected based on the center of the graspable member, i.e., center of the circular anchor, as well as the non-interference attitude ranges, i.e., radius of the circular anchor (see Fig. 15-16). The center and radius are determined as the first parameter of the circular anchor which estimates the grasping attitude. Furthermore, we will consider the center of the estimated midpoint of the finger on one side of the grip configuration. Further, S150 performs grasping based on the position and posture of the fingers determined when selecting the grasping target.); …
calculating a …GC of the handling tool (The grasp configuration is determined with respect to the position of the approximated finger positions within the circular anchor in [0055].)…
wherein the circular anchor is a circumscribed circle of the … GC (As the “possibility of interference J” is the circular anchor, it can be seen that said circle is circumscribed around the estimated grip configuration of fingers 5 which are equally separated.),
the first parameter includes a first sub-parameter indicating coordinates of a center of the circumscribed circle (Center O which indicates the center of the grasp configuration between fingers.) and a second sub-parameter indicating a radius of the circumscribed circle (The radius of “possibility of interference J” which is shown in Fig. 6 as r(i,j)+Rf which is the radius of the circle of the graspable member and associated diameter of the estimated finger.)…
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to have modified the estimation of a grasp pose as taught by Humayun by fixing a circular anchor to the image as taught by Maehara with a reasonable expectation of success. One of ordinary skill in the art would have been motivated to make such a modification because the inclusion of the circular anchor (“the possibility of interference”) in determining the grip configuration for the fingers in grasping the graspable member allows for steady grasps regardless of any attitude recognition errors which result from the visual sensing or positioning errors of the grasping attitude (Maehara, [0082]).
However, given that Maehara merely teaches an estimation for individual fingers, rather than the entire span of the grip configuration, Humayun as modified by Maehara still does not teach …calculating a position and a posture of a handling tool by a first parameter representing a polygon for which the circular anchor is a circumscribed circle…
…calculating a polygon GC of the handling tool by converting the position and the posture indicated by the first parameter into a second parameter indicating a position and a posture of the handling tool on the captured image,
wherein the circular anchor is a circumscribed circle of the polygon GC, … and
the GC calculation unit calculates the polygon GC by acquiring the second parameter from the first sub-parameter, the second sub-parameter, and a third sub-parameter indicating a midpoint of a side of the polygon GC.
Redmon, pertinent to the problem at hand, teaches an estimation of a polygon GC as shown in Figs. 2-4. The specification describes, “The graphical representation 260 is a “grasping rectangle” that defines a grasping coordinate 262, a grasping width 264, a grasping height 266, and an orientation parameter 268 of a robot end effector for the grasp. The grasping coordinate 262 defines a two-dimensional (e.g., “X” and “y”) coordinate that defines a “center” or other reference point of a robot end effector. The grasping width 264 defines a distance between two or more opposing actuable members of the robot end effector, such as a distance between opposing plates of a parallel plate gripper end effector. The grasping height 266 defines a span of each of one or more opposing actuable members of a robot end effector, such as the span of each plate of a parallel plate gripper. It is noted that some robot end effectors may have fixed heights and other end effectors may have adjustable heights (e.g., plates of adjustable sizes). The orientation parameter 268 defines an orientation angle of the end effector for the grasp of the object relative to a reference orientation such as an “X” axis (left to right in FIG. 2) of the image 250. In some implementations, the orientation parameter 268 may comprise two separate parameters to account for the two-fold rotationally symmetric nature of grasp angles” (C5,L65 through C6,L20). Thus, provided that the “grasping rectangle” details similar grasping parameters which describe the grasping configuration of Fig. 6 in Maehara, one of ordinary skill in the art may be adequately motivated to simply substitute the estimation of the finger positions as detailed by Maehara (one known element) with the estimation of the “grasping rectangle” as detailed by Redmon (another known element) to obtain predictable results. This substitution would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (see MPEP 2143.I(B)). Furthermore, one of ordinary skill in the art may be motivated to make this modification because the “grasping rectangle” of Redmon would not only indicate the position of the opposing gripping plates of the end effector, i.e., gripper, but additionally the potential occupancy of the opposing gripping plates as they close in around the object to be grasped.
Reliant on the substitution detailed above, Examiner concludes that the combination of references teaches …detecting a circular anchor on the feature map (Newly depicted “possibility of interference J” as shown in Fig. 1 of supplementary document.), and
calculating a position and a posture of a handling tool by a first parameter representing a polygon for which the circular anchor is a circumscribed circle (The position and posture of the handling tool is represented as the “grasping rectangle” of Redmon, i.e., first parameter representing a polygon, which is circumscribed by the “possibility of interference J” of Maehara as shown in Fig. 1 in supplementary document.), the position and the posture enabling the handling tool to grasp the grasping target object (S150 of Maehara and additionally 812 of Fig. 8 of Redmon detail that the estimated grasp configuration which is the associated position and posture enable the handling tool to grasp the grasping target object.)…
wherein the circular anchor is a circumscribed circle of the polygon GC (Again, as shown in Fig. 1 of supplementary document, “the possibility of interference J”, i.e., circular anchor, is a circumscribed circle of the “grasping rectangle”, i.e., polygon GC.)…
Although not explicit, the following limitations are implied such that the modified references teach …calculating a polygon GC of the handling tool by converting the position and the posture indicated by the first parameter into a second parameter indicating a position and a posture of the handling tool on the captured image (Figs. 2-3 of the supplementary document show how the variables indicated by Applicant as the first parameter (center of a circumscribed circle, radius of a circumscribed circle, and midpoint of a side of the polygon GC) are obvious variants of the second parameter (height, width, angle, and center) of the “grasping rectangle” disclosed by Redmon through the use of basic geometry.),
…the first parameter includes a first sub-parameter indicating coordinates of a center of the circumscribed circle (Center O as shown in Fig. 1 of supplementary document and additionally coordinate (x,y) as shown in Figs. 2-3 of supplementary document.) and a second sub-parameter indicating a radius of the circumscribed circle (Radius r shown in Fig. 2 of supplementary document.), and
the GC calculation unit calculates the polygon GC by acquiring the second parameter from the first sub-parameter, the second sub-parameter, and a third sub-parameter indicating a midpoint of a side of the polygon GC (As can be seen in Figs. 2-3 of supplementary document the second parameter (grasp height, grasp width, grasp orientation, and grasp coordinate) are obvious variants of the first sub-parameter, i.e., center (x,y), second sub-parameter, i.e., radius r, and third sub-parameter, i.e., midpoint dRx and dRy.).
Therefore, although the references do not explicitly determine that the polygon GC of the handling tool converts the position and posture indicated by the first parameter into a second parameter, both Maehara and Redmon provide the necessary information which would allow one of ordinary skill in the art to arrive at the claimed invention. This conclusion is motivated as a combination of variables in that the set of variables as taught by Redmon may be included with the desired variables by known methods of basic geometry to yield predictable results (MPEP 2143.I(A)). Such a modification would be obvious to try as this change of variables is one of a finite number of identified, predictable solutions, with a reasonable expectation of success (MPEP 2143.I(E)). Additionally, one of ordinary skill in the art may be prompted to vary the variables which were used based on design incentives as such variations are predictable to one of ordinary skill in the art (MPEP 2143.I(F)).
Regarding claim 7, Humayun further teaches a computer program product comprising a non-transitory computer-readable recording medium on which an executable program is recorded ("Alternative embodiments implement the above methods and/or processing modules in non-transitory computer-readable media, storing computer-readable instructions that, when executed by a processing system, cause the processing system to perform the method(s) discussed herein" [0124].), the program instructing a computer to execute the handling method according to claim 6 (See rejection of claim 6 above which is taught by Humayun as modified by Maehara further modified by Redmon.).
Regarding claim 8, Humayun as modified by Maehara further modified by Redmon teaches the handling method according to claim 6,
with Humayun further teaching wherein the calculating the feature map is performed by receiving input of a plurality of pieces of image sensor information ("The input to the graspability network can include: an RGB image, receptive field from image, optionally depth, optionally object detector output (e.g., object parameters, etc.), and/or any other suitable information. In a first variant, the input to the graspability network is a 2D image having 3 channels per pixel (i.e., red-green-blue; RGB). In a second variant, the input to the graspability network can be a 2.5D image having 4 channels per pixel (RGB-depth image). In a first example, the depth can be a sensed depth (e.g., from a lower-accuracy sensor or a higher-accuracy sensor such a Lidar). In a second example, the depth can be a ‘refined’ depth determined by a trained depth enhancement network (e.g., wherein the depth enhancement network can be a precursor neural network or form the initial layers of the graspability network; etc.). In a third variant, the input to the graspability network can include an object detection output as an input feature (e.g., an object parameter, such a characteristic axis of a detected object)" [0069]. Thus, the feature map receives a plurality of input from a plurality of image sensor information including color, object detection, and depth perception images.),
extracting a plurality of intermediate features by a plurality of feature extractors from the plurality of pieces of image sensor information, integrating the plurality of intermediate features, and fusing, by convolution calculation, features of the plurality of pieces of image sensor information including the plurality of intermediate features ("The graspability network is preferably trained based on the labelled images. The labelled images can include: the image (e.g., RGB, RGB-D, RGB and point cloud, etc.), grasp point (e.g., the image features depicting a 3D physical point to grasp in the scene), and grasp outcome; and optionally the object parameters (e.g., object pose, surface normal, etc.), effector parameters (e.g., end effector pose, grasp pose, etc.), and/or other information. In particular, the graspability network is trained to predict the outcome of a grasp attempt at the grasp point, given the respective image as the input. However, the network can additionally or alternatively be trained based on object parameters and/or robotic manipulator parameters, such as may be used to: train the graspability network to predict the object parameter values (or bins) and/or robotic manipulator parameter values (or bins), given the respective image as input" [0075]. "The graspability network can be a neural network (e.g., CNN, fully connected, etc.), such as a convolutional neural network (CNN), fully convolutional neural network (FCN), artificial neural network (ANN), a feed forward network, a clustering algorithm, and/or any other suitable neural network or ML model" [0039]. The image sensor information is labelled such that it may be fed into a training model in order to delineate features for each of the images. The features are integrated such that they are provided to a convolutional neural network model in order to fuse the image data and determine graspability outcomes.).
Regarding claim 9, Humayun as modified by Maehara further modified by Redmon teaches the computer program product according to claim 7,
with Humayun further teaching wherein the calculating the feature map is performed by receiving input of a plurality of pieces of image sensor information ("The input to the graspability network can include: an RGB image, receptive field from image, optionally depth, optionally object detector output (e.g., object parameters, etc.), and/or any other suitable information. In a first variant, the input to the graspability network is a 2D image having 3 channels per pixel (i.e., red-green-blue; RGB). In a second variant, the input to the graspability network can be a 2.5D image having 4 channels per pixel (RGB-depth image). In a first example, the depth can be a sensed depth (e.g., from a lower-accuracy sensor or a higher-accuracy sensor such a Lidar). In a second example, the depth can be a ‘refined’ depth determined by a trained depth enhancement network (e.g., wherein the depth enhancement network can be a precursor neural network or form the initial layers of the graspability network; etc.). In a third variant, the input to the graspability network can include an object detection output as an input feature (e.g., an object parameter, such a characteristic axis of a detected object)" [0069]. Thus, the feature map receives a plurality of input from a plurality of image sensor information including color, object detection, and depth perception images.),
extracting a plurality of intermediate features by a plurality of feature extractors from the plurality of pieces of image sensor information, integrating the plurality of intermediate features, and fusing, by convolution calculation, features of the plurality of pieces of image sensor information including the plurality of intermediate features ("The graspability network is preferably trained based on the labelled images. The labelled images can include: the image (e.g., RGB, RGB-D, RGB and point cloud, etc.), grasp point (e.g., the image features depicting a 3D physical point to grasp in the scene), and grasp outcome; and optionally the object parameters (e.g., object pose, surface normal, etc.), effector parameters (e.g., end effector pose, grasp pose, etc.), and/or other information. In particular, the graspability network is trained to predict the outcome of a grasp attempt at the grasp point, given the respective image as the input. However, the network can additionally or alternatively be trained based on object parameters and/or robotic manipulator parameters, such as may be used to: train the graspability network to predict the object parameter values (or bins) and/or robotic manipulator parameter values (or bins), given the respective image as input" [0075]. "The graspability network can be a neural network (e.g., CNN, fully connected, etc.), such as a convolutional neural network (CNN), fully convolutional neural network (FCN), artificial neural network (ANN), a feed forward network, a clustering algorithm, and/or any other suitable neural network or ML model" [0039]. The image sensor information is labelled such that it may be fed into a training model in order to delineate features for each of the images. The features are integrated such that they are provided to a convolutional neural network model in order to fuse the image data and determine graspability outcomes.).
Regarding claim 10, Humayun as modified by Maehara further modified by Redmon teaches the apparatus according to claim 1,
with Maehara further modified by Redmon teaching wherein the second parameter is constituted by elements {x, y, w, h, θ}, the elements x and y indicating a center position of the handling tool (Center (x,y) of “grasping rectangle” shown in Figs. 2-3 of supplementary document corresponding to “grasping coordinate 262” of Redmon.), the element w indicating an opening width of the handling tool (Element w shown in Fig. 2 of supplementary document corresponding to “grasping width 264” of Redmon.), the element h indicating a width of a finger of the handling tool (Element h shown in Fig. 2 of supplementary document corresponding to “grasping height 266” of Redmon.), and the element θ indicating an angle between the element w and an image horizontal axis (Element θ shown in Fig. 3 of supplementary document corresponding to “orientation parameter 268” of Redmon.), and
the second parameter is acquired by a following equation with the first sub- parameter as cx and cy (Center (x,y) shown in Figs. 2-3 of supplementary document.), the second sub-parameter as R (Radius r shown in Fig. 2 of supplementary document.), and the third sub-parameter as dRx and dRy (Translations from center dRx and dRy shown in Fig. 3 of supplementary document.)
x
=
c
x
(Center (x,y) corresponds to both the center of the circular anchor and the center of the “grasping rectangle”.)
y
=
c
y
(Center (x,y) corresponds to both the center of the circular anchor and the center of the “grasping rectangle”.)
w
=
2
*
d
R
x
2
+
d
R
y
2
(As can be seen in Fig. 3 of the supplementary document, half of the width forms the hypoteneuse of a right triangle with dRx and dRy, thus rendering this equation a mere implementation of the Pythagorean theorem.)
h
=
2
*
R
2
-
(
d
R
x
2
+
d
R
y
2
)
(As shown in Fig. 2 of the supplementary document half of the height forms a side of a right triangle with hypoteneuse r and other side which is half the width. As shown in Fig. 3 of the supplementary document, half of the width is the hypoteneuse of a right triangle formed by dRx and dRy. Thus, this is another mere implementation of the Pythagorean theorem.)
θ
=
a
r
c
t
a
n
d
R
y
d
R
x
(As shown in Fig. 3 of the supplementary document, θ is the tangent corresponding to dRy (opposite edge) and dRx (adjacent edge) of the right triangle, thus rendering this equation a mere implementation of basic geometry.).
Regarding claim 11, Humayun as modified by Maehara further modified by Redmon teaches the apparatus according to claim 10,
with Maehara further modified by Redmon teaching wherein the third sub-parameter indicates coordinates of a center of the element h (As shown in Fig. 3 of the supplementary document, the third sub-parameter which is indicated by dRx and dRy are exemplary of coordinates of a center of the element h.).
Regarding claim 12, Humayun as modified by Maehara further modified by Redmon teaches the handling method according to claim 6,
with Maehara further modified by Redmon teaching wherein the second parameter is constituted by elements {x, y, w, h, θ}, the elements x and y indicating a center position of the handling tool (Center (x,y) of “grasping rectangle” shown in Figs. 2-3 of supplementary document corresponding to “grasping coordinate 262” of Redmon.), the element w indicating an opening width of the handling tool (Element w shown in Fig. 2 of supplementary document corresponding to “grasping width 264” of Redmon.), the element h indicating a width of a finger of the handling tool (Element h shown in Fig. 2 of supplementary document corresponding to “grasping height 266” of Redmon.), and the element θ indicating an angle between the element w and an image horizontal axis (Element θ shown in Fig. 3 of supplementary document corresponding to “orientation parameter 268” of Redmon.), and
the second parameter is acquired by a following equation with the first sub- parameter as cx and cy (Center (x,y) shown in Figs. 2-3 of supplementary document.), the second sub-parameter as R (Radius r shown in Fig. 2 of supplementary document.), and the third sub-parameter as dRx and dRy (Translations from center dRx and dRy shown in Fig. 3 of supplementary document.)
x
=
c
x
(Center (x,y) corresponds to both the center of the circular anchor and the center of the “grasping rectangle”.)
y
=
c
y
(Center (x,y) corresponds to both the center of the circular anchor and the center of the “grasping rectangle”.)
w
=
2
*
d
R
x
2
+
d
R
y
2
(As can be seen in Fig. 3 of the supplementary document, half of the width forms the hypoteneuse of a right triangle with dRx and dRy, thus rendering this equation a mere implementation of the Pythagorean theorem.)
h
=
2
*
R
2
-
(
d
R
x
2
+
d
R
y
2
)
(As shown in Fig. 2 of the supplementary document half of the height forms a side of a right triangle with hypoteneuse r and other side which is half the width. As shown in Fig. 3 of the supplementary document, half of the width is the hypoteneuse of a right triangle formed by dRx and dRy. Thus, this is another mere implementation of the Pythagorean theorem.)
θ
=
a
r
c
t
a
n
d
R
y
d
R
x
(As shown in Fig. 3 of the supplementary document, θ is the tangent corresponding to dRy (opposite edge) and dRx (adjacent edge) of the right triangle, thus rendering this equation a mere implementation of basic geometry.).
Regarding claim 13, Humayun as modified by Maehara further modified by Redmon teaches the handling method according to claim 12,
with Maehara further modified by Redmon teaching wherein the third sub-parameter indicates coordinates of a center of the element h (As shown in Fig. 3 of the supplementary document, the third sub-parameter which is indicated by dRx and dRy are exemplary of coordinates of a center of the element h.).
Regarding claim 14, Humayun as modified by Maehara further modified by Redmon teaches the computer program product according to claim 7,
with Maehara further modified by Redmon teaching wherein the second parameter is constituted by elements {x, y, w, h, θ}, the elements x and y indicating a center position of the handling tool (Center (x,y) of “grasping rectangle” shown in Figs. 2-3 of supplementary document corresponding to “grasping coordinate 262” of Redmon.), the element w indicating an opening width of the handling tool (Element w shown in Fig. 2 of supplementary document corresponding to “grasping width 264” of Redmon.), the element h indicating a width of a finger of the handling tool (Element h shown in Fig. 2 of supplementary document corresponding to “grasping height 266” of Redmon.), and the element θ indicating an angle between the element w and an image horizontal axis (Element θ shown in Fig. 3 of supplementary document corresponding to “orientation parameter 268” of Redmon.), and
the second parameter is acquired by a following equation with the first sub- parameter as cx and cy (Center (x,y) shown in Figs. 2-3 of supplementary document.), the second sub-parameter as R (Radius r shown in Fig. 2 of supplementary document.), and the third sub-parameter as dRx and dRy (Translations from center dRx and dRy shown in Fig. 3 of supplementary document.)
x
=
c
x
(Center (x,y) corresponds to both the center of the circular anchor and the center of the “grasping rectangle”.)
y
=
c
y
(Center (x,y) corresponds to both the center of the circular anchor and the center of the “grasping rectangle”.)
w
=
2
*
d
R
x
2
+
d
R
y
2
(As can be seen in Fig. 3 of the supplementary document, half of the width forms the hypoteneuse of a right triangle with dRx and dRy, thus rendering this equation a mere implementation of the Pythagorean theorem.)
h
=
2
*
R
2
-
(
d
R
x
2
+
d
R
y
2
)
(As shown in Fig. 2 of the supplementary document half of the height forms a side of a right triangle with hypoteneuse r and other side which is half the width. As shown in Fig. 3 of the supplementary document, half of the width is the hypoteneuse of a right triangle formed by dRx and dRy. Thus, this is another mere implementation of the Pythagorean theorem.)
θ
=
a
r
c
t
a
n
d
R
y
d
R
x
(As shown in Fig. 3 of the supplementary document, θ is the tangent corresponding to dRy (opposite edge) and dRx (adjacent edge) of the right triangle, thus rendering this equation a mere implementation of basic geometry.).
Regarding claim 15, Humayun as modified by Maehara further modified by Redmon teaches the computer program product according to claim 14,
with Maehara further modified by Redmon teaching wherein the third sub-parameter indicates coordinates of a center of the element h (As shown in Fig. 3 of the supplementary document, the third sub-parameter which is indicated by dRx and dRy are exemplary of coordinates of a center of the element h.).
Regarding claim 16, Humayun as modified by Maehara further modified by Redmon teaches the apparatus according to claim 1,
with Maehara further modified by Redmon teaching wherein the polygon is a quadrangle, and the polygonal GC is a quadrangular GC (The circumscribed “grasping rectangle” is a quadrangle forming a quadrangular GC.).
Regarding claim 17, Humayun as modified by Maehara further modified by Redmon teaches the apparatus according to claim 16,
with Maehara further modified by Redmon teaching wherein the quadrangle is a rectangle, and the quadrangular GC is a rectangular GC (The circumscribed “grasping rectangle” is a rectangle forming a rectangular GC.).
Regarding claim 18, Humayun as modified by Maehara further modified by Redmon teaches the handling method according to claim 6,
with Maehara further modified by Redmon teaching wherein the polygon is a quadrangle, and the polygonal GC is a quadrangular GC (The circumscribed “grasping rectangle” is a quadrangle forming a quadrangular GC.).
Regarding claim 19, Humayun as modified by Maehara further modified by Redmon teaches the handling method according to claim 18,
with Maehara further modified by Redmon teaching wherein the quadrangle is a rectangle, and the quadrangular GC is a rectangular GC (The circumscribed “grasping rectangle” is a rectangle forming a rectangular GC.).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY L MOLNAR whose telephone number is (571)272-2276. The examiner can normally be reached 8 A.M. to 3 P.M. EST Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jonathan (Wade) Miles can be reached at (571) 270-7777. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.L.M./Examiner, Art Unit 3656
/WADE MILES/Supervisory Patent Examiner, Art Unit 3656