Prosecution Insights
Last updated: October 02, 2026
Application No. 18/876,272

VISION SYSTEM AND VISION DETECTION METHOD

Non-Final OA §102§103§112
Filed
Dec 18, 2024
Priority
Jul 05, 2022 — nonprovisional of PCTJP2022026701
Examiner
CULLEN, TANNER L
Art Unit
3656
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
FANUC Corporation
OA Round
1 (Non-Final)
72%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
125 granted / 174 resolved
+19.8% vs TC avg
Strong +16% interview lift
Without
With
+16.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
23 currently pending
Career history
212
Total Applications
across all art units

Statute-Specific Performance

§101
9.1%
-30.9% vs TC avg
§103
57.2%
+17.2% vs TC avg
§102
18.0%
-22.0% vs TC avg
§112
12.6%
-27.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 174 resolved cases

Office Action

§102 §103 §112
DETAILED CORRESPONDENCE This is the first office action regarding application number 18/876,272, filed on 18 December 2024. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: a. “acquisition unit” in claims 1, 3 and 10 b. “detection unit” in claims 1-3 and 7-8 c. “force measurement unit” in claim 4 d. “display unit” in claims 9-10 Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. The specification discloses the corresponding structure for a “force measurement unit” in Figure 8 and paragraph [0038] and for a “display unit” in Figure 1 and paragraph [0012] in the specification filed on 18 December 2024. Regarding the limitations reciting the “acquisition unit” and “detection unit”, the specification discloses a computer in Figure 2 and paragraph [0019] and an algorithm for performing the claimed functions in Figure 7 and its corresponding paragraphs. If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3 and 11 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chitta et al. (US 20170246744 A1 and Chitta hereinafter). Regarding Claim 1 Chitta teaches a vision system (see all Figs.; [0004]), comprising: an acquisition unit configured to acquire a first type of image in which a presence and surrounding region of a plurality of workpieces is captured (see Fig. 4, all; [0004 "According to a particular class of implementations, a first image of a loaded pallet is captured. The loaded pallet has a plurality of boxes stacked thereon."], [0010], [0024 "According to various implementations, the system uses computer vision and images (both 2D and/or 3D) from one or more visual sensors (e.g., one or more cameras) to determine the outermost corners of the top layer of a pallet of boxes."] and [0038]); a detection unit configured to output a result of detecting the plurality of workpieces (see Fig. 4, all; [0004 "According to a particular class of implementations, a first image of a loaded pallet is captured. The loaded pallet has a plurality of boxes stacked thereon. A first representation of a surface of the loaded pallet is generated using the first image."], [0024 "According to various implementations, the system uses computer vision and images (both 2D and/or 3D) from one or more visual sensors (e.g., one or more cameras) to determine the outermost corners of the top layer of a pallet of boxes."] and [0038]); and a robot configured to execute an operation to change at least one of positions of the plurality of workpieces (see Figs. 1 and 5, robot 28; [0004 "A first box at a corner of the surface of the loaded pallet is moved a programmable amount using a robotic arm."], [0024 "It then performs an exploratory pick using a robot and an attached gripper, lifting the outermost box from its corner by a small amount to separate it from the top layer."], [0027] and [0039]), wherein the robot executes a first operation to change the positions of the plurality of workpieces (see Fig. 5, all; [0004 "A first box at a corner of the surface of the loaded pallet is moved a programmable amount using a robotic arm."], [0012 "According to a particular implementation, all of the boxes are unloaded from the loaded pallet while performing the capturing, generating, lifting, capturing, generating, releasing, and processing for fewer than all of the boxes."]-[0013 "According to a particular implementation, moving the first box includes one of (1) lifting the first box the programmable amount relative to the loaded pallet using the robotic arm, (2) moving the first box laterally relative to the loaded pallet using the robotic arm, or (3) tilting the first box relative to the surface of the loaded pallet using the robotic arm."], [0024 "It then performs an exploratory pick using a robot and an attached gripper, lifting the outermost box from its corner by a small amount to separate it from the top layer."] and [0039]-[0040]), the acquisition unit acquires a second type of image in which the presence and surrounding region of the plurality of workpieces is captured after executing the first operation (see Fig. 6, all; [0004 "A second image of the loaded pallet is captured including the first box as moved by the robotic arm."], [0024 "The system takes a new image of the pallet and the boxes and using computer vision again…"] and [0041]), and the detection unit executes detection on the first type of image and the second type of image (see Figs. 8-9, all; [0004 "A second representation of the surface of the loaded pallet is generated using the second image. The first box is replaced with the robotic arm. The first and second representations of the surface of the loaded pallet are processed to generate a representation of a surface of the first box."], [0024 "The system takes a new image of the pallet and the boxes and using computer vision again, and computes a difference between the two images to determine the size of the picked box and its position in the layer."] and [0041]-[0043]), and outputs at least one information of a position, posture, or external shape information of the plurality of workpieces as the detection result, based on at least the information of an amount of change calculated between first type of image data which containing a detection result and second type of image data which containing a detection result (see Figs. 8-10, all; [0004 "The first and second representations of the surface of the loaded pallet are processed to generate a representation of a surface of the first box."]-[0005 "According to a particular implementation, processing of the first and second representations of the surface of the loaded pallet includes determining a difference between the first and second representations of the surface of the loaded pallet, and determining a size and a position of the surface of the first box using the difference."], [0012]-[0014 "This includes determining a location and orientation of each of the instances of the first box type. Each of the instances of the first box type is picked using the corresponding location and orientation."]-[0015], [0024 "The system takes a new image of the pallet and the boxes and using computer vision again, and computes a difference between the two images to determine the size of the picked box and its position in the layer."]-[0025] and [0042]-[0044]). Regarding Claim 2 Chitta teaches the vision system according to claim 1 (as discussed above in claim 1), wherein the detection unit outputs the detection result of the plurality of workpieces, based on a difference calculated by using at least one information of the position, posture, or external shape information of the plurality of workpieces included in the first type of image data and the second type of image data (see Figs. 8-10, all; [0004]-[0005 "According to a particular implementation, processing of the first and second representations of the surface of the loaded pallet includes determining a difference between the first and second representations of the surface of the loaded pallet, and determining a size and a position of the surface of the first box using the difference."], [0012]-[0015], [0024 "The system takes a new image of the pallet and the boxes and using computer vision again, and computes a difference between the two images to determine the size of the picked box and its position in the layer."]-[0025] and [0042 "Referring to FIG. 8, the Difference 54 between Image 38 and Image 50, corresponds to the Robot 28, the Gripper 30 and the Box 12 that the system has picked up."]-[0044 "Referring to the flowchart in FIG. 3 and the top view of the picked box in FIG. 10, the system also determines the Center 56 of the Box 12 using this Difference 54."]). Regarding Claim 3 Chitta teaches the vision system according to claim 1 (as discussed above in claim 1), wherein the robot executes a second operation, different from the first operation, to change the positions of the plurality of workpieces (see [0004 "The first and second representations of the surface of the loaded pallet are processed to generate a representation of a surface of the first box. The robotic arm is positioned relative to the first box using the representation of the surface of the first box. The first box is removed from the loaded pallet using the robotic arm."], [0008], [0012 "According to a particular implementation, all of the boxes are unloaded from the loaded pallet while performing the capturing, generating, lifting, capturing, generating, releasing, and processing for fewer than all of the boxes."] and [0024 "It puts the box down and, using the computed size and position, picks it up again at its center to complete the task.']), the acquisition unit acquires a third type of image in which the plurality of workpieces is captured after executing the second operation (see [0004 "The first and second representations of the surface of the loaded pallet are processed to generate a representation of a surface of the first box. The robotic arm is positioned relative to the first box using the representation of the surface of the first box. The first box is removed from the loaded pallet using the robotic arm."], [0014 "According to a particular implementation, the representation of the surface of the first box is stored. A third image of the loaded pallet is captured using the image capture device. The third image does not include the first box. A third representation of the surface of the loaded pallet is generated using the third image.']), and the detection unit executes detection on the third type of image and outputs the detection result, based on at least the information of an amount of change calculated between any two of the first type of image data, the second type of image data, and the third type of image data which containing a detection result (see [0005 "According to a particular implementation, processing of the first and second representations of the surface of the loaded pallet includes determining a difference between the first and second representations of the surface of the loaded pallet, and determining a size and a position of the surface of the first box using the difference."], [0014 "A third representation of the surface of the loaded pallet is generated using the third image. One or more instances of a first box type corresponding to the first box is/are identified by comparing the third representation of the surface of the loaded pallet and the stored representation of the surface of the first box. This includes determining a location and orientation of each of the instances of the first box type. Each of the instances of the first box type is picked using the corresponding location and orientation."]-[0015] and [0024]). Regarding Claim 11 Chitta teaches a vision detection method (see all Figs.; [0004]), comprising: an acquiring step of acquiring a first type of image in which a presence and surrounding region of a plurality of workpieces is captured (see Fig. 4, all; [0004 "According to a particular class of implementations, a first image of a loaded pallet is captured. The loaded pallet has a plurality of boxes stacked thereon."], [0010], [0024 "According to various implementations, the system uses computer vision and images (both 2D and/or 3D) from one or more visual sensors (e.g., one or more cameras) to determine the outermost corners of the top layer of a pallet of boxes."] and [0038]); a detecting step of outputting a result of detecting the plurality of workpieces (see Fig. 4, all; [0004 "According to a particular class of implementations, a first image of a loaded pallet is captured. The loaded pallet has a plurality of boxes stacked thereon. A first representation of a surface of the loaded pallet is generated using the first image."], [0024 "According to various implementations, the system uses computer vision and images (both 2D and/or 3D) from one or more visual sensors (e.g., one or more cameras) to determine the outermost corners of the top layer of a pallet of boxes."] and [0038]); and an executing step of causing a robot to execute an operation to change at least one of positions of the plurality of workpieces (see Figs. 1 and 5, robot 28; [0004 "A first box at a corner of the surface of the loaded pallet is moved a programmable amount using a robotic arm."], [0024 "It then performs an exploratory pick using a robot and an attached gripper, lifting the outermost box from its corner by a small amount to separate it from the top layer."], [0027] and [0039]), wherein in the executing step, the robot executes a first operation to change the positions of the plurality of workpieces (see Fig. 5, all; [0004 "A first box at a corner of the surface of the loaded pallet is moved a programmable amount using a robotic arm."], [0012 "According to a particular implementation, all of the boxes are unloaded from the loaded pallet while performing the capturing, generating, lifting, capturing, generating, releasing, and processing for fewer than all of the boxes."]-[0013 "According to a particular implementation, moving the first box includes one of (1) lifting the first box the programmable amount relative to the loaded pallet using the robotic arm, (2) moving the first box laterally relative to the loaded pallet using the robotic arm, or (3) tilting the first box relative to the surface of the loaded pallet using the robotic arm."], [0024 "It then performs an exploratory pick using a robot and an attached gripper, lifting the outermost box from its corner by a small amount to separate it from the top layer."] and [0039]-[0040]), in the acquiring step, a second type of image is acquired, in which a presence and surrounding region of the plurality of workpieces after executing the first operation is captured (see Fig. 6, all; [0004 "A second image of the loaded pallet is captured including the first box as moved by the robotic arm."], [0024 "The system takes a new image of the pallet and the boxes and using computer vision again…"] and [0041]), and in the detecting step, detection on the first type of image and the second type of image is executed (see Figs. 8-9, all; [0004 "A second representation of the surface of the loaded pallet is generated using the second image. The first box is replaced with the robotic arm. The first and second representations of the surface of the loaded pallet are processed to generate a representation of a surface of the first box."], [0024 "The system takes a new image of the pallet and the boxes and using computer vision again, and computes a difference between the two images to determine the size of the picked box and its position in the layer."] and [0041]-[0043]), and at least one information of a position, posture, or external shape information of the plurality of workpieces is output as the detection result, based on at least the information of an amount of change calculated between the first type of image data which containing the detection result and the second type of image data which containing the detection result (see Figs. 8-10, all; [0004 "The first and second representations of the surface of the loaded pallet are processed to generate a representation of a surface of the first box."]-[0005 "According to a particular implementation, processing of the first and second representations of the surface of the loaded pallet includes determining a difference between the first and second representations of the surface of the loaded pallet, and determining a size and a position of the surface of the first box using the difference."], [0012]-[0014 "This includes determining a location and orientation of each of the instances of the first box type. Each of the instances of the first box type is picked using the corresponding location and orientation."]-[0015], [0024 "The system takes a new image of the pallet and the boxes and using computer vision again, and computes a difference between the two images to determine the size of the picked box and its position in the layer."]-[0025] and [0042]-[0044]). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 4-8 are rejected under 35 U.S.C. 103 as being unpatentable over Chitta as applied to claim 1 above, and further in view of Umetsu (US 20120004774 A1 and Umetsu hereinafter). Regarding Claim 4 Chitta teaches the vision system according to claim 1 (as discussed above in claim 1), the robot executes an operation to contact a surface of the plurality of workpieces, based on at least one result of the detection on the first type of image and/or the second type of image (see Fig. 5, all; [0004 "A first box at a corner of the surface of the loaded pallet is moved a programmable amount using a robotic arm."], [0012 "According to a particular implementation, all of the boxes are unloaded from the loaded pallet while performing the capturing, generating, lifting, capturing, generating, releasing, and processing for fewer than all of the boxes."]-[0013 "According to a particular implementation, moving the first box includes one of (1) lifting the first box the programmable amount relative to the loaded pallet using the robotic arm, (2) moving the first box laterally relative to the loaded pallet using the robotic arm, or (3) tilting the first box relative to the surface of the loaded pallet using the robotic arm."], [0024 "It then performs an exploratory pick using a robot and an attached gripper, lifting the outermost box from its corner by a small amount to separate it from the top layer."] and [0039]-[0040]). Chitta is silent regarding wherein the robot includes a force measurement unit configured to measure contact information when contacting the plurality of workpieces, and the force measurement unit measures the contact information on the surface. Umetsu teaches a vision system (see all Figs.; [0007]), comprising: an acquisition unit configured to acquire a first type of image in which a presence and surrounding region of a workpiece is captured (see Fig. 2, step S1; [0007 "The information acquiring unit acquires at least location information on a target by an input from a user or by detection made by a detection unit."], [0029] and [0036 "As illustrated in FIG. 2, in step S1, the image processor 43 of the control device 40 acquires the location information, attitude information, and shape information on the gripping target 110 on the basis of data of an image obtained by the visual sensor 30."]); a detection unit configured to output a result of detecting the workpiece (see Fig. 2, steps S1-S2; [0007], [0029 "The image processor 43 can acquire location information, attitude (orientation) information, shape information on the gripping target 110 by capturing data of an image obtained by the visual sensor 30 and processing the image data."] and [0036 "As illustrated in FIG. 2, in step S1, the image processor 43 of the control device 40 acquires the location information, attitude information, and shape information on the gripping target 110 on the basis of data of an image obtained by the visual sensor 30. Next, in step S2, the acquired shape information on the gripping target 110 is compared against data stored in the target database 44 of the control device 40, thus acquiring the dimensional information on the gripping target 110 and gripping manner data (gripping location and gripping attitude) on the gripping target 110 corresponding to the shape information or other data."]); and a robot configured to execute an operation to change at least one of positions of the workpiece (see Fig. 1, robot arm 10; [0007], [0019 "As illustrated in FIG. 1, the robot apparatus 100 according to the first embodiment has the function of gripping and lifting an object (gripping target 110) with a multi-fingered hand 20 attached to end of a robot arm 10 in response to an instruction from a user and moving the object to a designated location."]-[0020], [0044] and [0047]), wherein the robot executes a first operation to change the position of the workpiece (see Fig. 2, step S8; [0044 "After that, in step S8, in a state where the multi-fingered hand 20 is arranged in the modified gripping location, the fingers 21, 22, and 23 are moved by the hand controller 42 and the multi-fingered hand 20 takes the modified gripping attitude, and the gripping target 110 is thus gripped."] and [0047 "For the first embodiment, as described above, the control device 40 modifies the gripping location of the multi-fingered hand 20 and controls gripping the gripping target 110 by the multi-fingered hand 20 on the basis of the location information modified using the contact location."]); wherein the robot includes a force measurement unit configured to measure contact information when contacting the workpiece (see [0007], [0022 "The three fingers 21, 22, and 23 have substantially the same configuration and include force sensors 21 a, 22 a, and 23 a, respectively, at their ends (fingertips)."]-[0023 "Each of the force sensors 21 a, 22 a, and 23 a is a pressure sensor for detecting forces in directions of mutually perpendicular three axes. The force sensors 21 a, 22 a, and 23 a detect external forces (loads) applied to the tips of the fingers 21, 22, and 23, respectively. A detection signal of each of the force sensors 21 a, 22 a, and 23 a is output to the control device 40. For the first embodiment, contact between the multi-fingered hand 20 ( fingers 21, 22, and 23) and an object (gripping target 110 and another peripheral object, such as a wall or table) can be detected on the basis of an output of each of the force sensors 21 a, 22 a, and 23 a."] and [0040]-[0041]), the robot executes an operation to contact a surface of the workpiece, based on at least one result of the detection on the first type of image (see Fig. 2, steps S5-S6; [0007 "The controller moves the robot arm to cause the multi-fingered hand to approach the target on the basis of the at least location information on the target acquired by the information acquiring unit, detects a contact location of actual contact with the target on the basis of an output of the at least one fingertip force sensor of the multi-fingered hand"], [0023], [0040 "...such that the multi-fingered hand 20 approaches the gripping attitude from the approaching attitude in the gripping location, each of the fingers 21, 22, and 23 of the multi-fingered hand 20 comes into contact with the gripping target 110. In response to this, the force sensors 21 a, 22 a, and 23 a at the ends of the fingers 21, 22, and 23 detect the contact with the gripping target 110.]-[0041]), and the force measurement unit measures the contact information on the surface (see Fig. 2, step S6; [0007], [0023 "For the first embodiment, contact between the multi-fingered hand 20 ( fingers 21, 22, and 23) and an object (gripping target 110 and another peripheral object, such as a wall or table) can be detected on the basis of an output of each of the force sensors 21 a, 22 a, and 23 a."] and [0040 "In response to this, the force sensors 21 a, 22 a, and 23 a at the ends of the fingers 21, 22, and 23 detect the contact with the gripping target 110."]-[0041 "In step S6, the control device 40 (hand controller 42) acquires the location of the contact with the gripping target 110 detected by each of the force sensors 21 a, 22 a, and 23 a (contact location)."]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify the vision system of Chitta to further include a force measurement unit configured to measure contact information on the surfaces of the workpieces when contacting the plurality of workpieces, as taught by Umetsu, in order to correct position and orientation information of the workpieces determined from images and to consequently modify a gripping position and orientation for the robot. Regarding Claim 5 Modified Chitta teaches the vision system according to claim 4 (as discussed above in claim 4), Chitta further teaches wherein the robot executes at least the first operation on the plurality of workpieces, based on the contact information (see Fig. 5, all; [0012 "According to a particular implementation, all of the boxes are unloaded from the loaded pallet while performing the capturing, generating, lifting, capturing, generating, releasing, and processing for fewer than all of the boxes."]-[0013 "According to a particular implementation, moving the first box includes one of (1) lifting the first box the programmable amount relative to the loaded pallet using the robotic arm, (2) moving the first box laterally relative to the loaded pallet using the robotic arm, or (3) tilting the first box relative to the surface of the loaded pallet using the robotic arm."], [0024 "It then performs an exploratory pick using a robot and an attached gripper, lifting the outermost box from its corner by a small amount to separate it from the top layer."] and [0039]-[0040]). Umetsu additionally teaches wherein the robot executes at least the first operation on the plurality of workpieces, based on the contact information (see Fig. 2, steps S7-S8; [0044 "After that, in step S8, in a state where the multi-fingered hand 20 is arranged in the modified gripping location, the fingers 21, 22, and 23 are moved by the hand controller 42 and the multi-fingered hand 20 takes the modified gripping attitude, and the gripping target 110 is thus gripped."] and [0047 "For the first embodiment, as described above, the control device 40 modifies the gripping location of the multi-fingered hand 20 and controls gripping the gripping target 110 by the multi-fingered hand 20 on the basis of the location information modified using the contact location."]). Regarding Claim 6 Modified Chitta teaches the vision system according to claim 4 (as discussed above in claim 4), Chitta is silent regarding wherein the contact information includes at least one of information on presence or absence of contact, information on a contact position, information on magnitude, direction, and distribution of a contact force, information on magnitude and direction of a contact moment, or information on slippage. Umetsu teaches wherein the contact information includes at least one of information on presence or absence of contact, information on a contact position, and information on magnitude, direction, and distribution of a contact force (see [0007], [0033 "That is, the control device 40 moves the robot arm 10 to cause the multi-fingered hand 20 to approach the gripping target 110 on the basis of the location information and attitude information on the gripping target 110 acquired by the visual sensor 30 and detects a contact location of actual contact with the gripping target 110 on the basis of the output of each of the force sensors 21 a, 22 a, and 23 a of the multi-fingered hand 20. Specifically, a contact location when the contact with the gripping target 110 is detected by the force sensors (21 a, 22 a, and 23 a) is calculated from a joint angle of the robot arm 10 and a joint angle of the multi-fingered hand 20. The control device 40 has the location correcting function of modifying the location information and attitude information on the gripping target 110 on the basis of information that indicates the detected contact location."], [0041] and [0043 "To address this, in step S7, the location information and attitude information on the gripping target 110 is modified on the basis of information indicating the detected three contact locations."]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify the vision system of Chitta to further include a force measurement unit configured to measure contact information on presence or absence of contact, information on a contact position and information on magnitude, direction, and distribution of a contact force, as taught by Umetsu, in order to correct position and orientation information of the workpieces determined from images and to consequently modify a gripping position and orientation for the robot. Regarding Claim 7 Modified Chitta teaches the vision system according claim 4 (as discussed above in claim 4), Chitta is silent regarding wherein the detection unit compensates the detection result, based on at least the information on a contact position among the contact information. Umetsu teaches wherein the detection unit compensates the detection result, based on at least the information on a contact position among the contact information (see Fig. 2, step S7; [0007 "...and modifies the location information on the target on the basis of information indicating the detected contact location."], [0042]-[0043 "To address this, in step S7, the location information and attitude information on the gripping target 110 is modified on the basis of information indicating the detected three contact locations. With this, the location of the robot arm 10 (i.e., the gripping location of the multi-fingered hand 20) is modified by the arm controller 41 and the gripping attitude of the multi-fingered hand 20 ( fingers 21, 22, and 23) is modified by the hand controller 42, on the basis of the modified location information and attitude information on the gripping target 110."] and [0046]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify the vision system of Chitta to compensate the detection result based on at least the information on a contact position with the workpieces, as taught by Umetsu, in order to correct position and orientation information of the workpieces determined from images and to consequently modify a gripping position and orientation for the robot. Regarding Claim 8 Modified Chitta teaches the vision system according to claim 6 (as discussed above in claim 6), Chitta is silent regarding wherein the detection unit compensates at least one of the first type of image, the second type of image, or the detection result, based on the information on an amount of change calculated with magnitude, direction or distribution of the contact force. Umetsu teaches wherein the detection unit compensates at least one of the first type of image or the detection result, based on the information on an amount of change calculated with magnitude, direction or distribution of the contact force (see Fig. 2, step S7; [0007 "...and modifies the location information on the target on the basis of information indicating the detected contact location."], [0042]-[0043 "To address this, in step S7, the location information and attitude information on the gripping target 110 is modified on the basis of information indicating the detected three contact locations. With this, the location of the robot arm 10 (i.e., the gripping location of the multi-fingered hand 20) is modified by the arm controller 41 and the gripping attitude of the multi-fingered hand 20 ( fingers 21, 22, and 23) is modified by the hand controller 42, on the basis of the modified location information and attitude information on the gripping target 110."] and [0046]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify the vision system of Chitta to compensate the first type of image or the detection result, based on the information on an amount of change calculated with magnitude, direction or distribution of the contact force, as taught by Umetsu, in order to correct position and orientation information of the workpieces determined from images and to consequently modify a gripping position and orientation for the robot. Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Chitta as applied to claim 1 above, and further in view of Diankov et al. (US 20200139553 A1 and Diankov hereinafter). Regarding Claim 9 Chitta teaches the vision system according to claim 1 (as discussed above in claim 1), further comprises a display unit (see [0035 "A rendering of the virtual pallet is then shown to the operator on the screen superimposed on the image of the actual pallet to see how well they are aligned."]). Chitta is silent regarding the display unit configured to display the detection result. Diankov teaches a vision system (see all Figs., especially Fig. 5), comprising: an acquisition unit configured to acquire a first type of image in which a presence and surrounding region of a plurality of workpieces is captured (see Fig. 5, step 202; [0076 "More specifically as an example, the capture module 202 can operate the image system 160 and/or interact with the image system 160 to capture and/or receive from the image system 160 the imaging data corresponding to a top view of a stack as illustrated in FIG. 3B and/or FIG. 4B."]); a detection unit configured to output a result of detecting the plurality of workpieces (see Fig. 5, steps 202-204; [0076 "The capture module 202 can process the resulting imaging data to capture the top surface of the package 112 of FIG. 1A as the first image data. In some embodiments, the capture module 202 can employ one or more image recognition algorithms (e.g., the contextual image classification, pattern recognition, and/or edge detection) to analyze the imaging data and identify edges and/or surfaces of the packages 112 therein. Based on the processing results (e.g., the identified edges and/or continuous surfaces), the capture module 202 can identify portions (e.g., sets of pixels values and/or depth readings) of the imaging data as representing top surfaces of individual packages."]-[0082]); and a robot configured to execute an operation to change at least one of positions of the plurality of workpieces (see Fig. 5, step 206; [0090 "The lift module 206 implements (via, e.g., communicating and/or executing) the command for the robotic arm 132 of FIG. 1A to lift the package 112. For example, the lift module 206 can operate the robotic arm 132 to lift the unregistered instance of the package 112 for the lift check distance as discussed above."]), wherein the robot executes a first operation to change the positions of the plurality of workpieces (see Fig. 5, step 206; [0090 "The lift module 206 implements (via, e.g., communicating and/or executing) the command for the robotic arm 132 of FIG. 1A to lift the package 112. For example, the lift module 206 can operate the robotic arm 132 to lift the unregistered instance of the package 112 for the lift check distance as discussed above."]-[0092]), the acquisition unit acquires a second type of image in which the presence and surrounding region of the plurality of workpieces is captured after executing the first operation (see Fig. 5, step 206; [0093 "More specifically as an example, the capture module 202 can capture the second image data based on the package 112 being lifted for the lift check distance to include the now visible two clear edges 112-44 g of FIG. 4H."]), and the detection unit executes detection on the first type of image and the second type of image (see Fig. 5, step 208; [0094 "For example, the extraction module 208 can extract the third image data based on the first image data, the second image data, or a combination thereof."]), and outputs at least one information of a position, posture, or external shape information of the plurality of workpieces as the detection result, based on at least the information of an amount of change calculated between first type of image data which containing a detection result and second type of image data which containing a detection result (see Fig. 5, step 208; [0065], [0095 "The extraction module 208 can extract the third image data in a number of ways. For example, the extraction module 208 can determine an image difference based comparing the first image data and the second image data. The image difference can represent the difference between the first image data and the second image data of the same instance of the unregistered instance of the package 112."]-[0098 "For a further example, the extraction module 208 can determine the length, width, or a combination thereof of the package 112 based on the third image data."] and [0106]); further comprises a display unit configured to display the detection result (see Figs. 3C, 4C, 4E, 4G, 4I, and 4L, all; [0040 "The image system 160 can include at least one display unit 164 configured to present an image of the package(s) 112 captured by the sensors 162 that may be viewed by one or more operators of the robotic system 100 as discussed in detail below."], [0051] and [0056 "As stated above, SIs for the packages 112 may be captured by the image system 160 of FIG. 1B and compared with the SIs of the registration records 172 of FIG. 1B stored in the RDS 170 of FIG. 1B. If registered, the robotic system can assign and/or display symbologies indicative of the SI being registered. As illustrated, symbologies 112-33 a, 112-43 a, and 112-51 a, can include rectangular outlines displayed by the display unit 164."]-[0069]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify the vision system of Chitta to further include a display unit configured to display the detection result, as taught by Diankov, in order to provide an operator with a view of the workpieces. Regarding Claim 10 Modified Chitta teaches the vision system according to claim 9 (as discussed above in claim 9), Chitta is silent regarding wherein the display unit displays the first type of image or the second type of image acquired by the acquisition unit, together with the detection result. Diankov teaches wherein the display unit displays the first type of image or the second type of image acquired by the acquisition unit, together with the detection result (see Figs. 3C, 4C, 4E, 4G, 4I, and 4L, all; [0040 "The image system 160 can include at least one display unit 164 configured to present an image of the package(s) 112 captured by the sensors 162 that may be viewed by one or more operators of the robotic system 100 as discussed in detail below."], [0051] and [0056 "As stated above, SIs for the packages 112 may be captured by the image system 160 of FIG. 1B and compared with the SIs of the registration records 172 of FIG. 1B stored in the RDS 170 of FIG. 1B. If registered, the robotic system can assign and/or display symbologies indicative of the SI being registered. As illustrated, symbologies 112-33 a, 112-43 a, and 112-51 a, can include rectangular outlines displayed by the display unit 164."]-[0069]). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify the vision system of Chitta to further include a display unit configured to display the first type of image or the second type of image acquired by the acquisition unit, together with the detection result, as taught by Diankov, in order to provide an operator with a view of the workpieces. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to TANNER LUKE CULLEN whose telephone number is (303)297-4384. The examiner can normally be reached Monday-Friday 9:00-5:00 MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Khoi Tran can be reached at (571) 272-6919. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TANNER L CULLEN/Examiner, Art Unit 3656 /KHOI H TRAN/Supervisory Patent Examiner, Art Unit 3656
Read full office action

Prosecution Timeline

Dec 18, 2024
Application Filed
Jul 16, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12734691
ROBOTIC MANIPULATOR HAVING A TASK NULL SPACE
2y 1m to grant Granted Sep 15, 2026
Patent 12735061
Driver Assistance System and Driver Assistance Method for a Vehicle
2y 0m to grant Granted Sep 15, 2026
Patent 12722293
SMART ROBOT TEACHING SYSTEM AND SMART ROBOT TEACHING METHOD
2y 4m to grant Granted Sep 01, 2026
Patent 12714522
STRUCTURAL ADJUSTMENT SYSTEMS AND METHODS FOR A TELEOPERATIONAL MEDICAL SYSTEM
2y 5m to grant Granted Aug 25, 2026
Patent 12709009
AUTONOMOUS VEHICLE SIMULATION USING MACHINE LEARNING
7y 3m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
72%
Grant Probability
88%
With Interview (+16.2%)
2y 12m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 174 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month