DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed July 14, 2026 have been fully considered but are not persuasive.
Applicant argues that the combination does not derive the image foreground seed from a point-cloud detection of the object's upper surface. Sharma does. Sharma removes the support surface from the point cloud and uses the base-removed point cloud “as a mask to
detect the object 104 in the image” (Sharma ¶17), which labels the image region corresponding
to the point-cloud-detected upper surface. Consistent with the specification, the claimed foreground region is a coarse seed that “need not accurately reflect the boundaries of the upper surface” and instead serves “as inputs to a segmentation algorithm” (Spec. ¶61). Sharma’s mask is such a region. Lee supplies the seeded foreground segmentation on that region. Lee projects a depth-extracted foreground region onto the color image and treats it as a graph-cut seed, “project[ing] them onto the color image and treat[ing] them as segmentation seeds,” with the “segmentation result in the color image … obtained by minimizing equation (2)” (Lee p. 3, FIG. 1(c)).
Applicant’s characterization of Lee as directed to only a human body skeleton mischaracterizes the references and is directed to a feature not relied upon in the rejection. The claimed foreground seed is mapped to Sharma’s point cloud mask. Lee is only cited for projecting a depth derived foreground region into the image and using it as a graph-cut seed (Lee p. 3, FIG. 1(c)), not for its skeleton, which is a separate refinement Lee adds because the projected region alone “is not accurate enough” (Lee p. 4). An argument against a teaching the rejection does not apply does not establish nonobviousness, and Lee cannot be considered in isolation from Sharma. In re Keller, 642 F.2d 413 (CCPA 1981).
Applicant's contention that there is no reason to combine is also not persuasive. Sharma and Lee both segment an object from a color image using depth-sensor data, and the rationale of record, improving segmentation accuracy at low-contrast object boundaries, is a proper rationale under KSR that does not rely on hindsight. The fact that Sharma performs contour analysis does not negate the benefit of Lee's more accurate seeded segmentation.
Applicant's traversal of the Official Notice on claims 5 and 15 is acknowledged, and documentary evidence is now provided in the rejection of claims 5 and 15 above. Poelman teaches assigning point-cloud points to a different segment where "the angle between their surface normals is smaller than some certain threshold" (Poelman col. 8), which is detecting a surface whose normal differs from the reference by at least a threshold.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3 and 11- 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma et al., "Three Dimensional Object Recognition," US20170308736A1, (hereinafter "Sharma") in view of Lee et al., "Graph Cut Based Human Body Segmentation in Color Images Using Skeleton Information from the Depth Sensor," (hereinafter "Lee"), and further in view of McCloskey et al. , "Dark Parcel Dimensioning," US20210209784A1, (hereinafter "McCloskey").
Claims 1 and 11.
Sharma and Lee disclose a method in a computing device (claim 1) and a computing device
(claim 11), the method comprising:
capturing, via a depth sensor, (i) a point cloud depicting, and (ii) a two-dimensional image depicting the object and the support surface (Sharma: "[T]he module 202 includes one or more depth sensors and one or more color sensors. A depth sensor is a visual sensor used to capture depth data of the object (¶11) ... A 3D image of an object 104 is received at 302. When an image taken with color sensor and an image taken with the depth sensor are used to create the 3D image, image information for each sensor is often calibrated to create an accurate 3D point cloud of the object 104 including coordinates such as (x, y, z)." (¶15). This teaches that a depth sensor and color sensor cooperate to capture both a three-dimensional point cloud and a two-dimensional image of an object resting on a support base surface.);
detecting, from the point cloud, the support surface and a portion of an upper surface of the object (Sharma: "The base, or generally planar surface, on which the object 104 is placed, is removed from the point cloud at 304. In one example, a plane fitting technique is used to remove the base from the point cloud. One such plane fitting technique can be found in tools applying RANSAC (Random sample consensus), which is an iterative method to estimate parameters of a mathematical model from a set of observed data that contains outliers. In
this case, the outliers can be the images of the objects 104 and the inliers can be the image
of the planar base." (¶16). This teaches detecting the planar support surface from the point cloud via RANSAC plane fitting; the non-base outlier points constitute the point cloud portion corresponding to the object's upper surface.);
(addressed below via Lee);
based on the first region, performing a foreground segmentation operation on the image to segment the upper surface of the object from the image (Sharma: "The point cloud with the base removed can be used as a mask to detect the object 104 in the image. The mask includes data points representing the object 104. Once the base has been subtracted from the image, the 3D point cloud is projected onto a 2D plane ... the 2D planar image of the object is subjected to a contour analysis for segmentation." (¶¶17-18). This teaches projecting the non-base point cloud as a 2D image mask and performing segmentation on the masked 2D image to determine the object boundary. Lee further teaches performing this segmentation as a seeded graph-cut operation, as addressed below.);
determining, based on the point cloud, a three-dimensional position of the upper surface segmented from the image (Sharma: "In one example, corrected depth data is used to find the object's height, orientation, or other characteristics of a 3D object." (¶19); Claim 9 of Sharma:
"applying, with the processor, the depth data to determine height of the object." This teaches using point cloud depth data to determine the three-dimensional position of the segmented upper surface.); and
determining dimensions of the object based on the three-dimensional position of the upper surface (Sharma: " corrected depth data is used to find the object's height, orientation, or other characteristics of a 3D object." (¶19); Sharma: Claim 9: "applying, with the processor, the depth data to determine height of the object"; Claim 8: "applying depth data includes determining the orientation of the detected object."; thereby determining object height from the point cloud depth data. McCloskey supplies the full length, width, and height dimensioning as addressed below.).
Sharma does not specifically teach labelling a first region of the image corresponding to
the portion of the upper surface as a foreground region for use as an explicit seed in a graph-cut segmentation operation. Sharma projects the non-base point cloud as a binary mask onto the 2D image and then applies contour analysis, but does not label the projected region as a named "foreground" seed in a graph-cut energy minimization framework. However, Lee, in the same field of RGB-D sensor-based object segmentation, explicitly teaches projecting depth-detected object regions onto the 2D color image and treating those projected regions as foreground segmentation seeds, "Once initial human body regions are obtained from the depth image, we can first project them onto the color image and treat them as segmentation seeds. The segmentation result in the color image can then be obtained by minimizing Equation (2) … the foreground regions are then extracted from the depth image and then projected the foreground region to the color image, as shown in Figure 1c." (Lee, p. 3). This teaches explicitly labelling a depth-derived region of the 2D image as a foreground seed and performing graph-cut energy-minimization segmentation seeded from that label.)
Sharma discloses finding the height of the object but does not specifically teach
determining multiple dimensions. However, McCloskey teaches determining dimensions
(McCloskey further teaches this limitation explicitly in the context of parcel dimensioning: "[D]etected points in a 3D point cloud which lie on top of the object (e.g., and hence lie over the detected void region) can be additionally or alternatively employed to measure height of the object." [0032]. McCloskey also teaches that "[d]imensions of the object can be inferred from a shape and/or a size of the void region for the object" including "length data for the object, width data for the object, height data for the object." [0030], [0041]. This teaches determining all three physical dimensions of the object-height from the 3D position of points on the upper surface relative to the support surface, and length and width from the segmented object boundary-which is the specific dimensional measurement claimed.).
It would have been obvious to one of ordinary skill in the art before the effective filing date to combine Sharma, Lee, and McCloskey. All three references relate to processing 3D point cloud and 2D image data of objects on a support surface. Lee's graph-cut seeding would improve Sharma's contour-based segmentation at low-contrast object surface boundaries, yielding a more accurate object boundary. McCloskey supplies the explicit dimensioning step that Sharma lacks, deriving height, width, and length from the point cloud. Incorporating McCloskey's dimensioning into Sharma's segmentation pipeline would yield the predictable improvement of converting the already-detected upper surface position and support surface into actual physical measurements, enabling automated parcel sizing without additional sensor hardware.
Claims 2 and 12.
Sharma, Lee, and McCloskey discloses the method of claim 1, further comprising: presenting the dimensions on a display of the computing device (Sharma:"[T]he computer 204 includes a display 206 to render images and/or interfaces of the object detection application." (¶9, FIG. 2 discussion). This teaches a computing device with a display configured to render outputs of the object detection application, including computed object dimensions.).
Claims 3 and 13.
Sharma, Lee, and McCloskey discloses the method of claim 1, further comprising: labelling a second region of the image corresponding to the support surface as a background region.
Sharma does not label this support-surface image region as a background region seed for graph cut segmentation. However, Lee teaches that graph-cut framework assigns a background label to image regions outside of the projected foreground seed (Lee: “The probability of x being labeled as l exponentially decreases as the distance from the skeleton of the l-th object increases” and "The probability of x being labeled as the background is defined as p(Lx = 0) = 1 – max p(Lx = l)” (Lee p. 5, Eq. 7)). Because Sharma detects and removes the support surface from the point cloud, the image region corresponding to that support surface lies outside the projected upper-surface foreground seed and is thereby labeled as the background region in the graph-cut energy minimization.).
It would have been obvious to explicitly label the image region corresponding to Sharma's detected support surface as a background region seed in the Lee-style graph-cut segmentation in order to constrain the foreground object region and prevent the segmentation from erroneously including the planar support surface within the segmented object boundary. Doing so is the direct and predictable extension of combining Sharma's explicit support-surface detection with Lee's background seed labeling mechanism.
Claims 4, 6, 14, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Sharma, Lee, and McCloskey, further in view of Xiao et al., "An Effective Graph and Depth Layer Based RGB-D Image Foreground Object Extraction Method," (hereinafter "Xiao").
Claims 4 and 14.
Claims 4 and 14 further require: detecting, in the point cloud, a further surface distinct from the upper surface and the support surface; and labelling a third region of the image corresponding to the further surface as a probably background region.
Sharma, Lee, and McCloskey collectively disclose the limitations of claims 1-3 and 11-13 as discussed above. Sharma, Lee, and McCloskey do not specifically teach detecting, in the point cloud, a further surface distinct from the upper surface and the support surface and labelling a third region of the image corresponding to the further surface as a probably background region. However, Xiao, in the same field of RGB-D image segmentation, teaches both limitations. Xiao partitions the depth map into multiple depth layers, where each layer contains pixels in a range of depth values: "[A] depth layer contains pixels in a range of depth values, and we consider these pixels as the foreground region (white) of the chosen depth layer." (Xiao, Section 1.3). Depth layers corresponding to surfaces at different distances from the sensor-such as a lateral face or a wall behind the object constitute further surfaces distinct from both the upper surface and the support surface. Xiao further teaches labelling regions outside the selected foreground depth layer as background under the regional continuity function: "C(k) = 1 if A_d(k) > T_A • A_c(k); 0 otherwise," where A_d(k) is the overlap of the foreground depth layer with region k. (Xiao, Eq. 6, Section 1.4). Regions with C(k) = 0 are labeled as probable background because they fall outside the foreground object's depth layer.
It would have been obvious to combine Sharma, Lee, McCloskey, and Xiao. Xiao's depth-layer partitioning detects surfaces at different depths and classifies them as foreground or background. Incorporating Xiao 's depth-layer classification into Sharma's pipeline would yield improved segmentation by identifying and excluding non-target surfaces from the foreground region.
Claims 6 and 16.
Sharma, Lee, McCloskey, and Xiao collectively disclose the limitations of claims 1-4 and 11-14 as discussed above. Claims 6 and 16 depend from claims 4 and 14, respectively, and further require labelling a remainder of the image as a probable foreground region. Xiao teaches that regions connected to positive seed regions through depth-layer constraints are merged as foreground: "[W]e maintain positive seed point regions as well as regions which are connected to them." (Xiao, Section 1.2). Xiao's regional continuity function classifies each region as either foreground (C(k) = 1, overlapping the selected depth layer) or background (C(k) = 0). (Xiao, Eq. 6, Section 1.4). Under this scheme, once the confirmed background region (support surface, claim 3) and the probable background region (further surface, claim 4) have been labeled, the remaining image regions that satisfy C(k) = 1 but are not yet confirmed as foreground carry a residual foreground classification (i.e., they are labeled as probable foreground). This is the direct result of Xiao's binary depth-layer classification applied after the surface detection steps of the combination.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Sharma in
view of Lee, McCloskey, and Xiao as applied to claims 4 and 14 above, further in view of Poelman et al., US 10,268,917 B2 (hereinafter "Poelman").
Claims 5 and 15.
Claims 5 and 15 depend from claims 4 and 14, respectively, and further recite that detecting the further surface includes detecting a portion of the point cloud with a normal vector different from a normal vector of the upper surface by at least a threshold. Sharma's RANSAC plane-fitting technique computes a surface normal to each detected plane, including the support surface and the object’s upper surface. However, Sharma, Lee, McCloskey, and Xiao do not specifically teach comparing point cloud surface normal against a threshold to detect a further surface. Poelman teaches this limitations. Poelman calculates "the surface normal of each 3D point" (Poelman col. 8, step 602) and, using a region growing methodology, assigns "two points ... to the same segment if they are spatially close (plane-to-point distances) and the angle between their surface normals is smaller than some certain threshold" (Poelman col. 8, step 606). A point of the point cloud whose surface normal differs from the upper surface normal by more than the threshold is thereby assigned to a different segment, that is, detected as a further surface distinct from the upper surface.
It would have been obvious to one of ordinary skill in the art before the effective filing date to apply Poelman's surface-normal-angle thresholding to the point cloud of Sharma to detect the further surface of claims 4 and 14, in order to distinguish a surface of a different orientation, such as a lateral face of the object or a wall, from the object's upper surface. This evidence is provided in response to Applicant's traversal of the Official Notice taken in the prior action and supports the finding that comparing point-cloud surface normals against a threshold to distinguish surfaces is known in the art.
Claims 7-8 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over of Sharma, Lee, and McCloskey as applied to claims 1 and 11 above, and further in view of Azam et al., "Removing Reflection Artifacts from Point Clouds," (US 20240054621 A1 – hereinafter “Azam”).
Claims 7 and 17.
Claims 7 and 17 depend from claims 1 and 11 , respectively, and further require, prior to
determining dimensions of the object, determining whether the point cloud exhibits multipath
artifacts by: (i) selecting a candidate point on the upper surface; (ii) determining a reflection
score for the candidate point; and (iii) comparing the reflection score to a threshold. Sharma,
Lee, and McCloskey disclose the limitations of claims 1 and 11 as discussed above, except for
specifically teaching determining whether the point cloud exhibits multipath artifacts by
selecting a candidate point on the upper surface, determining a reflection score for the candidate point, and comparing the reflection score to a threshold.
Sharma, Lee, and McCloskey do not specifically teach determining whether the point cloud exhibits multipath artifacts. However, Azam, in the same field of 3D point cloud processing using time-of-flight depth sensors, teaches detecting and removing reflection artifacts from point clouds prior to downstream use of the 3D data: "A computer-implemented method is provided that includes detecting at least one reflective surface in at least one two-dimensional (2D) image of an environment ... projecting the bounding coordinates of the 2D image into a three-dimensional (3D) space of the environment ... identifying a reflection artifact encompassed by the bounding coordinates in the 3D space." (Azam, Abstract). This teaches detecting multipath/reflection artifacts in a 3D point cloud by correlating 2D image reflectance detection with 3D spatial coordinates. Azam further teaches selecting candidate points and applying a threshold-based reflectance score to identify artifacts: "[S]electing candidate 3D points encompassed by the bounding coordinates in the 3D space; clustering the candidate 3D points by intensity values or reflectance values; and selecting at least one of the candidate 3D points as the reflection artifact based at least in part on a threshold associated with the intensity values or the reflectance values." (Azam, Claim 4). This teaches selecting candidate points in the 3D point cloud region corresponding to a reflective surface, computing a reflection-related score (intensity/reflectance value) for each candidate point, and comparing that score to a threshold to determine whether the point is a reflection artifact which is the substance of the claimed reflection score compared to a threshold.
It would have been obvious to combine Sharma, Lee, McCloskey, and Azam. Reflection artifacts are a known source of error in ToF depth sensors. Azam's artifact detection-evaluating candidate points against a reflectance threshold would naturally be incorporated prior to dimensioning to prevent corrupted depth measurements from affecting the output.
Claims 8 and 18.
Claims 8 and 18 depend from claims 7 and 17, respectively, and further recite that selecting the
candidate point includes identifying a non-planar region of the upper surface, and selecting the
candidate point from the non-planar region. Azam teaches a plane-fitting approach to bound the artifact region: "[F]it a plane on the bounding coordinates by using a normal vector of a
center point of the bounding coordinates; find a maximum depth within the bounding
coordinates; and fit a rectangular volume using the maximum depth." (Azam, Claim 6).
Points deviating from the fitted plane are non-planar anomalies, the natural candidates for
artifact evaluation. Focusing candidate selection on non-planar regions is obvious because ToF
multipath artifacts manifest as depth discontinuities on otherwise planar surfaces.
Allowable Subject Matter
Claims 9-10 and 19-20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ross Varndell whose telephone number is (571)270-1922. The examiner can normally be reached M-F, 9-5 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’Neal Mistry can be reached at (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Ross Varndell/Primary Examiner, Art Unit 2674