DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
35 U.S.C. 101 requires that a claimed invention must fall within one of the four eligible categories of invention (i.e. process, machine, manufacture, or composition of matter) and must not be directed to subject matter encompassing a judicially recognized exception as interpreted by the courts. MPEP 2106. The four eligible categories of invention include: (1) process which is an act, or a series of acts or steps, (2) machine which is an concrete thing, consisting of parts, or of certain devices and combination of devices, (3) manufacture which is an article produced from raw or prepared materials by giving to these materials new forms, qualities, properties, or combinations, whether by hand labor or by machinery, and (4) composition of matter which is all compositions of two or more substances and all composite articles, whether they be the results of chemical union, or of mechanical mixture, or whether they be gases, fluids, powders or solids. MPEP 2106(I).
Claims directed toward purely data such as alphanumeric data, image data, video data or music data are not directed toward a process since the data by itself does not perform any steps or acts. And, the claims are not directed toward machine, manufacture or composition of matter because those categories require a tangible object which is not satisfied by pure data. Furthermore, even if the data is placed on a non-transitory computer readable media (i.e. manufacture) the claims would still be ineligible because the thrust of the claims are directed toward the data with the non-transitory computer readable media acting as merely a carrier of that data.
Claim 21 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter since the claims are directed toward data which is not any of the four eligible categories of invention.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-4, 12, 14-16, 20 and 21 are rejected under 35 U.S.C. 102(a)(2) as being US 12,175,783 B2 by Jotwani.
Regarding claim 1, Jotwani discloses a system for generating three-dimensional (3D) point cloud data (column 2, lines 40-44, wherein systems, and methods for determining an estimate of a position of a camera. In various implementations, a method is performed at a device including one or more processors and non-transitory memory. The method includes obtaining a point cloud of a physical environment) the system comprising: at least one processor (column 2, lines 42-43, wherein a method is performed at a device including one or more processors and non-transitory memory) configured to: generate, based on two-dimensional (2D) image data, the 3D point cloud data, wherein each point of the 3D point cloud data (column 2, lines 43-51, wherein the method includes obtaining a point cloud of a physical environment including a plurality of points, wherein each of the plurality of points is associated with a set of three-dimensional coordinates in a three-dimensional coordinate system of the physical environment, wherein the plurality of points includes a first cluster of points associated with a first semantic label. The method includes obtaining a two-dimensional image of the physical environment with a camera associated with one or more intrinsic parameters of the camera) comprises: position information, the position information comprising at least three coordinates indicating a position of the point (column 6, lines 16-24, and Fig. 5, wherein the point cloud data object 500 includes a plurality of data elements (shown as rows in FIG. 5), wherein each data element is associated with a particular point of a point cloud. The data element for a particular point includes a point identifier field 510 that includes a point identifier of a particular point. As an example, the point identifier may be a unique number. The data element for the particular point includes a coordinate field 520 that includes a set of coordinates in a three-dimensional space of the particular point.); object information labeling the point as a first object selected from a plurality of objects (column 6, lines 25-27 and Fig. 5, wherein the data element for the particular point includes a cluster identifier field 530, as object information labeling, that includes an identifier of the cluster into which the particular point is spatially disambiguated); and category information labeling the point as belonging to a first category of a plurality of categories (column 6, lines 28-31, and Fig. 5, wherein the cluster identifier may be a letter or number. The data element for the particular point includes a semantic label field 540, as the category label, that includes a semantic label for the cluster into which the particular point is spatially disambiguated.), wherein each of the plurality of categories comprises at least one respective object of the plurality of objects (column 4, lines 30-34, wherein at least one of the plurality of points is further associated with a semantic label that represents an object type or identity of the surface of the object. For example, the semantic label may be “tabletop” or “table” or “wall”).
Regarding claim 2, Jotwani discloses wherein the plurality of categories are defined based on a task (column 6, lines 9-11, wherein In various implementations, the semantic label indicates an object type or identity of the object, inherently as different tasks of object either types or identity determination).
Regarding claim 3, Jotwani discloses wherein a number of the plurality of categories is less than a number of the plurality of objects (column 6, lines 25-31, and Fig. 5, wherein the data element for the particular point includes a cluster identifier field 530, as each object, that includes an identifier of the cluster into which the particular point is spatially disambiguated. As an example, the cluster identifier may be a letter or number. The data element for the particular point includes a semantic label field 540, as each category, that includes a semantic label for the cluster into which the particular point is spatially disambiguated, and wherein number of categories 540 are less than the number objects 530).
Regarding claim 4, Jotwani discloses the system further comprising a memory that stores a table comprising a label for each of the plurality of categories and a label for each of the plurality of objects (column 6, lines 12-16 and Fig. 5, wherein the handheld electronic device 110 stores the semantic label in association with each point of the cluster. FIG. 5 illustrates a point cloud data object 500, inherently as the table, in accordance with some implementations.).
Regarding claim 12, Jotwani discloses wherein the 2D image data is generated by a camera (column 5, lines 7-10, wherein the handheld electronic device 110 includes a single scene camera (or single rear-facing camera disposed on an opposite side of the handheld electronic device 110 as the display).
Regarding claim 14, Jotwani discloses wherein the at least one first processor is further configured to display the 3D point cloud data on a display (column 5, lines 52-55, wherein FIG. 4A illustrates the handheld electronic device 110 displaying the first image 211A overlaid with the representation of the point cloud 310 spatially disambiguated into a plurality of clusters 412-416).
Regarding claim 15, Jotwani discloses wherein the at least one first processor is configured to display the 3D point cloud data based on a selection of one or more of the plurality of objects and/or a selection of one or more of the plurality of categories (column 12, lines 23-32, wherein in block 840, with the device determining a plurality of sets of two-dimensional coordinates in a two-dimensional coordinate system of the two-dimensional image of the physical environment corresponding to the representation of the first object, inherently as selected predetermined or pretrained object. Once a representation of the first object is detected in the two-dimensional image, the device determines a plurality of sets of two-dimensional coordinates in the two-dimensional coordinate system of the two-dimensional image that would correspond to points in the point cloud).
Regarding claim 16, Please refer to the corresponding system claim 1 above for further teachings.
Regarding claim 20, at least one non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by at least one processor, cause the at least one processor (Column 3, lines 5-9, wherein a device includes one or more processors, a non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors) to perform a method (Please refer to the corresponding system claim 1 above for further teachings).
Regarding claim 21, Please refer to the corresponding system claim 1 and storage medium 20 above for further teachings.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 5-7 and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Jotwani in view of US 2022/0121852 A1 to Curtis et al (hereinafter ‘Curtis’).
Regarding claims 5 and 17, Jotwani does not specifically discloses wherein the at least one first processor is further configured to: perform 2D object recognition processing on the 2D image data to generate labeled 2D image data including the object information; classify the labeled 2D image data to generate classified 2D image data including the category information; and convert the classified 2D image data to generate the 3D point cloud data including the position information, the object information, and the category information. Curtis discloses perform 2D object recognition processing on the 2D image data to generate labeled 2D image data including the object information (Para [0105], wherein FIG. 4F shows an example of a labeled 2D image according to an embodiment. The image 470 captures a scene 472 including one or more shelves 473, 474, 475, arranged vertically, The image 470 shows one or more products 476, 478, 480, 482, which are identified according to the embodiments and labeled with a label 477, 479, 481, 483, respectively); classify the labeled 2D image data to generate classified 2D image data including the category information (Para [0100], wherein For example, while two classes or types of products 403, 404 are identified among the four total objects identified on the shelf 405, the system may indicate if other classes or types of products are present.); and convert the classified 2D image data to generate the 3D point cloud data including the position information, the object information, and the category information (Para [0073] and [0075], wherein the detected scene depth advantageously complements the RGB images of the 2D object detection conducted by the 2D object detection model. The 3D object detection process localizes and classifies objects in 3D space, and wherein the system 100 utilizes three distinct artificial neural networks 110, 120, 130 for synthesizing 2D images, for example RGB images 102, and 3D point cloud data 112, to generate a predicted count of objects in a detected scene). Jotwani and Curtis are combinable because they both disclose image object detection. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one ordinary skill in the art to combine the convert the classified 2D image data to generate the 3D point cloud data including the position information, the object information, and the category information of Curtis’ system with Jotwani’s in order to count objects in the three-dimensional space Para [0014]).
Regarding claim 6 and 18, in the combination of Jotwani and Curtis, Curtis further discloses wherein the at least one first processor is configured to perform the 2D object recognition processing at least in part using a machine learning model (Para [0072], wherein the 2D image recognition assessment may be a 2D object detection model that treats image classification as a regression problem, conducting one pass through an associated neural network to predict what front-facing objects are in the image and where they are present).
Regarding claims 7 and 19, in the combination of Jotwani and Curtis, Curtis further discloses wherein the at least one first processor is configured to classify the labeled 2D image data using a machine learning model (Para [0072], wherein the 2D image recognition assessment may be a 2D object detection model that treats image classification as a regression problem, conducting one pass through an associated neural network to predict what front-facing objects are in the image and where they are present).
Claims 8-11 are rejected under 35 U.S.C. 103 as being unpatentable over Jotwani in view of US 11,216,663 B1 to Ettinger et al (hereinafter ‘Ettinger’).
Regarding claim 8, Jotwani does not specifically discloses wherein the 2D image data comprises a plurality of 2D image data including a first set of 2D image data and a second set of 2D image data, wherein a field of view of the first set of 2D image data at least partially overlaps with a field of view of the second set of 2D image data, and the at least one processor is configured to generate the 3D point cloud data based on the plurality of 2D image data. Ettinger discloses wherein the 2D image data comprises a plurality of 2D image data including a first set of 2D image data and a second set of 2D image data, wherein a field of view of the first set of 2D image data at least partially overlaps with a field of view of the second set of 2D image data (column 36, lines 46-52, wherein viewports for the object, feature, scene, or location of interest can be generated from the same acquired sensor data, where the acquired data is processed to provide different type of visual representations. For example, 3D point clouds or 3D meshes can be derived from a plurality of overlapping 2D images acquired of an object). Jotwani and Ettinger are combinable because they both disclose image object detection. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one ordinary skill in the art to combine the convert the classified 2D image data to generate the 3D point cloud data including the position information, the object information, and the category information of Ettinger’s system with Jotwani’s in order to determine a canonical representation (e.g., orientation) for the given dataset (column 54, lines 60-61).
Regarding claim 9, in the combination of Jotwani and Ettinger, Ettinger further discloses wherein the at least one first processor is further configured to receive depth image data, a field of view of the depth image data at least partially overlaps with a field of view of the 2D image data (column 35, lines 53-57, wherein the set of 2D overlapping RGB imagery and their corresponding 3D point cloud could be used together to generate an RGBD dataset. This new dataset is generated via adding a new depth channel to the existing RGB image set, and.
Regarding claim 10, in the combination of Jotwani and Ettinger, Ettinger further discloses wherein the depth image data is generated by a depth sensor (column 36, lines 60-64, wherein sensor data can be acquired from different sensors operational in the same sensor data acquisition event. For example, a UAV can be outfitted with a first sensor such as a 2D image camera and a second sensor such as a 3D depth sensor or LIDAR).
Regarding claim 11, in the combination of Jotwani and Ettinger, Ettinger further discloses wherein the depth image data is generated based on the plurality of 2D image data (column 37, line 65 through column 38, line 1, wherein the synthetic image (i.e., 2D representation) provides a virtual—or synthetic—snapshot of the data from a certain location and direction; it could be presented as a depth, grayscale, or color (e.g., RGB) image.).
Regarding claim 13, Jotwani does not specifically discloses wherein the 2D image data comprises a plurality of pixels, and the at least one first processor is further configured to generate a category label map having the category information for each pixel of the 2D image data and wherein the at least one first processor is configured to generate the 3D point cloud data based on the category map. Ettinger discloses wherein the 2D image data comprises a plurality of pixels, and the at least one first processor is further configured to generate a category label map having the category information for each pixel of the 2D image data (column 52, lines 48-53, wherein Deep Convolutional Neural Networks (DCNNs) can be used to assigning a label to one or more portions of an image (e.g., bounding box, region enclosed by a contour, or a set of pixels creating a regular or irregular shape) that include a given object, feature) and wherein the at least one first processor is configured to generate the 3D point cloud data based on the category map (column 37, 16-23, wherein such separately generated sensor data can be when a sensor data acquisition event is conducted for an object of interest and a second sensor dataset was acquired at a different time for that same object of interest, for example sensor data available from an image library that includes sensor data for the object of interest, such as that available from the NearMap® library, Google Maps®, or a GIS library). Jotwani and Ettinger are combinable because they both disclose image object detection. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one ordinary skill in the art to combine a category label map of Ettinger’s system with Jotwani’s in order to detect an object of interest at a different time for that same object of interest (column 37, lines 17-19).
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHERVIN K NAKHJAVAN whose telephone number is (571)272-5731. The examiner can normally be reached Monday-Friday 9:00-12:00 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sue Lefkowitz can be reached at (571)272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHERVIN K NAKHJAVAN/Primary Examiner, Art Unit 2672