978Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Double Patenting
1. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1-24 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-50 of U.S. Patent No. 12205319. Although the claims at issue are not identical, they are not patentably distinct from each other because patented claims recite features/limitations of instant application claims.
For Example:
Current Application Claims
Patent Application Claims
1. A method for estimating a three-dimensional (3-D) location of an object captured in a two-dimensional (2-D) image, said method comprising: acquiring position data indicative of 3-D positions of objects in a space; acquiring a 2-D image of said space, said 2-D image including 2-D representations of said objects in said space; associating said position data and said 2-D image to create training data; providing a machine learning framework; using said training data to train said machine learning framework to create a trained machine learning framework capable of determining 3-D positions of target objects represented in 2-D images of target spaces including said target objects; obtaining a subsequent 2-D image of a subsequent space; and utilizing said trained machine learning framework to determine a 3-D position of an object in said subsequent space.
3. The method of claim 2, wherein: said first portion of said trained machine learning framework is configured to generate an encoded tensor based at least in part on said subsequent 2-D image; said step of providing said encoded 2-D image to a second portion of said trained machine learning framework includes providing said encoded tensor to said second portion of said trained machine learning framework; and said step of providing said encoded 2-D image to a third portion of said trained machine learning framework includes providing said encoded tensor to said third portion of said trained machine learning framework.
4. The method of claim 3, wherein said first portion of said trained machine learning framework is a deep learning aggregation network configured to encode image features of said subsequent 2-D image to create encoded image features, said encoded image features existing at varying scales.
2. The method of claim 1, wherein said step of utilizing said trained machine learning framework to determine a 3-D position of an object in said subsequent space includes: providing said subsequent 2-D image to a first portion of said trained machine learning framework configured to encode said subsequent 2-D image to generate an encoded 2-D image; providing said encoded 2-D image to a second portion of said trained machine learning framework configured to determine a 2-D position of said object in said subsequent space; providing said encoded 2-D image to a third portion of said trained machine learning framework configured to estimate a depth of said object within said subsequent space; and combining said 2-D position of said object and said depth of said object to estimate a 3-D position of said object within said subsequent space.
5. The method of claim 2, wherein said steps of providing said encoded 2-D image to a second portion of said trained machine learning framework and providing said encoded 2-D image to a third portion of said trained machine learning framework occur in parallel.
6. The method of claim 2, further comprising determining a 2-D, real-world position of said object in said subsequent space, and wherein: said step of obtaining a subsequent 2-D image of a subsequent space includes capturing said subsequent 2-D image with an image capture device, said image capture device being associated with an intrinsic matrix; said intrinsic matrix represents a relationship between points in an image space of said image capture device and locations in a 3-D world space corresponding to said subsequent space; and said step of determining a 2-D, real-world position of said object in said subsequent space includes associating points in said image space of said image capture device with locations in said 3-D world space based at least in part on said intrinsic matrix.
7. The method of claim 6, wherein: said step of combining said 2-D position of said object and said depth of said object to estimate a 3-D position of said object within said subsequent space includes utilizing the Pythagorean theorem to relate said estimated depth of said object, a coordinate of said 2-D position of said object, and a corrected depth of said object; said estimated depth is a distance between said image capture device and said object as estimated by said third portion of said trained machine learning framework; said corrected depth is a distance between a first plane perpendicular to an optical axis of said image capture device and intersecting said image capture device and a second plane perpendicular to said optical axis and intersecting said object, said depth of said object representing an estimate of the distance between said image capture device and said object; and said step of combining said 2-D position of said object and said depth of said object includes calculating said corrected depth based at least in part on said estimated depth, said coordinate, and the Pythagorean theorem.
8. The method of claim 1, wherein: said step of obtaining a subsequent 2-D image of a subsequent space includes capturing said subsequent image with an image capture device; said image capture device is coupled to a vehicle; said subsequent 2-D image is captured at a particular time; and said subsequent 2-D image is at least partially representative of the surroundings of said vehicle at said particular time.
9. The method of claim 8, further comprising providing said trained machine learning framework to said vehicle, and wherein: said vehicle is an autonomous vehicle; and movements of said autonomous vehicle are informed at least in part by said 3-D position of said object.
10. The method of claim 1, wherein said step of using said training data to train said machine learning framework includes: utilizing said machine learning framework to estimate depths of said objects in said space; utilizing said machine learning framework to estimate 2-D positions of said objects in said space; comparing said estimated depths of said objects to observed depths of said objects obtained from said position data; comparing said estimated 2-D positions of said objects to observed 2-D positions of said objects obtained from position data; generating a loss function based at least in part on said comparison between said estimated depths of said objects and said observed depths of said objects and based at least in part on said comparison between said estimated 2-D positions of said objects and said observed 2-D positions of said objects; and altering said machine learning framework based at least in part on said loss function.
11. The method of claim 10, wherein said step of altering said machine learning framework based at least in part on said loss function includes: calculating a contribution of a node of said machine learning framework to said loss function; and altering at least one value corresponding to said node of said machine learning framework based at least in part on said contribution of said node to said loss function.
12. The method of claim 1, wherein said position data is light detection and ranging (LiDAR) data.
13. A system configured to estimate a three-dimensional (3-D) location of an object captured in a two-dimensional (2-D) image, said system comprising: a hardware processor configured to execute code, said code including a set of native instructions for causing said hardware processor to perform a corresponding set of operations when executed by said hardware processor; and memory for storing data and said code, said data including position data indicative of 3-D positions of objects in a space, a 2-D image of said space, said 2-D image including 2-D representations of said objects in said space, and a subsequent 2-D image of a subsequent space; and wherein said code includes a machine learning framework, a first subset of said set of native instructions configured to associate said position data and said 2-D image to create training data, a second subset of said set of native instructions configured to use said training data to train said machine learning framework to create a trained machine learning framework capable of determining 3-D positions of target objects represented in 2-D images of target spaces including said target objects, and a third subset of said set of native instructions configured to cause said trained machine learning framework to determine a 3-D position of an object in said subsequent space.
14. The system of claim 13, wherein said third subset of said set of native instructions is additionally configured to: provide said subsequent 2-D image to a first portion of said trained machine learning framework configured to encode said subsequent 2-D image to generate an 4 encoded 2-D image; provide said encoded 2-D image to a second portion of said trained machine learning framework configured to determine a 2-D position of said object in said subsequent space; provide said encoded 2-D image to a third portion of said trained machine learning framework configured to estimate a depth of said object within said subsequent 10 space; and combine said 2-D position of said object and said depth of said object to estimate a 3-D position of said object within said subsequent space.
15. The system of claim 14, wherein: said first portion of said trained machine learning framework is configured to generate an encoded tensor based at least in part on said subsequent 2-D image; said third subset of said set of native instructions is additionally configured to provide said encoded tensor to said second portion of said trained machine learning framework; and said third subset of said set of native instructions is additionally configured to provide said encoded tensor to said third portion of said trained machine learning 8 framework.
16. The system of claim 15, wherein said first portion of said trained machine learning framework is a deep learning aggregation network configured to encode image features of said subsequent 2-D image to create encoded image features, said encoded image features existing at varying scales.
17. The system of claim 14, wherein said third subset of said set of native instructions is additionally configured to provide said encoded 2-D image to said second portion of said trained machine learning framework and to said third portion of said trained machine learning framework in parallel.
18. The system of claim 14, further comprising an image capture device associated with an intrinsic matrix, and wherein: said code includes a fourth subset of said set of native instructions configured to determine a 2-D, real-world position of said object in said subsequent space; said subsequent 2-D image is captured by said image capture device; said intrinsic matrix represents a relationship between points in an image space of said image capture device and locations in a 3-D world space corresponding to said subsequent space; and said fourth subset of said set of native instructions is additionally configured to associate points in said image space of said image capture device with locations in said 3-D world space based at least in part on said intrinsic matrix.
19. The system of claim 18, wherein: said third subset of said set of native instructions is additionally configured to utilize the Pythagorean theorem to relate said estimated depth of said object, a coordinate of said 2-D position of said object, and a corrected depth of said object; said estimated depth is a distance between said image capture device and said object as estimated by said third portion of said trained machine learning framework; said corrected depth is a distance between a first plane perpendicular to an optical axis of said image capture device and intersecting said image capture device and a second plane perpendicular to said optical axis and intersecting said object, said depth of said object representing an estimate of the distance between said image capture device and said object; and said third subset of said set of native instructions is additionally configured to calculate said corrected depth based at least in part on said estimated depth, said coordinate, and the Pythagorean theorem.
20. The system of claim 13, further comprising: a vehicle; and an image capture device coupled to said vehicle; and wherein said subsequent 2-D image is captured by said image capture device at a particular time; and said subsequent 2-D image is at least partially representative of the surroundings of said vehicle at said particular time.
21. The system of claim 20, further comprising a network adapter configured to establish a data connection between said hardware processor and said vehicle, and wherein: said code includes a fourth subset of said set of native instructions configured to provide said trained machine learning framework to said vehicle; said vehicle is an autonomous vehicle; and movements of said autonomous vehicle are informed at least in part by said 3-D position of said object.
22. The system of claim 13, wherein said second subset of said set of native instructions is additionally configured to: utilize said machine learning framework to estimate depths of said objects in said space; utilize said machine learning framework to estimate 2-D positions of said objects in said space; compare said estimated depths of said objects to observed depths of said objects obtained from said position data; compare said estimated 2-D positions of said objects to observed 2-D positions of said objects obtained from position data; generate a loss function based at least in part on said comparison between said estimated depths of said objects and said observed depths of said objects and at least in part on said comparison between said estimated 2-D positions of said objects and said observed 2-D positions of said objects; and alter said machine learning framework based at least in part on said loss function.
23. The system of claim 22, wherein said second subset of said set of native instructions is additionally configured to: calculate a contribution of a node of said machine learning framework to said loss function; and alter at least one value corresponding to said node of said machine learning framework based at least in part on said contribution of said node to said loss function.
24. The method of claim 13, wherein said position data is light detection and ranging (LiDAR) data.
1. A method for estimating a three-dimensional (3-D) location of an object captured in a two-dimensional (2-D) image, said method comprising: acquiring position data indicative of 3-D positions of objects in a space; acquiring a 2-D image of said space, said 2-D image including 2-D representations of said objects in said space; associating said position data and said 2-D image to create training data; providing a machine learning framework; using said training data to train said machine learning framework to create a trained machine learning framework capable of determining 3-D positions of target objects represented in 2-D images of target spaces including said target objects; obtaining a subsequent 2-D image of a subsequent space; and utilizing said trained machine learning framework to determine a 3-D position of an object in said subsequent space; and wherein said step of utilizing said trained machine learning framework to determine said 3-D position of an object in said subsequent space includes providing said subsequent 2-D image to a first portion of said trained machine learning framework configured to encode said subsequent 2-D image to generate an encoded 2-D image, providing said encoded 2-D image to a second portion of said trained machine learning framework configured to determine a 2-D position of said object in said subsequent space, providing said encoded 2-D image to a third portion of said trained machine learning framework configured to estimate a depth of said object within said subsequent space, and combining said 2-D position of said object and said depth of said object to estimate a 3-D position of said object within said subsequent space.
2. The method of claim 1, wherein: said first portion of said trained machine learning framework is configured to generate an encoded tensor based at least in part on said subsequent 2-D image; said step of providing said encoded 2-D image to said second portion of said trained machine learning framework includes providing said encoded tensor to said second portion of said trained machine learning framework; and said step of providing said encoded 2-D image to said third portion of said trained machine learning framework includes providing said encoded tensor to said third portion of said trained machine learning framework.
3. The method of claim 2, wherein said first portion of said trained machine learning framework is a deep learning aggregation network configured to encode image features of said subsequent 2-D image to create encoded image features, said encoded image features existing at varying scales.
24. The method of claim 23, wherein: said step of utilizing said trained machine learning framework to determine said 3-D position of said object in said subsequent space includes providing said subsequent 2-D image to a first portion of said trained machine learning framework configured to encode said subsequent 2-D image to generate an encoded 2-D image, providing said encoded 2-D image to a second portion of said trained machine learning framework configured to determine a 2-D position of said object in said subsequent space, providing said encoded 2-D image to a third portion of said trained machine learning framework configured to estimate a depth of said object within said subsequent space, and combining said 2-D position of said object and said depth of said object to estimate a 3-D position of said object within said subsequent space; said first portion of said trained machine learning framework is configured to generate an encoded tensor based at least in part on said subsequent 2-D image; said step of providing said encoded 2-D image to said second portion of said trained machine learning framework includes providing said encoded tensor to said second portion of said trained machine learning framework; and said step of providing said encoded 2-D image to said third portion of said trained machine learning framework includes providing said encoded tensor to said third portion of said trained machine learning framework.
26. The method of claim 23, wherein: said step of utilizing said trained machine learning framework to determine said 3-D position of said object in said subsequent space includes providing said subsequent 2-D image to a first portion of said trained machine learning framework configured to encode said subsequent 2-D image to generate an encoded 2-D image, providing said encoded 2-D image to a second portion of said trained machine learning framework configured to determine a 2-D position of said object in said subsequent space, providing said encoded 2-D image to a third portion of said trained machine learning framework configured to estimate a depth of said object within said subsequent space, and combining said 2-D position of said object and said depth of said object to estimate a 3-D position of said object within said subsequent space; and said steps of providing said encoded 2-D image to said second portion of said trained machine learning framework and providing said encoded 2-D image to said third portion of said trained machine learning framework occur in parallel.
5. The method of claim 1, further comprising determining a 2-D, real-world position of said object in said subsequent space, and wherein: said step of obtaining said subsequent 2-D image of said subsequent space includes capturing said subsequent 2-D image with an image capture device, said image capture device being associated with an intrinsic matrix; said intrinsic matrix represents a relationship between points in an image space of said image capture device and locations in a 3-D world space corresponding to said subsequent space; and said step of determining said 2-D, real-world position of said object in said subsequent space includes associating points in said image space of said image capture device with locations in said 3-D world space based at least in part on said intrinsic matrix.
6. The method of claim 5, wherein: said step of combining said 2-D position of said object and said depth of said object to estimate said 3-D position of said object within said subsequent space includes utilizing the Pythagorean theorem to relate said estimated depth of said object, a coordinate of said 2-D position of said object, and a corrected depth of said object; said estimated depth is a distance between said image capture device and said object as estimated by said third portion of said trained machine learning framework; said corrected depth is a distance between a first plane perpendicular to an optical axis of said image capture device and intersecting said image capture device and a second plane perpendicular to said optical axis and intersecting said object, said depth of said object representing an estimate of the distance between said image capture device and said object; and said step of combining said 2-D position of said object and said depth of said object includes calculating said corrected depth based at least in part on said estimated depth, said coordinate, and the Pythagorean theorem.
7. The method of claim 1, wherein: said step of obtaining said subsequent 2-D image of said subsequent space includes capturing said subsequent image with an image capture device; said image capture device is coupled to a vehicle; said subsequent 2-D image is captured at a particular time; and said subsequent 2-D image is at least partially representative of the surroundings of said vehicle at said particular time.
8. The method of claim 7, further comprising provided said trained machine learning framework to said vehicle, and wherein: said vehicle is an autonomous vehicle; and movements of said autonomous vehicle are informed at least in part by said 3-D position of said object.
9. The method of claim 1, wherein said step of using said training data to train said machine learning framework includes: utilizing said machine learning framework to estimate depths of said objects in said space; utilizing said machine learning framework to estimate 2-D positions of said objects in said space; comparing said estimated depths of said objects to observed depths of said objects obtained from said position data; comparing said estimated 2-D positions of said objects to observed 2-D positions of said objects obtained from position data; generating a loss function based at least in part on said comparison between said estimated depths of said objects and said observed depths of said objects and based at least in part on said comparison between said estimated 2-D positions of said objects and said observed 2-D positions of said objects; and altering said machine learning framework based at least in part on said loss function.
10. The method of claim 9, wherein said step of altering said machine learning framework based at least in part on said loss function includes: calculating a contribution of a node of said machine learning framework to said loss function; and altering at least one value corresponding to said node of said machine learning framework based at least in part on said contribution of said node to said loss function.
11. The method of claim 1, wherein said position data is light detection and ranging (LiDAR) data.
12. A system configured to estimate a three-dimensional (3-D) location of an object captured in a two-dimensional (2-D) image, said system comprising: a hardware processor configured to execute code, said code including a set of native instructions for causing said hardware processor to perform a corresponding set of operations when executed by said hardware processor; and memory for storing data and said code, said data including position data indicative of 3-D positions of objects in a space, a 2-D image of said space, said 2-D image including 2-D representations of said objects in said space, and a subsequent 2-D image of a subsequent space; and wherein said code includes a machine learning framework, a first subset of said set of native instructions configured to associate said position data and said 2-D image to create training data, a second subset of said set of native instructions configured to use said training data to train said machine learning framework to create a trained machine learning framework capable of determining 3-D positions of target objects represented in 2-D images of target spaces including said target objects, and a third subset of said set of native instructions configured to cause said trained machine learning framework to determine a 3-D position of an object in said subsequent space, provide said subsequent 2-D image to a first portion of said trained machine learning framework configured to encode said subsequent 2-D image to generate an encoded 2-D image, provide said encoded 2-D image to a second portion of said trained machine learning framework configured to determine a 2-D position of said object in said subsequent space, provide said encoded 2-D image to a third portion of said trained machine learning framework configured to estimate a depth of said object within said subsequent space, and combine said 2-D position of said object and said depth of said object to estimate a 3-D position of said object within said subsequent space.
13. The system of claim 12, wherein: said first portion of said trained machine learning framework is configured to generate an encoded tensor based at least in part on said subsequent 2-D image; said third subset of said set of native instructions is additionally configured to provide said encoded tensor to said second portion of said trained machine learning framework; and said third subset of said set of native instructions is additionally configured to provide said encoded tensor to said third portion of said trained machine learning framework.
14. The system of claim 13, wherein said first portion of said trained machine learning framework is a deep learning aggregation network configured to encode image features of said subsequent 2-D image to create encoded image features, said encoded image features existing at varying scales.
15. The system of claim 12, wherein said third subset of said set of native instructions is additionally configured to provide said encoded 2-D image to said second portion of said trained machine learning framework and to said third portion of said trained machine learning framework in parallel.
16. The system of claim 12, further comprising an image capture device associated with an intrinsic matrix, and wherein: said code includes a fourth subset of said set of native instructions configured to determine a 2-D, real-world position of said object in said subsequent space; said subsequent 2-D image is captured by said image capture device; said intrinsic matrix represents a relationship between points in an image space of said image capture device and locations in a 3-D world space corresponding to said subsequent space; and said fourth subset of said set of native instructions is additionally configured to associate points in said image space of said image capture device with locations in said 3-D world space based at least in part on said intrinsic matrix.
17. The system of claim 16, wherein: said third subset of said set of native instructions is additionally configured to utilize the Pythagorean theorem to relate said estimated depth of said object, a coordinate of said 2-D position of said object, and a corrected depth of said object; said estimated depth is a distance between said image capture device and said object as estimated by said third portion of said trained machine learning framework; said corrected depth is a distance between a first plane perpendicular to an optical axis of said image capture device and intersecting said image capture device and a second plane perpendicular to said optical axis and intersecting said object, said depth of said object representing an estimate of the distance between said image capture device and said object; and said third subset of said set of native instructions is additionally configured to calculate said corrected depth based at least in part on said estimated depth, said coordinate, and the Pythagorean theorem.
18. The system of claim 12, further comprises: a vehicle; and an image capture device coupled to said vehicle; and wherein said subsequent 2-D image is captured by said image capture device at a particular time; and said subsequent 2-D image is at least partially representative of the surroundings of said vehicle at said time.
19. The system of claim 18, further comprising a network adapter configured to establish a data connection between said hardware processor and said vehicle, and wherein: said code includes a fourth subset of said set of native instructions configured to provide said trained machine learning framework to said vehicle; said vehicle is an autonomous vehicle; and movements of said autonomous vehicle are informed at least in part by said 3-D position of said object.
20. The system of claim 12, wherein said second subset of said set of native instructions is additionally configured to: utilize said machine learning framework to estimate depths of said objects in said space; utilize said machine learning framework to estimate 2-D positions of said objects in said space; compare said estimated depths of said objects to observed depths of said objects obtained from said position data; compare said estimated 2-D positions of said objects to observed 2-D positions of said objects obtained from position data; generate a loss function based at least in part on said comparison between said estimated depths of said objects and said observed depths of said objects and at least in part on said comparison between said estimated 2-D positions of said objects and said observed 2-D positions of said objects; and alter said machine learning framework based at least in part on said loss function.
21. The system of claim 20, wherein said second subset of said set of native instructions is additionally configured to: calculate a contribution of a node of said machine learning framework to said loss function; and alter at least one value corresponding to said node of said machine learning framework based at least in part on said contribution of said node to said loss function.
22. The method of claim 12, wherein said position data is light detection and ranging (LiDAR) data.
Claim Rejections - 35 USC § 103
2. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1,8-9,12-13,20 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Rad et al. US 20180268601 A1, in view of Satzoda et al. US 20180365888 A1.
Regarding claim 1, Rad provides for a method for estimating a three- dimensional (3-D) location of an object captured in a two-dimensional (2-D) image ( see [0028] The system determines a three-dimensional ("3D") pose of an object from a two-dimensional ("2D") input image which contains the object"), said method comprising: acquiring position data indicative of 3-D positions of objects in a space; acquiring a 2-D image of said space, said 2-D image including 2-D representations of said objects in said space ( see [0079], see" The corners of mirror image 940m, i.e., 941m, 942m, 943m, and 944m have changed their relative positions. This change can be compensated for upon determining the 3D bounding box of the mirror image 940m"), associating said position data and said 2-D image to create training data; providing a machine learning framework; using said training data to train said machine learning framework to create a trained machine learning framework capable of determining 3-D positions of target objects represented in 2-D images of target spaces including said target objects ( see Fig.3 and Fig.11, see "[0092] In blocks 1114 and 1116, the 2D projections are refined by inputting the image or the patch of the image, and the estimated 3D pose of the object to a neural network in the refiner block 337 (FIG. 3)". Rad does not provide for obtaining a subsequent 2-D image of a subsequent space; and utilizing said trained machine learning framework to determine a 3-D position of an object in said subsequent space. Satzoda teaches the above missing limitation of Rad (see [0079], see "estimating an expected future object pose based on the determined object pose and kinematics, tracking the object through one or more subsequent frames, and verifying the determined object pose and kinematics in response to the subsequently-determined object pose (determined from the subsequent frames) substantially matching the expected object pose (e.g., within an error threshold))" it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention, to combine the teaching of Satzoda with the system and method of Rod, in order for verifying the determined object pose, via estimating an expected future object pose based on the determined object pose and kinematics, tracking the object through one or more subsequent frames.
Regarding claim 8, Rad does not provide for, wherein: said step of obtaining a subsequent 2-D image of a subsequent space includes capturing said subsequent image with an image capture device; said image capture
device is coupled to a vehicle; said subsequent 2-D image is captured at a
time: and said subsequent 2-D image is at least partially
representative of the surroundings of said vehicle at said time.
Satzoda teaches the above missing limitation of Rad (see [0079], see claim
9 of Satzoda, see "capturing image data using a camera of an onboard
vehicle system coupled to a vehicle, wherein the image data depicts an
exterior environment of the vehicle; determining that a rare vehicle event
is depicted by the image data, determining an event type of the rare vehicle
event, and tagging the image data with the event type to generate tagged
image data"). it would have been obvious to one of ordinary skill in the art
before the effective filing data of the claimed invention, to combine the
teaching of Satzoda with the system and method of Rod, for
verifying the determined object pose, via estimating an expected future
object pose based on the determined object pose and kinematics, tracking
the object through one or more subsequent frames.
Regarding claim 9, Rad provides for trained machine learning
framework to said vehicle (see "[0029] The system uses a classifier (e.g., a
neural network) to determine whether an image shows an object having arotation angle within a first predetermined range (corresponding to objects on a first side of the plane of symmetry)"). Rad does not provide for, wherein: said vehicle is an autonomous vehicle; and movements of said autonomous vehicle are informed at least in part by said 3-D position of said object. Satzoda teaches the above missing limitation of Rad (see [0020], see "In a second example, the method can generate 3D reconstructions (e.g., exterior obstacles and travel paths, internal vehicle conditions, agent parameters during a vehicle event, etc.) of vehicle accidents (e.g., static scene, moments leading up to the accident, moments after the accident, etc.). The method can also function to generate abstracted representations (e.g., virtual models) of vehicle events (e.g., driving events, events that occur with respect to a vehicle while driving, etc.), which can include agents (e.g., objects that operate in the traffic environment, moving objects, controllable objects in the vehicle environment, etc.) and agent parameters; provide abstracted representations of vehicle events to a third party (e.g., a vehicle controller, an autonomous vehicle operator, an insurer, a claim adjuster, and any other third party"). it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention, to combine the teaching of Satzoda with the system and method of Rod, in order for
verifying the determined object pose, via estimating an expected future
object pose based on the determined object pose and kinematics, tracking
the object through one or more subsequent frames.
Regarding claim 12, Rad does not provide for, wherein said position
data is light detection and ranging 2 (LiDAR) data. Satzoda teaches the
above missing limitation of Rad (see [0020], see "[0025] Second, variations
of the system and method can reconstruct 3D models from 2D imagery,
which can be translated and/or transformed for use with other vehicle
sensor systems (e.g., LIDAR, stereo cameras, collections of monocular
cameras, etc.). These 3D models can be augmented, verified, and/or
otherwise processed using auxiliary sensor data"). it would have been
obvious to one of ordinary skill in the art before the effective filing data of
the claimed invention, to combine the teaching of Satzoda with the system
and method of Rod, in order for verifying the determined object pose, via
estimating an expected future object pose based on the determined object
pose and kinematics, tracking the object through one or more subsequent
frames.
Regarding claim 13, see the rejection of claim 1. It recites similar
limitations as claim 13. Except for a hardware processor and a set of
native instructions (see Rad, see" [0088] Additionally, each of the
processes 1100 and/or 1200 may be performed under the control of
one or more computer systems configured with executable
instructions and may be implemented as code (e.g., executable
instructions, one or more computer programs". Hence it is similarly
analyzed and rejected.
Regarding claim 20, see the rejection of claim 8. It recites similar
limitations as claim 20. Hence it is similarly analyzed and rejected.
Regarding claim 24, see the rejection of claim 12. It recites similar
limitations as claim 24. Hence it is similarly analyzed and rejected.
Allowable Subject Matter
3. Claims 2-7,10-11,14-19 and 21-23 are objected to as
being dependent upon a rejected base claim, but would be allowable
if rewritten in independent form including all of the limitations of the
base claim and any intervening claims.
Reasons for Allowance
4. The following is an examiner’s statement of reasons for allowance: the closest prior arts of Rad et al. US20180268601 A1, in view of Satzoda et al. US 20180365888 A1, failed to teach or suggest for features/limitations of claims 2-7,10-11,14-19 and 21-23.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Rad et al. US 20180137644 A1, is cited because the reference teaches "determining, for each of the plurality of training images, corresponding two-dimensional projections of the three-dimensional bounding box of the object, the corresponding two-dimensional projections being determined by projecting the corresponding three-dimensional locations of the three- dimensional bounding box onto an image plane of each of the plurality of training images", in [0015].
Hess et al. US 20200125845 A1, is cited because the reference teaches "[0023] Autonomous, semi-autonomous, or manually-driven vehicles can include and/or utilize one or more trained machine learning models. For example, one or more machine learning models can be trained to identify objects in a vehicle's surrounding environment based on data received from one or more sensors".
Mao et al. US 20210365707 A1, is cited because the reference teaches "At block 1312, the process 1310 includes obtaining the one or more subsequent frame. In some cases, a single iteration of the process 1310 can be performed for one frame at a time from the one or more subsequent frames", in [0231].
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALI BAYAT whose telephone number is (571)272-7444. The examiner can normally be reached 9:00-5:00 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicants are encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at 571-2705183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALI BAYAT/Primary Examiner, Art Unit 2677