Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 8-9 and 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Chaturvedi (US 11,126,845) in view of Watanabe et al. (US 2021/0019906) in view of Maschmeyer et al. (US 2023/0260203) in view of Bauer et al. (US 2022/0302091).
Regarding claim 1, Chaturvedi discloses a method (Chaturvedi, col 2. 26, “Systems and methods”) comprising:
obtaining, from an object detection engine trained to recognize a plurality of objects (Chaturvedi, col 2. 49-51, “analyzing the image data using a trained neural network and the selection to determine one or more types of items for the representation”. In addition, in col 3. 36-38, “a trained neural network may recognize corners of a table, of chairs, and of a television from the image data”), an image representing a space (Chaturvedi, Fig. 1) and including an object of interest located in the space (Chaturvedi, Fig. 1);
Chaturvedi does not expressly disclose “converting, based on the location of the object within the image, a source location of an image capture device which captured the image and a three-dimensional representation, the location of the object to a three-dimensional location of the object within the three-dimensional representation of the space”;
Watanabe et al. (hereinafter Watanabe) discloses converting, based on an object within an image, a source location of an image capture device which captured the image, the location of the object to a three-dimensional location of the object within the three-dimensional representation of the space (Watanabe, [0026], “object detection is performed on a 2D still image, wherein the detection results are projected onto a 3D reconstructed space, and 3D positions of objects are specified”. In addition, in paragraph [0031], “by projecting from 2D coordinates of a detected object to 3D data (e.g. point cloud, voxel, mesh)”. In addition, in paragraph [0046], “the 2D object detection recognizing an object as illustrated with the bounding of FIGS. 4(a)”. Fig. 4(b) illustrates a source location of an image capture device which captured the image);
a location of the object of interest within the image (Watanabe, [0007], “determining a location of the object in three dimensional (3D) space from the point cloud”);
an indication of the three-dimensional location of the object of interest (Watanabe, [0031], “estimating the 3D position of detected objects by raycasting to point clouds”. Fig. 11(a) illustrates indication).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the augmented reality view of the physical environment of Chaturvedi to incorporate the projection of 2D to 3D space using 3D constructed data to accurate identify objects, as taught by Watanabe in order to improve the accuracy of object recognition and placement of augmented reality content.
Chaturvedi as modified by Watanabe does not expressly disclose “a three-dimensional representation”;
Maschmeyer et al. (hereinafter Maschmeyer) discloses a three-dimensional representation (Maschmeyer, [0145], “provide one or more 3D scans of a real-world space. Obtaining a 3D scan may include moving or rotating a sensor with a real-world space to capture multiple angles of the real-world space. LIDAR, radar, and photogrammetry (creating a 3D model from a series of 2D images) are example methods for generating a 3D scan of a real-world space”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Watanabe to determine 3D position of an object using a 3D scan of a real-world environment, as taught by Maschmeyer in order to improve detection accuracy to determine the object’s location within the physical environment.
Chaturvedi as modified by Watanabe and Maschmeyer does not expressly disclose “updating the three-dimensional representation of the space”;
Bauer et al. (hereinafter Bauer) discloses updating a three-dimensional representation of a space (Bauer, [0044], “measuring an environment using one or more measuring devices and automatically updating a digital geometric representation of the environment”. In addition, in paragraph [0092], “A method 900 for updating the digital twin includes recognition and locating of changes in the environment 500, at block 902. The locating of the change includes determining a region of interest in which the change has been detected…the method 900 includes updating the existing digital twin with the new scan of the region of interest”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chaturvedi as modified by Watanabe and Maschmeyer to update the 3D representation of the environment by updating the existing digital twin with the new scan of the region of interest, as taught by Bauer in order to maintain an accurate and current representation of the physical environment for improved object localization.
Regarding claim 2, Chaturvedi discloses obtaining captured data representing the space (Chaturvedi, col 3. 64-66, “capture a physical environment 100/118 in accordance with an embodiment”); extracting the image from the captured data (Chaturvedi, col 12. 58-62, “the image data from an image, a video, or a live camera view of a physical environment, is provided as input to a trained NN that is able to determine scene information, color information, and other information capable of being used as aspects from the image data”); and feeding the image to the object detection engine to recognize the object of interest and identify the location of the object of interest (col 3. 36-38, “a trained neural network may recognize corners of a table, of chairs, and of a television from the image data”).
Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses identify the location of the object of interest (Watanabe, [0007], “determining a location of the object in three dimensional (3D) space from the point cloud”).
Regarding claim 3, Chaturvedi discloses selecting a representative video frame from video data (Chaturvedi, col 12. 58-62, “the image data from an image, a video, or a live camera view of a physical environment, is provided as input to a trained NN that is able to determine scene information, color information, and other information capable of being used as aspects from the image data”).
Regarding claim 4, Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses the location of the object is represented by a bounding box about the object (Watanabe, [0030], “FIG. 4(a) illustrates an example of detecting objects from 2D images by representing the coordinates of the detected objects with bounding boxes”).
Regarding claim 8, Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses a marker located a predefined distance above the three-dimensional location of the object (Watanabe, Fig. 11(d))
Regarding claim 9, Chaturvedi discloses in real-time (col 2. 32-35, “An approach may include receiving image data (e.g., image still and/or video data) of a live camera view from the camera, where the image data includes a representation of a physical environment”).
Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses presenting the indication of the three-dimensional location of the object as an overlay in a current capture view of a data capture device (Watanabe, [0054], “FIG. 11(a) illustrates an example interface panel that can be shown on a user device such as a mobile device (e.g., smartphone), for conducting 2D image recognition”. Fig. 11(a) illustrates presenting the indication of the three-dimensional location of the object as an overlay in a current capture view of a data capture device).
Regarding claim 14, Chaturvedi discloses a server (Chaturvedi, col 5. 38-42, “FIG. 2 illustrates an example data flow diagram 200 of a computing device 204 interacting with a product or item search system or server 220, 222 to obtain types of items or products 208 for an aspect 226A, 226B selected in a live camera view 224 in accordance with various embodiments”) comprising:
a memory (Chaturvedi, Fig. 9 illustrates a webserver including memory)
a communications interface (Chaturvedi, Fig. 9 illustrates a webserver including a communication interface); and
a processor interconnected with the memory and the communications interface (Chaturvedi, Fig. 9 illustrates a webserver including a processor interconnected with the memory and the communications interface).
The remaining limitations recite in claim 14 are similar in scope to the method recited in claim 1 and therefore are rejected under the same rationale.
Regarding claim 15, Chaturvedi teaches the object detection engine; Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses implemented by the server (Watanabe, [0036], “image decode unit 701, image recognition unit 703, object management unit 705, 3D reconstruction unit 702, 3D data storage 704 and display control unit 706 can reside on an external server configured to conduct background processing on the images received from data capturing device 700”).
Regarding claim 16, Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses one or more neural networks, machine learning, or artificial intelligence algorithms (Chaturvedi, col 3. 30-33, “the image data can be analyzed using a trained neural network (NN), machine learning approach, or other such approach”).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Chaturvedi (US 11,126,845) in view of Watanabe et al. (US 2021/0019906) in view of Maschmeyer et al. (US 2023/0260203) in view of Bauer et al. (US 2022/0302091), as applied to claim 1, in further view of Aucsmith et al. (US 6,873,723).
Regarding claim 7, Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses define a boundary of the object (Watanabe, [0030], “FIG. 4(a) illustrates an example of detecting objects from 2D images by representing the coordinates of the detected objects with bounding boxes”);
Chaturvedi as modified by Watanabe, Maschmeyer and Bauer does not expressly disclose “cross-correlating the image to one or more further images including the object”;
Aucsmith et al. (hereinafter Aucsmith) discloses cross-correlating the image to one or more further images including an object (Aucsmith, col 4. 8-13, “The correspondence analyzer 226 matches segmented regions or points from the left image to the right image or vice versa. The objective of the correspondence analysis is to established the correspondence between the left and the right images so that depth or range of regions or points can be computed”. Fig. 3).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to determine the bounding box of the object in Watanabe by using the correspondence between matching regions in multiple images, as taught by Aucsmith in order to improve object localization accuracy.
Claims 10-13 are rejected under 35 U.S.C. 103 as being unpatentable over Chaturvedi (US 11,126,845) in view of Watanabe et al. (US 2021/0019906) in view of Maschmeyer et al. (US 2023/0260203) in view of Bauer et al. (US 2022/0302091), as applied to claim 1, in further view of Adeel et al. (US 2024/0046515).
Regarding claim 10, Chaturvedi discloses during a data capture operation at a data capture device (Chaturvedi, col 12. 58-59, “the image data from an image, a video, or a live camera view of a physical environment”):
extracting image representing a current capture view of the data capture device (Chaturvedi, col 12. 58-59, “the image data from an image, a video, or a live camera view of a physical environment”);
sending image to the object detection engine for training (Chaturvedi, Fig. 6 illustrates sub-process 616 analyzes the image to find the scene information and then process 626 the image is processed to train the neural network);
Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses identifying location within the image of the object of interest (Watanabe, [0006], “determining a location of the object in three dimensional (3D) space from the point cloud”);
Chaturvedi as modified by Watanabe, Maschmeyer and Bauer does not expressly disclose “receiving an indication of a further object of interest”;
Adeel et al. (hereinafter Adeel) discloses receiving an indication of a further object of interest (Adeel, [0025], “the user 112 may select such an image 402 of an object (e.g., a canine) in a video frame 118 that is of interest to the user but for which no current machine-learning models of the content processing engine 104 are able to recognize”);
sending the further image to an object detection engine for training (Adeel, [0025], “submit information that results in the training of a new machine-learning model to recognize a new type of object…the user 112 may select such an image 402 of an object (e.g., a canine) in a video frame 118 that is of interest to the user but for which no current machine-learning models of the content processing engine 104 are able to recognize”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to train the neural network to identify object as taught by Chaturvedi based on a machine learning training process that determine whether the object is detectable based on user’s selection, as taught by Adeel in order to improve accuracy of object identification.
Regarding claim 11, Chaturvedi as modified by Watanabe with the same motivation from claim 1 discloses receiving a single point of input (Watanabe, [0044], “an interface is provided as shown at FIG. 11(a) or FIG. 11(d) in which objects can be added or deleted by tapping on a touch screen”).
Chaturvedi as modified by Watanabe, Maschmeyer, Bauer and Adeel with the same motivation from claim 10 discloses applying one or more image processing algorithms based on user’s input (Adeel, [0012], “the user may use the user interface controls of the content processing engine to apply a machine-learning model that is trained to detect the objects of the specific object type to the video frame”).
Regarding claim 12, Chaturvedi as modified by Watanabe, Maschmeyer, Bauer and Adeel with the same motivation from claim 10 discloses receiving a boundary about the further object of interest (Adeel, [0025], “In order to select the image 402, the user 112 may draw a perimeter border or apply a perimeter shape that surrounds the image 402”).
Regarding claim 13, Chaturvedi as modified by Watanabe, Maschmeyer, Bauer and Adeel with the same motivation from claim 10 discloses presenting the further image with the further location of the further object of interest at the data capture device for confirmation (Adeel, [0026], “The new machine-learning module is then integrated into the content processing engine 104 for use. In this way, the content processing engine 104 may be used to detect the objects of the specific object type that the content processing engine 104 was previously unable to detect”. After the training of the machine-learning model is completed, an image including the previously undetectable object may be displayed for confirmation).
Allowable Subject Matter
Claims 5-6 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYLE ZHAI whose telephone number is (571)270-3740. The examiner can normally be reached 9AM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ke Xiao can be reached at (571) 272 - 7776. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KYLE ZHAI/Primary Examiner, Art Unit 2611