DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The IDS dated 10/30/2024 has been considered and placed in the application file.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claim(s) 1-20 are rejected under 35 U.S.C. 112(b), as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim 1 (and correspondingly claim 8 and 15) recites “determine a vertical position encoding and a horizontal position encoding for the at least one annotated object, wherein the vertical position encoding is determined based on the vertical position maximum of the reference image and the dimension size of the at least one embedding and the horizontal position encoding is determined based on the horizontal position maximum of the reference image and the dimension size of the at least one embedding”. It is unclear what the Applicant is claiming when stating a “vertical position encoding” and “horizontal position encoding”. The definition of image encoding (for one of ordinary skill in the art) is the process of converting data into another format. However, the Applicant has claimed these encodings are based on the “maximum of the reference image and dimension size of the embedding”, a meaning that is different from what is standardly known. It is also unclear what the Applicant is referring to when stating maximum, as maximum by definition means “most”. For examination purposes, the examiner will interpret this limitation as determining the boundaries of an object.
Claim 1 (and correspondingly claim 8 and 15) also additionally recites “And identify an unannotated object in image data based on a comparison of the first cluster centroid and a second cluster centroid generated for the unannotated object”. It is unclear what the Applicant is claiming when stating “cluster centroid”. Cluster centroid by definition (for one of ordinary skill in the art) is the center point/average of a cluster group. However, the Applicant has stated that the cluster centroid is generated by embeddings and the position encodings of an object. Thus, it is indefinite as to what the Applicant is claiming when stating a “cluster centroid”. For examination purposes, the examiner will interpret this limitation as generating a cluster from the embeddings and position encodings. It is also unclear how the “second cluster centroid” is being generated for the unannotated object, as the process prior to this only applies to annotated objects.
Claim 5 (and correspondingly claim 12 and 19) recites “wherein the image data including the unannotated object is the reference image”. It is unclear what the Applicant is claiming when stating “the unannotated object is the reference image”, as an object is presumably residing in an image. Independent claim 1 also states “receive a reference image including at least one annotated object”, which is a direct contradiction to what claim 5 is claiming. For examination purposes, the examiner will interpret this limitation as the unannotated object additionally residing in the reference image.
Claims 2-7, 9-14, and 16-20 are all rejected for their dependencies on claims 1, 8, and 15 respectively.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kharbanda (US 11978271 B1) in view of Lin (US 20240404268 A1).
Regarding claim 1, Kharbanda discloses a system (Kharbanda, Col. 5, Lines 52-65, "The systems and methods disclosed herein can process an image with a vision language model and a fine-grained object recognition model in parallel to generate an output that is scene-aware and object-aware while being formatted in a natural language format"), comprising:
a processor (Kharbanda, Col. 20, Lines 62-67, "The user computing system 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, a FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected"),
And a non-transitory memory storing instructions, that when executed, cause the processor to (Kharbanda, Col. 21, Lines 1-4, "The memory 114 can include one or more non-transitory computer-readable storage mediums, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof."),
receive a reference image including at least one annotated object (Kharbanda, Col. 2, Lines 65-67, "In some implementations, the input image can be descriptive of the object in an environment with one or more additional objects"),
generate at least one embedding representative of at least one feature of the at least one annotated object (Kharbanda, Col. 8, Lines 46-49, "The one or more embedding models may process the image data 12 and/or the image segments to generate one or more image embeddings"),
determine a dimension size of the at least one embedding, a vertical position maximum of the reference image, and a horizontal position maximum of the reference image (Kharbanda, Col. 2, Lines 2-7, "Generating the object embedding can include generating a bounding box associated with a position of the object within the input image, generating an image segment based on the bounding box, and processing the image segment with an embedding model to generate the object embedding."),
determine a vertical position encoding and a horizontal position encoding for the at least one annotated object, wherein the vertical position encoding is determined based on the vertical position maximum of the reference image and the dimension size of the at least one embedding and the horizontal position encoding is determined based on the horizontal position maximum of the reference image and the dimension size of the at least one embedding (Kharbanda, Col. 10, Lines 48-50, "The object recognition model may include a detection model that processes the input images to generate bounding boxes indicating a position of the detected objects"),
combine the at least one feature embedding, the vertical position encoding, and the horizontal position encoding to generate a first cluster centroid for the at least one annotated object (Kharbanda, Col. 13, Lines 35-37, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding", as stated above, the object embedding generates a position of the object within the image. Thus, the embedding is combined with the information it contains to create a cluster).
While Kharbanda discloses identifying an object in image data based on a comparison of the first cluster centroid and a second cluster centroid generated for the annotated object (Kharbanda, Col. 13, Lines 35-46, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding. The neighbor embeddings, similar embeddings, and/or embedding clusters may be associated with image embeddings, text embeddings, document embeddings, multimodal embeddings, and/or other embeddings. Content items and/or web resources associated with the neighbor embedding(s), similar embedding(s), and/or embedding cluster(s) can be obtained and/or processed to determine the identification details. The identification details can include a precise name (and/or classification) for the particular object", neural networks are naturally trained to compare data with previous training data. In this scenario, web resources with clusters from previous training data is compared to the current cluster of the object), they do not do so through an “unannotated object in image data”.
However, Lin teaches identifying an unannotated object in image data (Lin, paragraph [0018], "Once the CNN model 104 is trained, the CNN model 104 can be utilized to detect objects in unannotated images").
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to use unannotated image data in place of Kharbanda’s annotated images, as taught by Lin.
The suggestion/motivation for doing so would have been to save storage and require less computational resources.
Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Kharbanda in view of Lin to obtain the invention as specified in claim 1.
Regarding claim 2, Kharbanda in view of Lin discloses the system of claim 1, wherein the vertical position encoding and the horizontal position encoding comprises position embeddings (Kharbanda, Col. 2, Lines 2-7, "Generating the object embedding can include generating a bounding box associated with a position of the object within the input image, generating an image segment based on the bounding box, and processing the image segment with an embedding model to generate the object embedding.").
Regarding claim 3, Kharbanda in view of Lin discloses the system of claim 2, wherein combining the at least one feature embedding, the vertical position encoding, and the horizontal position encoding includes concatenating the at least one feature embedding, the vertical position encoding, and the horizontal position encoding (Kharbanda, Col. 13, Lines 35-37, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding", as stated in claim 1, the object embeddings contain the position of the object).
Regarding claim 4, Kharbanda in view of Lin discloses the system of claim 1, wherein the reference image is a first image and the image data including the unannotated object is a second image (Kharbanda, Col. 25, Lines 59-63, "The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images and, for each region, a likelihood that region depicts an object of interest").
Regarding claim 5, Kharbanda in view of Lin discloses the system of claim 1, wherein the image data including the unannotated object is the reference image (Kharbanda, Col. 25, Lines 59-63, "The input image can be descriptive of the object in an environment with one or more additional objects", through the modification with Lin, the image is unannotated).
Regarding claim 6, Kharbanda in view of Lin discloses the system of claim 1, where the instructions cause the processor to: generate a second reference image from the image data including an annotation identifying the unannotated object as the at least one annotated object based on the comparison of the first cluster centroid and a second cluster centroid; and identify a second unannotated object in second image data based on a comparison of the second cluster centroid and a third cluster centroid generated for the unannotated object based on the second reference image (Kharbanda, Col. 13, Lines 35-46, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding. The neighbor embeddings, similar embeddings, and/or embedding clusters may be associated with image embeddings, text embeddings, document embeddings, multimodal embeddings, and/or other embeddings. Content items and/or web resources associated with the neighbor embedding(s), similar embedding(s), and/or embedding cluster(s) can be obtained and/or processed to determine the identification details. The identification details can include a precise name (and/or classification) for the particular object", this process of finding a second reference image to find a second unannotated object is a repeat of the method of claim 1. Due to the nature of training neural networks i.e. being continuously trained based on the input they receive, the neural network can repeat the steps needed to reach claim 6.).
Regarding claim 7, Kharbanda in view of Lin discloses the system of claim 1, wherein the identification of the unannotated object in the image data is provided for training a computer vision task (Kharbanda, Col. 25, Line 52-53, "In some cases, the input includes visual data and the task is a computer vision task.").
Claims 8-14 corresponds to claims 1-7, additionally reciting a computer-implemented method (Kharbanda, Col. 1, Lines 39-40, “One example aspect of the present disclosure is directed to a computer-implemented method”). Thus, they are rejected for the same reasons of obviousness as claims 1-7.
Claims 15-20 corresponds to claims 1-6, additionally reciting a non-transitory computer readable medium having instructions stored thereon (Kharbanda, Col. 2, Lines 42-48, “The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining image data”). Thus, they are rejected for the same reasons of obviousness as claims 1-7.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WAYNE ZHANG whose telephone number is (571) 272-0245. The examiner can normally be reached Monday-Friday 10:00-6:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Sumati Lefkowitz can be reached on (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WAYNE ZHANG/Examiner, Art Unit 2672
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672