Prosecution Insights
Last updated: August 17, 2026
Application No. 18/925,395

IMAGE ANNOTATION USING LOCALIZED EMBEDDINGS

Non-Final OA §103§112
Filed
Oct 24, 2024
Examiner
ZHANG, WAYNE
Art Unit
2672
Tech Center
2600 — Communications
Assignee
Walmart Apollo LLC
OA Round
1 (Non-Final)
56%
Grant Probability
Moderate
1-2
OA Rounds
1y 1m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
14 granted / 25 resolved
-6.0% vs TC avg
Strong +40% interview lift
Without
With
+40.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
21 currently pending
Career history
44
Total Applications
across all art units

Statute-Specific Performance

§101
17.1%
-22.9% vs TC avg
§103
44.9%
+4.9% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
24.9%
-15.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 25 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The IDS dated 10/30/2024 has been considered and placed in the application file. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claim(s) 1-20 are rejected under 35 U.S.C. 112(b), as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claim 1 (and correspondingly claim 8 and 15) recites “determine a vertical position encoding and a horizontal position encoding for the at least one annotated object, wherein the vertical position encoding is determined based on the vertical position maximum of the reference image and the dimension size of the at least one embedding and the horizontal position encoding is determined based on the horizontal position maximum of the reference image and the dimension size of the at least one embedding”. It is unclear what the Applicant is claiming when stating a “vertical position encoding” and “horizontal position encoding”. The definition of image encoding (for one of ordinary skill in the art) is the process of converting data into another format. However, the Applicant has claimed these encodings are based on the “maximum of the reference image and dimension size of the embedding”, a meaning that is different from what is standardly known. It is also unclear what the Applicant is referring to when stating maximum, as maximum by definition means “most”. For examination purposes, the examiner will interpret this limitation as determining the boundaries of an object. Claim 1 (and correspondingly claim 8 and 15) also additionally recites “And identify an unannotated object in image data based on a comparison of the first cluster centroid and a second cluster centroid generated for the unannotated object”. It is unclear what the Applicant is claiming when stating “cluster centroid”. Cluster centroid by definition (for one of ordinary skill in the art) is the center point/average of a cluster group. However, the Applicant has stated that the cluster centroid is generated by embeddings and the position encodings of an object. Thus, it is indefinite as to what the Applicant is claiming when stating a “cluster centroid”. For examination purposes, the examiner will interpret this limitation as generating a cluster from the embeddings and position encodings. It is also unclear how the “second cluster centroid” is being generated for the unannotated object, as the process prior to this only applies to annotated objects. Claim 5 (and correspondingly claim 12 and 19) recites “wherein the image data including the unannotated object is the reference image”. It is unclear what the Applicant is claiming when stating “the unannotated object is the reference image”, as an object is presumably residing in an image. Independent claim 1 also states “receive a reference image including at least one annotated object”, which is a direct contradiction to what claim 5 is claiming. For examination purposes, the examiner will interpret this limitation as the unannotated object additionally residing in the reference image. Claims 2-7, 9-14, and 16-20 are all rejected for their dependencies on claims 1, 8, and 15 respectively. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kharbanda (US 11978271 B1) in view of Lin (US 20240404268 A1). Regarding claim 1, Kharbanda discloses a system (Kharbanda, Col. 5, Lines 52-65, "The systems and methods disclosed herein can process an image with a vision language model and a fine-grained object recognition model in parallel to generate an output that is scene-aware and object-aware while being formatted in a natural language format"), comprising: a processor (Kharbanda, Col. 20, Lines 62-67, "The user computing system 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, a FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected"), And a non-transitory memory storing instructions, that when executed, cause the processor to (Kharbanda, Col. 21, Lines 1-4, "The memory 114 can include one or more non-transitory computer-readable storage mediums, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof."), receive a reference image including at least one annotated object (Kharbanda, Col. 2, Lines 65-67, "In some implementations, the input image can be descriptive of the object in an environment with one or more additional objects"), generate at least one embedding representative of at least one feature of the at least one annotated object (Kharbanda, Col. 8, Lines 46-49, "The one or more embedding models may process the image data 12 and/or the image segments to generate one or more image embeddings"), determine a dimension size of the at least one embedding, a vertical position maximum of the reference image, and a horizontal position maximum of the reference image (Kharbanda, Col. 2, Lines 2-7, "Generating the object embedding can include generating a bounding box associated with a position of the object within the input image, generating an image segment based on the bounding box, and processing the image segment with an embedding model to generate the object embedding."), determine a vertical position encoding and a horizontal position encoding for the at least one annotated object, wherein the vertical position encoding is determined based on the vertical position maximum of the reference image and the dimension size of the at least one embedding and the horizontal position encoding is determined based on the horizontal position maximum of the reference image and the dimension size of the at least one embedding (Kharbanda, Col. 10, Lines 48-50, "The object recognition model may include a detection model that processes the input images to generate bounding boxes indicating a position of the detected objects"), combine the at least one feature embedding, the vertical position encoding, and the horizontal position encoding to generate a first cluster centroid for the at least one annotated object (Kharbanda, Col. 13, Lines 35-37, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding", as stated above, the object embedding generates a position of the object within the image. Thus, the embedding is combined with the information it contains to create a cluster). While Kharbanda discloses identifying an object in image data based on a comparison of the first cluster centroid and a second cluster centroid generated for the annotated object (Kharbanda, Col. 13, Lines 35-46, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding. The neighbor embeddings, similar embeddings, and/or embedding clusters may be associated with image embeddings, text embeddings, document embeddings, multimodal embeddings, and/or other embeddings. Content items and/or web resources associated with the neighbor embedding(s), similar embedding(s), and/or embedding cluster(s) can be obtained and/or processed to determine the identification details. The identification details can include a precise name (and/or classification) for the particular object", neural networks are naturally trained to compare data with previous training data. In this scenario, web resources with clusters from previous training data is compared to the current cluster of the object), they do not do so through an “unannotated object in image data”. However, Lin teaches identifying an unannotated object in image data (Lin, paragraph [0018], "Once the CNN model 104 is trained, the CNN model 104 can be utilized to detect objects in unannotated images"). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to use unannotated image data in place of Kharbanda’s annotated images, as taught by Lin. The suggestion/motivation for doing so would have been to save storage and require less computational resources. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Kharbanda in view of Lin to obtain the invention as specified in claim 1. Regarding claim 2, Kharbanda in view of Lin discloses the system of claim 1, wherein the vertical position encoding and the horizontal position encoding comprises position embeddings (Kharbanda, Col. 2, Lines 2-7, "Generating the object embedding can include generating a bounding box associated with a position of the object within the input image, generating an image segment based on the bounding box, and processing the image segment with an embedding model to generate the object embedding."). Regarding claim 3, Kharbanda in view of Lin discloses the system of claim 2, wherein combining the at least one feature embedding, the vertical position encoding, and the horizontal position encoding includes concatenating the at least one feature embedding, the vertical position encoding, and the horizontal position encoding (Kharbanda, Col. 13, Lines 35-37, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding", as stated in claim 1, the object embeddings contain the position of the object). Regarding claim 4, Kharbanda in view of Lin discloses the system of claim 1, wherein the reference image is a first image and the image data including the unannotated object is a second image (Kharbanda, Col. 25, Lines 59-63, "The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images and, for each region, a likelihood that region depicts an object of interest"). Regarding claim 5, Kharbanda in view of Lin discloses the system of claim 1, wherein the image data including the unannotated object is the reference image (Kharbanda, Col. 25, Lines 59-63, "The input image can be descriptive of the object in an environment with one or more additional objects", through the modification with Lin, the image is unannotated). Regarding claim 6, Kharbanda in view of Lin discloses the system of claim 1, where the instructions cause the processor to: generate a second reference image from the image data including an annotation identifying the unannotated object as the at least one annotated object based on the comparison of the first cluster centroid and a second cluster centroid; and identify a second unannotated object in second image data based on a comparison of the second cluster centroid and a third cluster centroid generated for the unannotated object based on the second reference image (Kharbanda, Col. 13, Lines 35-46, "In some implementations, one or more embedding clusters can be determined to be associated with the object embedding. The neighbor embeddings, similar embeddings, and/or embedding clusters may be associated with image embeddings, text embeddings, document embeddings, multimodal embeddings, and/or other embeddings. Content items and/or web resources associated with the neighbor embedding(s), similar embedding(s), and/or embedding cluster(s) can be obtained and/or processed to determine the identification details. The identification details can include a precise name (and/or classification) for the particular object", this process of finding a second reference image to find a second unannotated object is a repeat of the method of claim 1. Due to the nature of training neural networks i.e. being continuously trained based on the input they receive, the neural network can repeat the steps needed to reach claim 6.). Regarding claim 7, Kharbanda in view of Lin discloses the system of claim 1, wherein the identification of the unannotated object in the image data is provided for training a computer vision task (Kharbanda, Col. 25, Line 52-53, "In some cases, the input includes visual data and the task is a computer vision task."). Claims 8-14 corresponds to claims 1-7, additionally reciting a computer-implemented method (Kharbanda, Col. 1, Lines 39-40, “One example aspect of the present disclosure is directed to a computer-implemented method”). Thus, they are rejected for the same reasons of obviousness as claims 1-7. Claims 15-20 corresponds to claims 1-6, additionally reciting a non-transitory computer readable medium having instructions stored thereon (Kharbanda, Col. 2, Lines 42-48, “The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining image data”). Thus, they are rejected for the same reasons of obviousness as claims 1-7. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WAYNE ZHANG whose telephone number is (571) 272-0245. The examiner can normally be reached Monday-Friday 10:00-6:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Sumati Lefkowitz can be reached on (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WAYNE ZHANG/Examiner, Art Unit 2672 /SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672
Read full office action

Prosecution Timeline

Oct 24, 2024
Application Filed
Jul 22, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705709
METHOD AND DEVICE FOR GENERATING ALL-IN-FOCUS IMAGE
3y 10m to grant Granted Aug 11, 2026
Patent 12694539
METHOD FOR TRACKING POSITION OF OBJECT AND SYSTEM FOR TRACKING POSITION OF OBJECT
2y 10m to grant Granted Jul 28, 2026
Patent 12688614
POINT CLOUD ENCODING AND DECODING METHOD AND DEVICE BASED ON TWO-DIMENSIONAL REGULARIZATION PLANE PROJECTION
3y 1m to grant Granted Jul 21, 2026
Patent 12670570
3D VOLUME INSPECTION METHOD AND METHOD OF CONFIGURING OF A 3D VOLUME INSPECTION METHOD
3y 4m to grant Granted Jun 30, 2026
Patent 12591990
METHOD AND APPARATUS FOR GENERATING SPATIAL GEOMETRIC INFORMATION ESTIMATION MODEL
3y 6m to grant Granted Mar 31, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
56%
Grant Probability
96%
With Interview (+40.0%)
2y 11m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 25 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month