Prosecution Insights
Last updated: August 17, 2026
Application No. 18/885,173

Image Recognition for the Identification of Incorrect Items

Non-Final OA §101§103§Other
Filed
Sep 13, 2024
Priority
Sep 13, 2023 — provisional 63/538,248
Examiner
SHIN, SOO JUNG
Art Unit
Tech Center
Assignee
Maplebear Inc.
OA Round
1 (Non-Final)
87%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
540 granted / 620 resolved
+27.1% vs TC avg
Strong +16% interview lift
Without
With
+16.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
31 currently pending
Career history
643
Total Applications
across all art units

Statute-Specific Performance

§101
8.4%
-31.6% vs TC avg
§103
38.5%
-1.5% vs TC avg
§102
18.1%
-21.9% vs TC avg
§112
26.0%
-14.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 620 resolved cases

Office Action

§101 §103 §Other
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Statement Regarding 35 USC § 101 The claims are directed to a machine learning algorithm using a weight to identify items depicted in a captured image for a mobile order application, e.g., Instacart. The examiner considers the weights of the machine learning model to improve current machine learning technologies because one cannot use random weights for identifying specific objects. Therefore, the examiner is of the opinion that the claims are eligible under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 5, 6, 7, 10, 11, 15, 16, 17, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Talbot et al. (US 2020/0219043 A1), in view of van Horne et al. (US 2022/0114648 A1), hereinafter referred to as Talbot and van Horne, respectively. Regarding claim 1, Talbot teaches a method comprising: providing, to a client device of a picker, a prompt to capture one or more images of items (Talbot ¶¶0004: “prompting a user to capture an image of an array of physical items, capturing the image with the integrated camera, and sending the captured image to a remote server”); applying a machine learning model to the one or more images to classify the items in the one or more images to one or more products (Talbot ¶¶0004: “applying one or more machine learning models … an image annotation data set defining an array of segments, each segment corresponding to a physical item in the array of physical items and having an associated product information, a given associated product information determined using a trained product model that identifies a product identifier based on a portion of the image that corresponds to a given segment of the image”; Talbot ¶¶0126: “associate product information with each identified physical item and may produce an annotated image that includes respective product information for respective identified physical items. (The product information may be any suitable product information, such as a brand name, a product name, a product type, a container size, a product class or category, or any other suitable information (or combinations thereof))”); for each of the one or more users, matching the classified items to the user’s order (Talbot Fig. 5D & ¶¶0182: “an ‘audit product tags’ interface in which users are tasked with reviewing multiple segments that have been associated with given product information to ensure that the products in each of those segments were correctly identified … The user may review the segments 550 and select those that do not match the image 548. As shown, the top row of segments 552 all appear to be the same product as the image 548, while the bottom row of segments 554 appear different”); for at least one of the one or more users, responsive to identifying one or more classified items which do not match the user’s order, highlighting the one or more classified items to produce an annotated image (Talbot ¶¶0143: “Product information 264 (e.g., 264-1, . . . , 264-n) may also be displayed for each item in the annotated image 260 … selecting one of the visual indicators will cause associated product information 264 to be prominently displayed (e.g., highlighted, shown in a separate page or interface or popup window, or the like)”; Talbot Fig. 5D & ¶¶0182 discussed above); and causing the client device to display the annotated image to the picker (Talbot Fig. 5D discussed above; also see Talbot Fig. 5H: 578; Talbot ¶¶0188: “an interface for reviewing potentially mislabeled segments. These segments may be automatically identified by the image analysis engine, or by operators who encounter segments that may be mislabeled and flag them for further review. The interface may display segments that have been identified as possibly mislabeled (e.g., segments 580, 582), as well as information 581, 583 indicating the current product information associated with those segments”). However, Talbot does not appear to explicitly teach fulfilling orders for one or more users of an online system and applying weights of a machine learning model to the images. Pertaining to the same field of endeavor, van Horne teaches fulfilling orders for one or more users of an online system (van Horne Abstract: “A customer places an order of items to be purchased with an online concierge system … requests an image of a receipt of the order from the picker”); and applying weights of a machine learning model to the images (van Horne ¶¶0003: “a delivery system generates and uses machine-learned models to identify items and corresponding actual amount purchased in an image of a receipt of an order”; van Horne ¶¶0034: “trained by the image processing module 216 on the training images to determine relative weights of kernel functions within each machine learned model to provide a desired output”). Talbot and van Horne are considered to be analogous art because they are directed to image processing for recognizing objects. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the technique for identifying products within an image captured using a terminal device (as taught by Talbot) to fulfill users’ orders and applying weights of a machine learning model (as taught by van Horne) because the combination provides more convenience to the users and the machine learning weights can be adjusted to provide a desired output (van Horne ¶¶0002 & ¶¶0034). Regarding claim 5, Talbot, in view of van Horne, teaches the method of claim 1, wherein highlighting the one or more classified items to produce the annotated image comprises, for the at least one of the one or more users: identifying whether each product included in the user’s order has a corresponding matching classified item in the one or more images (Talbot ¶¶0142: “The status identifier may indicate whether or not each item that was identified in the captured image was successfully associated with product information or a product identifier. For example, the status identifier 256-1 shows that the status is incomplete, thus prompting the user to select that scene to provide additional information about the products in the image”); in response to identifying that a product included in the user’s order has no corresponding matching classified item in the one or more images, providing a notification to the client device of the picker of a missing product that is included in the user’s order (Talbot ¶¶0144: “items without a product identifier or product information (e.g., segments whose contents were not able to be identified with a sufficient confidence metric to satisfy a confidence condition) may be associated with distinct visual indicators 262 (e.g., 262-1, . . . , 262-n). These indicators may prompt the user to select those items and provide the missing product data (and/or confirm or reject suggested product information)”), and in response to identifying that a classified item in the one or more images has no matching product in the user’s order, highlighting the classified item in the annotated image (Talbot ¶¶0142-¶¶0144 discussed above). Regarding claim 6, Talbot, in view of van Horne, teaches the method of claim 1, wherein the machine learning model is trained using a training set of a plurality of images of items, the images in the training set being labeled with information identifying products that match the items captured in the images (Talbot ¶¶0072: “the training data used to generate the machine learning model(s) of the image analysis engine 110 may include photographs of items, each associated with a product identifier or product information that identifies the product in the photograph”; Talbot ¶¶0192: “features are extracted from the image and/or from the individual segments of the image. At operation 608, the features are analyzed with an image analysis engine. The image analysis engine may be or include a machine learning model trained with images (or segments of images) that have been manually tagged with product identifiers”). Regarding claim 7, Talbot, in view of van Horne, teaches the method of claim 1, further comprising: performing an action responsive to identifying, for the at least one of the one or more users, the one or more classified items which do not match the user’s order (Talbot Fig. 5, ¶¶0142-¶¶0144, ¶¶0182, ¶¶0188 discussed above). Regarding claim 10, Talbot, in view of van Horne, teaches the method of claim 1, further comprising: responsive to displaying the annotated image on the client device of the picker, receiving from the client device of the picker, feedback indicating that the classification by the machine learning model is incorrect; and training the machine learning model based on the feedback (Talbot ¶¶0082: “results of the machine learning model(s) that have been confirmed to be correct (e.g., segments whose contents were confirmed by a human operator to have been accurately identified) may be used to periodically (or continuously) retrain the model(s)”; Talbot Fig. 5I & ¶¶0189: “an interface for reviewing a corpus of segments that are all labeled or otherwise associated with the same product identifier (e.g., the same UPC) to allow a user to remove segments that have been incorrectly associated with that UPC. For example, an array 586 of segments may be displayed. The user may review the array 586 and select segments (e.g., segments 587-1, 587-2) that do not correspond to the UPC. The user may then select an affordance 584 to remove the selected tags from the corpus. This process may disassociate the UPC in question from the selected segments. In this way, the corpus of segments for the UPC in question may be updated so that it does not include incorrect segments that may negatively affect or impact the effectiveness of the models that are trained using the corpus”). Regarding claim 11, Talbot, in view of van Horne, further teaches a non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform operations of claim 1 (Talbot ¶¶0004: “a mobile device comprising a processor, a memory, a display, and an integrated camera”). Therefore, claim 11 is rejected using the same rationale as applied to claim 1 discussed above. Claim 15 is rejected using the same rationale as applied to claim 5 discussed above. Claim 16 is rejected using the same rationale as applied to claim 6 discussed above. Claim 17 is rejected using the same rationale as applied to claim 7 discussed above. Regarding claim 19, Talbot, in view of van Horne, further teaches a computer system comprising: a computer processor; and a non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform operations of claim 1 (Talbot ¶¶0004 discussed above). Therefore, claim 19 is rejected using the same rationale as applied to claim 1 discussed above. Claim(s) 2, 3, 12, and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Talbot et al. (US 2020/0219043 A1), in view of van Horne et al. (US 2022/0114648 A1), and further in view of Dhar et al. (US 2022/0327965 A1), hereinafter referred to as Talbot, van Horne, and Dhar, respectively. Regarding claim 2, Talbot, in view of van Horne, teaches the method of claim 1, wherein the one or more images are images of items on retail shelf (Talbot Fig. 7). However, Talbot, in view of van Horne, does not appear to explicitly teach that the images are images of items on a checkout belt of a retailer. Pertaining to the same field of endeavor, Dhar teaches that the images are images of items on a checkout belt of a retailer (Dhar ¶¶0053: “assume that all movable objects 402 that are purchasable within the venue 400 include a customized image (e.g., customized image 104) that is encoded with an encrypted payload, as described herein. A customer 404 may gather multiple objects as the customer 404 traverses the venue 400, and may attempt to checkout of the grocery store by placing all gathered objects 406 onto a checkout aisle conveyor belt 406. The gathered objects 406 may proceed down the conveyor belt 406 until the objects 406 reach a checkout clerk 410, who may scan the gathered objects 406 by capturing an image of the customized image associated with each respective gathered object 406”). Talbot, in view of van Horne, and Dhar are considered to be analogous art because they are directed to image processing for recognizing objects. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the technique for identifying products within an image captured using a terminal device for fulfilling users’ orders (as taught by Talbot, in view of van Horne) to image the items on the checkout belt (as taught by Dhar) because the combination allows the cashier to scan all of the items to be purchased and further enables prevention of ticket switching (Dhar ¶¶0053-¶¶0054). Regarding claim 3, Talbot, in view of van Horne and Dhar, teaches the method of claim 2, further comprising: identifying a first image region of the captured one or more images, the first image region corresponding to an order of a first user (Talbot Fig. 5 & ¶¶0182 discussed above; Talbot ¶¶0126: “The accompanying data file may include information such as the location or region within the image file where each physical item is shown, as well as the product information for each physical item”); wherein the machine learning model classifies items in the first image region to the one or more products, and wherein the method further comprises highlighting the classified items in the first image region which do not match the first user’s order to produce the annotated image for the first user’s order (Talbot Fig. 5 & ¶¶0182, ¶¶0188 discussed above). Claim 12 is rejected using the same rationale as applied to claim 2 discussed above. Claim 13 is rejected using the same rationale as applied to claim 3 discussed above. Claim(s) 8 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Talbot et al. (US 2020/0219043 A1), in view of van Horne et al. (US 2022/0114648 A1), and further in view of Adapa et al. (US 2023/0297947 A1), hereinafter referred to as Talbot, van Horne, and Adapa, respectively. Regarding claim 8, Talbot, in view of van Horne, teaches the method of claim 7, wherein the action comprises at least one of: providing visual feedback to a device of the picker to indicate a potential error associated with the user’s order (Talbot Fig. 5 & ¶¶0182, ¶¶0188 discussed above); and providing to the client device of the picker another prompt to capture an image of a checkout receipt of the user’s order (van Horne ¶¶0066: “The user interface 615 prompts the picker 108 to select one of the interactive elements 617 to either retake the captured image, which leads back to the user interface 612 shown in FIG. 6C, add another image of the receipt, submit the captured image 616, or cancel the uploading the captured image 616 of the receipt”). However, Talbot, in view of van Horne, does not appear to explicitly teach providing haptic feedback. Pertaining to the same field of endeavor, Adapa teaches providing haptic feedback (Adapa ¶¶0078: “Herein, the alert could be in the form of at least one of: a text, an image, an audio, a light signal, a haptic signal and so forth”). Talbot, in view of van Horne, and Adapa are considered to be analogous art because they are directed to image processing for recognizing objects. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the technique for identifying products within an image captured using a terminal device for fulfilling users’ orders (as taught by Talbot, in view of van Horne) to provide haptic feedback (as taught by Adapa) because the combination is more accessible (Adapa ¶¶0078). Claim 18 is rejected using the same rationale as applied to claim 8 discussed above. Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Talbot et al. (US 2020/0219043 A1), in view of van Horne et al. (US 2022/0114648 A1), and further in view of Gordon-Carroll et al. (US 2017/0355076 A1), hereinafter referred to as Talbot, van Horne, and Gordon-Carroll, respectively. Regarding claim 20, Talbot, in view of van Horne, teaches the computer system of claim 19, wherein the one or more images are images captured at the retail store (Talbot Fig. 7). However, Talbot, in view of van Horne, does not appear to teach or suggest that the images are captured at a delivery drop-off location associated with the user. Pertaining to the same field of endeavor, Gordon-Carroll teaches that the images are captured at a delivery drop-off location associated with the user (Gordon-Carroll ¶¶0018: “The system and method may further include capturing at least one image of the person placing the package in the first location and including the at least one image in the request to transport the package to the drop-off location”; Gordon-Carroll ¶¶0070: “action instruction deriving module 210 may include one or more images associated with the package (e.g., identity of the delivery person, image of the package, image of the location where the package is left by the delivery person) in the request to transport the package.”). Talbot, in view of van Horne, and Gordon-Carroll are considered to be analogous art because they are directed to image processing for recognizing objects. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the technique for identifying products within an image captured using a terminal device for fulfilling users’ orders (as taught by Talbot, in view of van Horne) to take an image at the delivery drop-off location (as taught by Gordon-Carroll) because the combination allows the user to confirm the delivery (Gordon-Carroll ¶¶0072). Allowable Subject Matter Claims 4, 9, and 14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding claim 4, the prior art of record teaches that it was known at the time the application was filed to use the method of claim 3, but does not appear to teach or suggest that the one or more images include a second image region separate from the first image region, the second image region corresponding to an order of a second user. Regarding claim 9, the prior art of record teaches that it was known at the time the application was filed to use the method of claim 1, but does not appear to teach or suggest: for the at least one of the one or more users, identifying an estimated risk of error in fulfilling the user’s order; and providing the prompt to capture the one or more images to the client device of the picker in response to the estimated risk of the error being higher than a threshold. Claim 14 is objected to for the same reason as claim 4 discussed above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOO J SHIN whose telephone number is (571)272-9753. The examiner can normally be reached M-F; 10-6. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella can be reached at (571)272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Soo Shin/Primary Examiner, Art Unit 2667 571-272-9753 soo.shin@uspto.gov
Read full office action

Prosecution Timeline

Sep 13, 2024
Application Filed
Jul 16, 2026
Non-Final Rejection mailed — §101, §103, §Other (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705917
AMBIGUITY RESOLUTION FOR OBJECT SELECTION AND FASTER APPLICATION LOADING FOR CLUTTERED SCENARIOS
3y 1m to grant Granted Aug 11, 2026
Patent 12705764
ELECTRONIC DEVICE FOR ESTIMATING OPTICAL FLOW AND OPERATING METHOD THEREOF
2y 9m to grant Granted Aug 11, 2026
Patent 12694661
EXPLAINABLE VISUAL ATTENTION FOR DEEP LEARNING
2y 9m to grant Granted Jul 28, 2026
Patent 12694662
REINFORCEMENT LEARNING AGENT TO MEASURE ROBUSTNESS OF BLACK-BOX IMAGE CLASSIFICATION MODELS
2y 9m to grant Granted Jul 28, 2026
Patent 12688927
IMAGE FEATURE CLASSIFICATION
3y 2m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
87%
Grant Probability
99%
With Interview (+16.3%)
2y 2m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 620 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month