Prosecution Insights
Last updated: August 17, 2026
Application No. 18/976,916

Dynamic Triggering and Processing of a Purchase Based on Computer Detection of Media Object

Non-Final OA §DP
Filed
Dec 11, 2024
Priority
Apr 15, 2022 — continuation of 12/198,434
Examiner
BHATNAGAR, ANAND P
Art Unit
Tech Center
Assignee
Roku Inc.
OA Round
1 (Non-Final)
91%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 91% — above average
91%
Career Allowance Rate
662 granted / 724 resolved
+31.4% vs TC avg
Minimal +2% lift
Without
With
+2.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
18 currently pending
Career history
740
Total Applications
across all art units

Statute-Specific Performance

§101
21.1%
-18.9% vs TC avg
§103
29.0%
-11.0% vs TC avg
§102
32.6%
-7.4% vs TC avg
§112
6.9%
-33.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 724 resolved cases

Office Action

§DP
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections 2. Claims 19 and 20 are objected to because of the following informalities: These claims are “non-transitory computer-readable medium” claims but are dependent from claim 17 which is a “system” claim which is improper. Appropriate correction is required. Claims 19 and 20 will be treated as dependent on claim 18. Double Patenting 3. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Application 18/976,916 U.S. patent 12,198,434 B2 1. A method for processing a purchase based on image recognition in a video stream being presented by a computing system, the method comprising: receiving, by the computing system, a first user-input defining a first user-request to pause presentation of the video stream, and, responsive to the first user-input, pausing by the computing system the presentation of the video stream at a video frame; responsive to the pausing, detecting, by the computing system, based on computer-vision analysis of the video frame, at least one object depicted by the video frame, wherein detecting based on computer-vision analysis of the video frame the object depicted by the video frame comprises (a) determining an identity of the video stream being presented and (b) based on data correlating the identity of the video stream being presented with a set of coordinates of the object in the video frame, receiving, from data storage, the set of coordinates of the object in the video frame; responsive to the detecting, (i) correlating, by the computing system, the detected object with at least one purchasable item and (ii) presenting, by the computing system, a prompt for purchase of the at least one purchasable item; receiving, by the computing system, in response to presenting the prompt, a second user- input requesting to purchase a given one of the at least one purchasable item; and processing, by the computing system, responsive to receiving the second user-input, a purchase of the given purchasable item for the user. 1. A method for processing a purchase based on image recognition in a video stream being presented by a computing system, the method comprising: receiving, by the computing system, a first user-input defining a first user-request to pause presentation of the video stream, and, responsive to the first user-input, pausing by the computing system the presentation of the video stream at a video frame; responsive to the pausing, detecting, by the computing system, based on computer-vision analysis of the video frame, at least one object depicted by the video frame, wherein detecting based on computer-vision analysis of the video frame the object depicted by the video frame comprises (a) determining an identity of the video stream being presented and (b) based on data correlating the identity of the video stream being presented with a set of coordinates of the object in the video frame, receiving, from data storage, the set of coordinates of the object in the video frame, wherein the set of coordinates was determined based on applying a pre-trained machine-learning model to the video frame of the video stream; responsive to the detecting, (i) correlating, by the computing system, the detected object with at least one purchasable item and (ii) presenting, by the computing system, a prompt for purchase of the at least one purchasable item; receiving, by the computing system, in response to presenting the prompt, a second user-input requesting to purchase a given one of the at least one purchasable item; and processing, by the computing system, responsive to receiving the second user-input, a purchase of the given purchasable item for the user. 2. The method of claim 1, wherein detecting, based on computer-vision analysis of the video frame, the at least one object depicted by the video frame comprises: detecting, based on computer-vision analysis of the video frame, a plurality of objects depicted by the video frame. 2. The method of claim 1, wherein detecting, based on computer-vision analysis of the video frame, the at least one object depicted by the video frame comprises: detecting, based on computer-vision analysis of the video frame, a plurality of objects depicted by the video frame. 3. The method of claim 1, wherein detecting based on computer-vision analysis of the video frame, the object depicted by the video frame comprises: determining the set of coordinates of the object in the video frame. 3. The method of claim 1, wherein detecting based on computer-vision analysis of the video frame, the object depicted by the video frame comprises: determining, based on applying the pre-trained machine-learning model to the video frame, the set of coordinates of the object in the video frame. 4. The method of claim 1, wherein the method further comprises: before receiving the first user-input defining the first user-request to pause the presentation of the video stream, engaging in an object detection process including (i) determining the set of coordinates of the object in the video frame, and (ii) storing, in the data storage, the set of coordinates of the object in the video frame. 4. The method of claim 1, wherein the method further comprises: before receiving the first user-input defining the first user-request to pause the presentation of the video stream, engaging in an object detection process including (i) determining, based on applying the pre-trained machine-learning model to the video frame, the set of coordinates of the object in the video frame, and (ii) storing, in the data storage, the set of coordinates of the object in the video frame. 5. The method of claim 4, wherein the method further comprises: determining that the video stream has been presented at least a predefined threshold number of times, wherein engaging in the object detection process is responsive to the determining that the video stream has been presented at least the predefined threshold number of times. 5. The method of claim 4, wherein the method further comprises: determining that the video stream has been presented at least a predefined threshold number of times, wherein engaging in the object detection process is responsive to the determining that the video stream has been presented at least the predefined threshold number of times. 6. The method of claim 1, wherein the set of coordinates defines a location of the object within the video frame, wherein correlating the object with at least one purchasable item comprises: extracting an image region of the video frame based on the set of coordinates; determining, based on the extracted image region, a feature value representative of the extracted image region; accessing a plurality of feature values each representative of a respective stored purchasable-item image of a plurality of stored object images; determining, based on the feature value of the extracted image region and each of the plurality of feature values, a plurality of similarity values; and selecting, based on the plurality of values, at least one stored object image from the plurality of stored object images based on the selected purchasable item having a highest similarity value, wherein the selected at least one stored object image corresponds to the at least one purchasable item. 6. The method of claim 1, wherein the set of coordinates defines a location of the object within the video frame, wherein correlating the object with at least one purchasable item comprises: extracting an image region of the video frame based on the set of coordinates; determining, based on applying a further pre-trained machine-learning model to the extracted image region, a feature value representative of the extracted image region; accessing a plurality of feature values each representative of a respective stored purchasable-item image of a plurality of stored object images; determining, based on the feature value of the extracted image region and each of the plurality of feature values, a plurality of similarity values; and selecting, based on the plurality of values, at least one stored object image from the plurality of stored object images based on the selected purchasable item having a highest similarity value, wherein the selected at least one stored object image corresponds to the at least one purchasable item. 7. The method of claim 6, wherein the method further comprises: determining, based on the plurality of stored object images, the plurality of feature values. 7. The method of claim 6, wherein the method further comprises: determining, by applying the further pre-trained machine-learning model to the plurality of stored object images, the plurality of feature values. 8. The method of claim 1, wherein the at least one purchasable item is a plurality of purchasable items, wherein each purchasable item is from a different vendor, wherein presenting the prompt for purchase comprises listing the plurality of purchasable items in the prompt as user- selectable options for purchase. 9. The method of claim 1, wherein the at least one purchasable item is a plurality of purchasable items, wherein each purchasable item is from a different vendor, wherein presenting the prompt for purchase comprises listing the plurality of purchasable items in the prompt as user-selectable options for purchase. 9. The method of claim 1, wherein presenting the prompt for purchase of the at least one purchasable item comprises: superimposing, in the video frame, (i) a bounding box at the set of coordinates within the video frame, and (ii) a prompt for purchase of the at least one purchasable item. 10. The method of claim 1, wherein presenting the prompt for purchase of the at least one purchasable item comprises: superimposing, in the video frame, (i) a bounding box at the set of coordinates within the video frame, and (ii) a prompt for purchase of the at least one purchasable item. 10. The method of claim 1, wherein presenting the prompt for purchase of the at least one purchasable item comprises: superimposing, in the video frame, a listing of the at least one purchasable item. 11. The method of claim 1, wherein presenting the prompt for purchase of the at least one purchasable item comprises: superimposing, in the video frame, a listing of the at least one purchasable item. 11. The method of claim 1, wherein correlating the detected object with the at least one purchasable item is based on a profile of the user. 12. The method of claim 1, wherein correlating the detected object with the at least one purchasable item is based on a profile of the user. 12. The method of claim 1, wherein correlating the detected object with at least one purchasable item is based on a price of each of the at least one purchasable item. 13. The method of claim 1, wherein correlating the detected object with at least one purchasable item is based on a price of each of the at least one purchasable item. 13. The method of claim 1, wherein the computing system is a provider of the video stream to the user, and wherein processing the purchase of the purchasable item for the user comprises transmitting, from the computing system to a vendor associated with the purchasable item, a purchase request, wherein the purchase request includes user-payment information. 14. The method of claim 1, wherein the computing system is a provider of the video stream to the user, and wherein processing the purchase of the purchasable item for the user comprises transmitting, from the computing system to a vendor associated with the purchasable item, a purchase request, wherein the purchase request includes user-payment information. 14. A computing system comprising: a network communication interface; one or more processors; non-transitory data storage; and program instructions stored in the non-transitory data storage and executable by the one or more processors to carry out operations including: receiving a first user-input defining a first user-request to pause presentation of a video stream, and, responsive to the first user-input, pausing the presentation of the video stream at a video frame, responsive to the pausing, detecting based on computer-vision analysis of the video frame, at least one object depicted by the video frame, wherein detecting based on computer-vision analysis of the video frame the object depicted by the video frame comprises (a) determining an identity of the video stream being presented and (b) based on data correlating the identity of the video stream being presented with a set of coordinates of the object in the video frame, receiving, from the non-transitory data storage, the set of coordinates of the object in the video frame, responsive to the detecting, (i) correlating the detected object with at least one purchasable item and (ii) presenting a prompt for purchase of the at least one purchasable item, receiving, in response to presenting the prompt, a second user-input requesting to purchase a given one of the at least one purchasable item, and processing, responsive to receiving the second user-input, a purchase of the given purchasable item for the user. 15. A computing system comprising: a network communication interface; one or more processors; non-transitory data storage; and program instructions stored in the non-transitory data storage and executable by the one or more processors to carry out operations including: receiving a first user-input defining a first user-request to pause presentation of a video stream, and, responsive to the first user-input, pausing the presentation of the video stream at a video frame; responsive to the pausing, detecting based on computer-vision analysis of the video frame, at least one object depicted by the video frame, wherein detecting based on computer-vision analysis of the video frame the object depicted by the video frame comprises (a) determining an identity of the video stream being presented and (b) based on data correlating the identity of the video stream being presented with a set of coordinates of the object in the video frame, receiving, from the non-transitory data storage, the set of coordinates of the object in the video frame, wherein the set of coordinates was determined based on applying a pre-trained machine-learning model to the video frame of the video stream; responsive to the detecting, (i) correlating the detected object with at least one purchasable item and (ii) presenting a prompt for purchase of the at least one purchasable item; receiving, in response to presenting the prompt, a second user-input requesting to purchase a given one of the at least one purchasable item; and processing, responsive to receiving the second user-input, a purchase of the given purchasable item for the user. 15. The computing system of claim 14, wherein detecting, based on computer-vision analysis of the video frame, the at least one object depicted by the video frame comprises detecting, based on computer-vision analysis of the video frame, a plurality of objects depicted by the video frame. 16. The computing system of claim 15, wherein detecting, based on computer-vision analysis of the video frame, the at least one object depicted by the video frame comprises detecting, based on computer-vision analysis of the video frame, a plurality of objects depicted by the video frame. 16. The computing system of claim 14, wherein detecting based on computer-vision analysis of the video frame, the object depicted by the video frame comprises: determining the set of coordinates of the object in the video frame. 17. The computing system of claim 15, wherein detecting based on computer-vision analysis of the video frame, the object depicted by the video frame comprises: determining, based on applying the pre-trained machine-learning model to the video frame, the set of coordinates of the object in the video frame. 18. A non-transitory computer-readable medium having stored thereon program instructions executable by one or more processors to cause a media presentation system to carry out operations including: receiving a first user-input defining a first user-request to pause presentation of a video stream, and, responsive to the first user-input, pausing the presentation of the video stream at a video frame; responsive to the pausing, detecting based on computer-vision analysis of the video frame, at least one object depicted by the video frame, wherein detecting based on computer-vision analysis of the video frame the object depicted by the video frame comprises (a) determining an identity of the video stream being presented and (b) based on data correlating the identity of the video stream being presented with a set of coordinates of the object in the video frame, receiving, from data storage, the set of coordinates of the object in the video frame; responsive to the detecting, (i) correlating the detected object with at least one purchasable item and (ii) presenting a prompt for purchase of the at least one purchasable item; receiving in response to presenting the prompt, a second user-input requesting to purchase a given one of the at least one purchasable item; and processing responsive to receiving the second user-input, a purchase of the given purchasable item for the user. 18. A non-transitory computer-readable medium having stored thereon program instructions executable by one or more processors to cause a media presentation system to carry out operations including: receiving a first user-input defining a first user-request to pause presentation of a video stream, and, responsive to the first user-input, pausing the presentation of the video stream at a video frame; responsive to the pausing, detecting based on computer-vision analysis of the video frame, at least one object depicted by the video frame, wherein detecting based on computer-vision analysis of the video frame the object depicted by the video frame comprises (a) determining an identity of the video stream being presented and (b) based on data correlating the identity of the video stream being presented with a set of coordinates of the object in the video frame, receiving, from data storage, the set of coordinates of the object in the video frame, wherein the set of coordinates was determined based on applying a pre-trained machine-learning model to the video frame of the video stream; responsive to the detecting, (i) correlating the detected object with at least one purchasable item and (ii) presenting a prompt for purchase of the at least one purchasable item; receiving in response to presenting the prompt, a second user-input requesting to purchase a given one of the at least one purchasable item; and processing responsive to receiving the second user-input, a purchase of the given purchasable item for the user. 19. The non-transitory computer-readable medium of claim 17, wherein detecting, based on computer-vision analysis of the video frame, the at least one object depicted by the video frame comprises detecting, based on computer-vision analysis of the video frame, a plurality of objects depicted by the video frame. 19. The non-transitory computer-readable medium of claim 17, wherein detecting, based on computer-vision analysis of the video frame, the at least one object depicted by the video frame comprises detecting, based on computer-vision analysis of the video frame, a plurality of objects depicted by the video frame. Claim 1 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of U.S. Patent No. 12,198,434 B2. Claim 2 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 2 of U.S. Patent No. 12,198,434 B2. Claim 3 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 3 of U.S. Patent No. 12,198,434 B2. Claim 4 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 4 of U.S. Patent No. 12,198,434 B2. Claim 5 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 5 of U.S. Patent No. 12,198,434 B2. Claim 6 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 6 of U.S. Patent No. 12,198,434 B2. Claim 7 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 7 of U.S. Patent No. 12,198,434 B2. Claim 8 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 9 of U.S. Patent No. 12,198,434 B2. Claim 9 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 10 of U.S. Patent No. 12,198,434 B2. Claim 10 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 11 of U.S. Patent No. 12,198,434 B2. Claim 11 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 12 of U.S. Patent No. 12,198,434 B2. Claim 12 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 13 of U.S. Patent No. 12,198,434 B2. Claim 13 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 14 of U.S. Patent No. 12,198,434 B2. Claim 14 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 15 of U.S. Patent No. 12,198,434 B2. Claim 15 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 16 of U.S. Patent No. 12,198,434 B2. Claim 16 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 17 of U.S. Patent No. 12,198,434 B2. Claim 18 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 18 of U.S. Patent No. 12,198,434 B2. Claim 19 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 19 of U.S. Patent No. 12,198,434 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because the scope of the claims of this instant invention are encompassed by the patented claims. Regarding claims 1-20: There is no prior are made on these claims since prior art was not found on the claimed subject matter. Allowable Subject Matter 4. Claims 17 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Contact Information 5. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANAND BHATNAGAR whose telephone number is (571)272-7416. The examiner can normally be reached on M-F 7:30am-4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached on 571-272-4650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANAND P BHATNAGAR/ Primary Examiner, Art Unit 2668 July 25, 2026
Read full office action

Prosecution Timeline

Dec 11, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694473
ELECTRONIC DEVICE AND METHOD WITH IMAGE PROCESSING.
2y 9m to grant Granted Jul 28, 2026
Patent 12676973
IMAGE DECODING DEVICE, IMAGE DECODING METHOD, AND PROGRAM
2y 4m to grant Granted Jul 07, 2026
Patent 12664661
IMAGE PROCESSING APPARATUS AND IMAGE PROCESSING METHOD
2y 0m to grant Granted Jun 23, 2026
Patent 12657704
TIME PHASE DETERMINATION APPARATUS AND TIME PHASE DETERMINATION METHOD
3y 2m to grant Granted Jun 16, 2026
Patent 12657779
DATA PROCESSING METHOD AND APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM
2y 6m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
91%
Grant Probability
94%
With Interview (+2.3%)
2y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 724 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month