Prosecution Insights
Last updated: August 17, 2026
Application No. 18/940,461

REGION-TEXT CAPTION GENERATION USING GLOBAL CAPTION INFORMATION

Non-Final OA §101§102§103§112
Filed
Nov 07, 2024
Examiner
ELLIOTT, JORDAN MCKENZIE
Art Unit
2666
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
41%
Grant Probability
Moderate
1-2
OA Rounds
1y 2m
Est. Remaining
15%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
11 granted / 27 resolved
-21.3% vs TC avg
Minimal -26% lift
Without
With
+-25.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
23 currently pending
Career history
67
Total Applications
across all art units

Statute-Specific Performance

§101
8.6%
-31.4% vs TC avg
§103
53.0%
+13.0% vs TC avg
§102
25.4%
-14.6% vs TC avg
§112
13.1%
-26.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION Claims 1-20 are pending in this application. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 11/07/2024 and 05/08/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: Processing unit in claims 17 and 19. A system for performing simulation operations of claims 9 and 20. A system for performing simulation operations to test or validate autonomous machine learning applications of claims 9 and 20. a system for performing digital twin operations of claims 9 and 20. a system for performing light transport simulation of claims 9 and 20. a system for rendering graphical output of claims 9 and 20. a system for performing deep learning operations of claims 9 and 20. a system implemented using an edge device of claims 9 and 20. a system for generating or presenting virtual reality (VR) content of claims 9 and 20. a system for generating or presenting augmented reality (AR) content of claims 9 and 20. a system for generating or presenting mixed reality (MR) content of claims 9 and 20. a system incorporating one or more Virtual Machines (VMs) of claims 9 and 20. a system for performing operations for a conversational AI application of claims 9 and 20. a system for performing operations for a generative AI application of claims 9 and 20. a system for performing operations using a language model of claims 9 and 20. a system for performing one or more operations using a large language model (LLM) of claims 9 and 20. a system for performing one or more operations using a vision language model (VLM) of claims 9 and 20. a system implemented at least partially in a data center of claims 9 and 20. a system for performing hardware testing using simulation of claims 9 and 20. a system for synthetic data generation of claims 9 and 20. a collaborative content creation platform for 3D assets of claims 9 and 20. a system implemented at least partially using cloud computing resources of claims 9 and 20. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 17 and 19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as failing to set forth the subject matter which the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the applicant regards as the invention. Regarding claim 17 and 19, the claims recite the use of “the one or more processing units”. However, the component of the one or more processing units are not introduced in any claims from which 17 and 19 depend. Therefore, it is unclear how a component which has not been established and defined as part of the system may be called on to perform an action. Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 as being drawn to an abstract idea or mental process without significantly more. Regarding claim 1, the claim is drawn to an abstract idea, mental process or step of mere data gathering. Claims 1 recites the following limitations; A processor, comprising: one or more circuits to: (Additional elements) generate an object list for an input image based on the input image and a first caption corresponding to the input image; (Step of mere data gathering, which could be reasonably accomplished by a human assessing an image and creating a list of objects in it) generate one or more bounding boxes for objects at least partially depicted in the input image; (step of data gathering/manipulation which a human could perform manually) generate a first selected caption for a selected bounding box of the one or more bounding boxes; (step of data gathering/manipulation which a human could perform manually) determine an object, within the selected bounding box, based on the first selected caption; (mental process where a human could view and object and generate a caption) generate a second selected caption, based on the first caption and the selected bounding box; (mental process where a human could view and object and generate a caption) and generate a merged caption based on the first selected caption, the second selected caption, and the object. (mental process where a human could view and object and generate a caption) Under step 2A prong 1, the limitations recited are drawn to abstract ideas, mental processes or steps of mere data gathering as noted above by the examiner. Further, under step 2A prong 2, the claim recites the additional elements of a processor and one or more circuits, which neither constitute judicial exceptions nor integrate the claim into practical application. Further, under step 2B, the claim does not include any additional elements which translate the claim into practical application or amount to significantly more than an abstract idea. (See MPEP section 2106) Dependent claims 2-9 do not add limitations that meaningfully translate the abstract idea into practical application or add significantly more. Regarding claim 2, claim 2 recites the following limitations; “wherein at least one of the first caption, the first selected caption, or the second selected caption, are generated using a vision language model (Mere data gathering/generation)” The limitations are drawn to a step of mere data gathering without significantly more. Further, the additional element recited of a “vision language model” is recited with a high level of generality and does not translate the claim into practical application. Regarding claim 3, claim 3 recites the following limitations; “wherein the one or more circuits are further to: generate a first object list from the input image; (step of mere data gathering) generate a second object list from the first caption; (step of mere data gathering) and combine the first object list and the second object list to form the object list. (step of mere data gathering/manipulation) ” The limitations are drawn to steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional element recited of one or more circuits is recited with a high level of generality and does not translate the claim into practical application. Regarding claim 4, claim 4 recites the following limitations; “wherein the one or more circuits are further to: provide the object list to one or more trained machine learning systems to generate the one or more bounding boxes. (abstract idea in which data is arbitrarily passed/provided to other system parts)” The limitations are drawn to abstract ideas or steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional elements recited of one or more circuits and one or more machine learning systems are recited with a high level of generality and do not translate the claim into practical application. Regarding claim 5, claim 5 recites the following limitations; “wherein the one or more circuits are further to: identify a plurality of bounding boxes, of the one or more bounding boxes, having a common label; (Mental process of assessing data visually) determine at least a portion of the plurality of bounding boxes are within a threshold distance; (mental process of decision making from assessing data) and combine the portion of the plurality of bounding boxes within a single bounding box. (step of mere data gathering/manipulation) ” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional elements recited of one or more circuits are recited with a high level of generality and do not translate the claim into practical application. Regarding claim 6, claim 6 recites the following limitations; “wherein the one or more circuits are further to: determine a first label associated with the selected bounding box (Mental process of assessing an image visually); determine a second label associated with the object (Mental process of assessing an image visually); determine a similarity metric between the first label and the second label is below a threshold; (Math which could be determined manually by a human) and identify the merged caption for review. (Mental process of determination) ” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional elements recited of one or more circuits are recited with a high level of generality and do not translate the claim into practical application. Regarding claim 7, claim 7 recites the following limitations; “wherein an input to a trained machine learning system used to generate the first caption includes raw caption data for the input image. (Data gathering/transmission) ” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional element recited of a trained machine learning is recited with a high level of generality and does not translate the claim into practical application. Regarding claim 8, claim 8 recites the following limitations; “wherein the one or more circuits are further to: receive one or more prompts associated with an output configuration for at least one of the first caption, the first selected caption, or the second selected caption. (Data gathering/transmission) ” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional elements recited of one or more circuits are recited with a high level of generality and do not translate the claim into practical application. Regarding claim 9, the claim recites the following limitations; “wherein the processor is comprised in at least one of: A system for performing simulation operations; (Data gathering/transmission) A system for performing simulation operations to test or validate autonomous machine learning applications; (Data gathering/transmission) a system for performing digital twin operations; (Data gathering/transmission) a system for performing light transport simulation; (Data gathering/transmission) a system for rendering graphical output; (Data gathering/transmission) a system for performing deep learning operations; (Data gathering/transmission) a system implemented using an edge device; (Data gathering/transmission) a system for generating or presenting virtual reality (VR) content; (Data gathering/transmission) a system for generating or presenting augmented reality (AR) content; (Data gathering/transmission) a system for generating or presenting mixed reality (MR) content; (Data gathering/transmission) a system incorporating one or more Virtual Machines (VMs); (Data gathering/transmission) a system for performing operations for a conversational AI application; (Data gathering/transmission) a system for performing operations for a generative AI application; (Data gathering/transmission) a system for performing operations using a language model; (Data gathering/transmission) a system for performing one or more operations using a large language model (LLM); (Data gathering/transmission) a system for performing one or more operations using a vision language model (VLM); (Data gathering/transmission) a system implemented at least partially in a data center; (Data gathering/transmission) a system for performing hardware testing using simulation; (Data gathering/transmission) a system for performing one or more generative content operations using a language model; (Data gathering/transmission) a system for synthetic data generation; (Data gathering/transmission) a collaborative content creation platform for 3D assets; (Data gathering/transmission) or a system implemented at least partially using cloud computing resources. (Data gathering/transmission)” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional elements recited of a processor and multiple systems for processing the data, multiple models for processing the data and a collaborative content creation platform are recited with a high level of generality and do not translate the claim into practical application. Regarding claim 10, Claim 10 recites the following limitations; A computer-implemented method, comprising: obtaining a set of bounding boxes and labels for an image based on an object list that includes one or more identified objects at least partially depicted in the image and one or more described objects from a first caption of the image; (Step of mere data gathering, which could be reasonably accomplished by a human assessing an image and creating a list of objects in it) determining, for a selected bounding box of the set of bounding boxes, a second caption; (step of data gathering/manipulation which a human could perform manually) determining, from the second caption, a bounding box object; (step of data gathering/manipulation which a human could perform manually) determining, for the selected bounding box, a third caption, based on the first caption; (mental process where a human could view and object and generate a caption) and generating a fourth caption based on the second caption, the third caption, and the bounding box object. (mental process where a human could view and object and generate a caption) Under step 2A prong 1, the limitations recited are drawn to abstract ideas, mental processes or steps of mere data gathering as noted above by the examiner. Further, under step 2A prong 2, the claim does not recite any additional elements which constitute judicial exceptions nor integrate the claim into practical application. (See MPEP section 2106) Dependent claims 11- 15 do not add limitations that meaningfully translate the abstract idea into practical application or add significantly more. Regarding claim 11, claim 11 recites the following limitations; “generating, using a raw caption associated with the image, the first caption; (Mere data gathering/generation) and generating a first object list based on the first caption. (Mere data gathering/generation)” The limitations are drawn to a step of mere data gathering without significantly more. Further, the claim does not recite additional elements that translate the claim into practical application. Regarding claim 12, claim 12 recites the following limitations; “further comprising: identifying the one or more identified objects in the image using a trained machine learning model; (step of mere data gathering) generating, a second object list based on the one or more identified objects; (step of mere data gathering) and generating the object list based on the first object list and the second object list. (step of mere data gathering/manipulation) ” The limitations are drawn to steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional element recited of a trained machine learning model is recited with a high level of generality and does not translate the claim into practical application. Regarding claim 13, claim 13 recites the following limitations; “determining a plurality of bounding boxes, of the set of bounding boxes, have a common label; (Mental process of assessing data visually) determining at least a portion of the plurality of bounding boxes are within a threshold distance; (mental process of decision making from assessing data) and combining the portion of the plurality of bounding boxes within a single bounding box with the common label. (step of mere data gathering/manipulation) ” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the claim does not recite additional elements that translate the claim into practical application. Regarding claim 14, claim 14 recites the following limitations; “receiving a prompt corresponding to an output format for the fourth caption. (Data gathering/transmission) ” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the claim does not recite additional elements that translate the claim into practical application. Regarding claim 15, claim 15 recites the following limitations; “further comprising: determining a first label associated with the selected bounding box; (Mental process of assessing an image visually) determining a second label associated with the bounding box object; (Mental process of assessing an image visually) determining a similarity metric between the first label and the second image label is below a threshold; (Math which could be determined manually by a human) and identifying the fourth caption for review. (Mental process of determination) ” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the claim does not recite additional elements that translate the claim into practical application. Regarding claim 16, Claim 16 recites the following limitations; “A system, comprising: processing circuitry to generate a caption for an input image based on a set of region captions, (Step of mere data gathering, which could be reasonably accomplished by a human assessing an image and creating a list of objects in it) wherein individual region captions of the set of region captions are generated using a first region caption based on a labeled bounding box for an object within the image (step of data gathering/manipulation which a human could perform manually) and a second region caption based on a global description of the image.(step of data gathering/manipulation which a human could perform manually) Under step 2A prong 1, the limitations recited are drawn to abstract ideas, mental processes or steps of mere data gathering as noted above by the examiner. Further, under step 2A prong 2, the claim recites the additional element of processing circuitry, which neither constitutes judicial exception nor integrates the claim into practical application. Further, under step 2B, the claim does not include any additional elements which translate the claim into practical application or amount to significantly more than an abstract idea. (See MPEP section 2106) Dependent claims 17-20 do not add limitations that meaningfully translate the abstract idea into practical application or add significantly more. Regarding claim 17, claim 17 recites the following limitations; “where the one or more processing units are further to generate an object list for the input image based on a first object list corresponding to a first output for an object recognition model and a second object list corresponding to a second output for a large language model. (Mere data gathering/generation)” The limitations are drawn to a step of mere data gathering without significantly more. Further, the additional elements recited of “one or more processing units”, a “object recognition model” and a “large language model” is recited with a high level of generality and does not translate the claim into practical application. Regarding claim 18, claim 18 recites the following limitations; “wherein the object recognition model processes the input image and the large language model processes an image caption generated by a vision language model based on the input image and raw caption data. (Mere data gathering/generation)” The limitations are drawn to a step of mere data gathering without significantly more. Further, the additional elements recited of a “vision language model” and a “large language model” are recited with a high level of generality and do not translate the claim into practical application. Regarding claim 19, claim 19 recites the following limitations; “wherein the one or more processing units are further to provide the caption to a human-in-the-loop review engine responsive to determining a similarity metric between labeled objects in the input image is below a threshold. (Mere data gathering/generation)” The limitations are drawn to a step of mere data gathering without significantly more. Further, the additional elements recited of a “human in the loop review engine” and multiple processing units are recited with a high level of generality and do not translate the claim into practical application. Regarding claim 20, claim 20 recites the following limitations; “A system for performing simulation operations; (Data gathering/transmission) A system for performing simulation operations to test or validate autonomous machine learning applications; (Data gathering/transmission) a system for performing digital twin operations; (Data gathering/transmission) a system for performing light transport simulation; (Data gathering/transmission) a system for rendering graphical output; (Data gathering/transmission) a system for performing deep learning operations; (Data gathering/transmission) a system implemented using an edge device; (Data gathering/transmission) a system for generating or presenting virtual reality (VR) content; (Data gathering/transmission) a system for generating or presenting augmented reality (AR) content; (Data gathering/transmission) a system for generating or presenting mixed reality (MR) content; (Data gathering/transmission) a system incorporating one or more Virtual Machines (VMs); (Data gathering/transmission) a system for performing operations for a conversational AI application; (Data gathering/transmission) a system for performing operations for a generative AI application; (Data gathering/transmission) a system for performing operations using a language model; (Data gathering/transmission) a system for performing one or more operations using a large language model (LLM); (Data gathering/transmission) a system for performing one or more operations using a vision language model (VLM); (Data gathering/transmission) a system implemented at least partially in a data center; (Data gathering/transmission) a system for performing hardware testing using simulation; (Data gathering/transmission) a system for performing one or more generative content operations using a language model; (Data gathering/transmission) a system for synthetic data generation; (Data gathering/transmission) a collaborative content creation platform for 3D assets; (Data gathering/transmission) or a system implemented at least partially using cloud computing resources. (Data gathering/transmission)” The limitations are drawn to abstract ideas, mental processes, steps of mere data gathering or data manipulation which could reasonably be manually performed by a human without significantly more. Further, the additional elements recited of a processor and multiple systems for processing the data, multiple models for processing the data and a collaborative content creation platform are recited with a high level of generality and do not translate the claim into practical application Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-2, 6-11, 14-18 and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Xie (US 20230394855 A1). Regarding claim 1 Xie discloses; A processor, comprising: one or more circuits to (Xie,[0044] the system may be executed on a processor to execute the model): generate an object list for an input image based on the input image and a first caption corresponding to the input image (Xie, [0019] the model takes an input image in, and generates a first caption and set of object visual information, which includes a list of objects, [0020] the model generates one or more tags, which identify objects in the image based on the caption and input image, which is analogous to a list of objects); generate one or more bounding boxes for objects at least partially depicted in the input image (Xie, [0036] the captioner model shows the locations of objects depicted in the image in the form of bounding boxes, bounding box coordinates are given for each object); PNG media_image1.png 124 384 media_image1.png Greyscale (Xie, [0036]) generate a first selected caption for a selected bounding box of the one or more bounding boxes (Xie, [0039] local descriptions/captions (at least a first selected caption) are generated for each selected bounding box); PNG media_image2.png 146 372 media_image2.png Greyscale (Xie, [0039]) determine an object, within the selected bounding box, based on the first selected caption (Xie, [0029] the objects listed in the caption (first caption) are verified to determine whether they occur in the image for each semantic proposition (determining and object based on the first selected caption), [0020] where the semantic proposition includes a region (bounding box) for where the object occurs, where [0036] object locations are provided in the form of bounding boxes); generate a second selected caption, based on the first caption and the selected bounding box (Xie, [0020]-[0021] local descriptions/captions (first selected caption) for each object in the region (shown as visual information 126 in figure 2) are used to generate visual clues which include multiple descriptions/captions/written attributes about the image (second selected caption, shown as 130 in figure 2)); and generate a merged caption based on the first selected caption, the second selected caption, and the object (Xie, [0019] as shown in figure 1, an image is taken as input generating a variety of captions/tags for objects in the image, [0020]-[0021] as shown in figure 2, multiple caption options for the image and the objects within the image (first (Figure 2, 126) and second selected caption (figure 2, 130)) are generated based on the objects detected, and multiple caption candidates (merged captions) are generated using the visual information, which includes the first and second selected captions). PNG media_image3.png 524 872 media_image3.png Greyscale (Xie, figure 2, emphasis added) PNG media_image4.png 610 378 media_image4.png Greyscale (Xie [0019]-[0021]) Regarding claim 2 Xie discloses; The processor of claim 1 (Xie,[0044] the system may be executed on a processor to execute the model), wherein at least one of the first caption, the first selected caption, or the second selected caption, are generated using a vision language model (Xie, [0015] a first vision language model generates an initial caption for objects appearing in an image). Regarding claim 6 Xie discloses; The processor of claim 1 (Xie,[0044] the system may be executed on a processor to execute the model), wherein the one or more circuits are further to: determine a first label associated with the selected bounding box (Xie, [0039] the model generates visual information which includes local descriptions for each bounding box (first label)); determine a second label associated with the object (Xie, [0031] visual clues about the objects in the images are generated (second label)); determine a similarity metric between the first label and the second label is below a threshold (Xie, [0040]-[0041] the local descriptions (second label) and the visual clues (second label) are used to generate multiple image captions, [0042]-[0043] a threshold is used to determine similarity and determine if the similarity is below a threshold, it may be removed); and identify the merged caption for review (Xie, [0053]-[0055] the merged caption generated from the descriptions is scored and reviewed by the model, the review of the caption takes place based on the scoring, therefore it is identified for review based on this score/thresholding). Regarding claim 7 Xie discloses; The processor of claim 1 (Xie,[0044] the system may be executed on a processor to execute the model), wherein an input to a trained machine learning system used to generate the first caption includes raw caption data for the input image (Xie, [0020]-[0021] an initial caption (raw caption data) is provided to the model along with other semantic details about the image to generate story captions (first and second captions)). Regarding claim 8 Xie discloses; The processor of claim 1, wherein the one or more circuits are further to (Xie,[0044] the system may be executed on a processor to execute the model): receive one or more prompts associated with an output configuration for at least one of the first caption, the first selected caption, or the second selected caption (Xie, [0040] the visual information including an initial caption, generated local descriptions and visual clues are used as prompts for the generative language model to generate further captions). Regarding claim 9. The processor of claim 1, wherein the processor is comprised in at least one of (Xie,[0044] the system may be executed on a processor to execute the model): a system for performing simulation operations (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing simulation operations to test or validate autonomous machine learning applications (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing digital twin operations(Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing light transport simulation (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for rendering graphical output (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing deep learning operations (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system implemented using an edge device (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for generating or presenting virtual reality (VR) content (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for generating or presenting augmented reality (AR) content (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for generating or presenting mixed reality (MR) content (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system incorporating one or more Virtual Machines (VMs) (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing operations for a conversational AI application (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing operations for a generative AI application(Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing operations using a language model (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing one or more operations using a large language model (LLM) (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing one or more operations using a vision language model (VLM) (Xie, [0044] a vision language model maybe executed using a block-based image processor, [0063] the system is a computing system for executing multiple vision language models); a system implemented at least partially in a data center (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing hardware testing using simulation (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing one or more generative content operations using a language model (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for synthetic data generation (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a collaborative content creation platform for 3D assets (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); or a system implemented at least partially using cloud computing resources (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed). Regarding claim 10, Xie discloses; A computer-implemented method, comprising: obtaining a set of bounding boxes and labels for an image based on an object list that includes one or more identified objects at least partially depicted in the image and one or more described objects from a first caption of the image (Xie, [0019] the model takes an input image in, and generates a first caption and set of object visual information, which includes a list of objects in the caption, [0020] the model generates one or more tags, which identify objects in the image based on the caption and input image, which is analogous to a list of objects, [0036] the captioner model shows the locations of objects depicted in the image in the form of bounding boxes, bounding box coordinates are given for each object); determining, for a selected bounding box of the set of bounding boxes, a second caption (Xie, [0039] local descriptions/captions (at least a second caption) are generated for each selected bounding box); determining, from the second caption, a bounding box object (Xie, [0029] the objects listed in the caption (first caption) are verified to determine whether they occur in the image for each semantic proposition (second caption) which is analogous to determining and object based on the second caption, [0020] where the semantic proposition includes a region (bounding box) for where the object occurs, where [0036] object locations are provided in the form of bounding boxes); determining, for the selected bounding box, a third caption, based on the first caption (Xie, [0040] the collected visual information which includes the initial caption (first caption) are used to generate visual clues, which include a local descriptions/captions for each bounding box location (third caption)); and generating a fourth caption based on the second caption, the third caption, and the bounding box object (Xie [0020]-[0021] as shown in figure 2, multiple caption options for the image and the objects within the image (second (Figure 2, 126) and third caption (figure 2, 130)) are generated based on the objects detected, and multiple caption candidates (fourth caption (figure 2, 144) are generated from this information). Regarding claim 11, Xie discloses; The computer-implemented method of claim 10, further comprising: generating, using a raw caption associated with the image, the first caption (Xie, [0056]-[0057] an initial image story caption is generated (raw caption data) from objects detected in the image, from there object information, which includes a first generated caption is formed); and generating a first object list based on the first caption (Xie, [0057] a caption graph (object list) is generated from the caption based on objects in the caption). Regarding claim 14, Xie discloses; The computer-implemented method of claim 10, further comprising: receiving a prompt corresponding to an output format for the fourth caption (Xie, [0040] the visual information including an initial caption, generated local descriptions and visual clues are used as prompts for the generative language model to generate further captions). Regarding claim 15, Xie discloses; The computer-implemented method of claim 10, further comprising: determining a first label associated with the selected bounding box(Xie, [0039] the model generates visual information which includes local descriptions for each bounding box (first label)); determining a second label associated with the bounding box object(Xie, [0031] visual clues about the objects in the images are generated (second label)); determining a similarity metric between the first label and the second image label is below a threshold(Xie, [0040]-[0041] the local descriptions (second label) and the visual clues (second label) are used to generate multiple image captions, [0042]-[0043] a threshold is used to determine similarity and determine if the similarity is below a threshold, it may be removed); and identifying the fourth caption for review (Xie, [0053]-[0055] the merged caption generated from the descriptions is scored and reviewed by the model, the review of the caption takes place based on the scoring, therefore it is identified for review based on this score/thresholding). Regarding claim 16 Xie discloses; A system, comprising: processing circuitry to generate a caption for an input image based on a set of region captions (Xie, [0020]-[0021] as shown in figure 2, multiple caption options for the image and the objects within the image (Region captions (Figure 2, 126 and figure 2, 130)) are generated based on the objects detected, and multiple caption candidates are generated from this information), wherein individual region captions of the set of region captions are generated using a first region caption based on a labeled bounding box for an object within the image and a second region caption based on a global description of the image (Xie, [0020]-[0021] as shown in figure 2, multiple caption options for the image and the objects within the image (Region captions (Figure 2, 126 and figure 2, 130)) are generated based on the objects detected, and multiple caption candidates are generated from this information, where [0039] local descriptions/captions (first region captions) are generated for each selected bounding box and [0040] the collected visual information which includes the initial caption (global description of the image) are used to generate visual clues, which include captions (second region caption)). Regarding claim 17 Xie discloses; The system of claim 16, where the one or more processing units are further to generate an object list for the input image based on a first object list corresponding to a first output for an object recognition model (Xie, [0019] an object detector model generates a list of a plurality of objects) and a second object list corresponding to a second output for a large language model (Xie, [0018] a set of sematic information is generated for the image using a large language model). Regarding claim 18 Xie discloses; The system of claim 17, wherein the object recognition model processes the input image (Xie[0019] an object detector model generates a list of a plurality of objects from the input image) and the large language model processes an image caption generated by a vision language model based on the input image and raw caption data (Xie, [0018] a vision language model generates an initial caption (raw caption) and information about the input image, then a large language model processes this output). Regarding claim 20, Xie discloses; The system of claim 16, wherein the system is comprised in at least one of: a system for performing simulation operations (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing simulation operations to test or validate autonomous machine learning applications (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing digital twin operations(Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing light transport simulation (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for rendering graphical output (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing deep learning operations (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system implemented using an edge device (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for generating or presenting virtual reality (VR) content (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for generating or presenting augmented reality (AR) content (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for generating or presenting mixed reality (MR) content (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system incorporating one or more Virtual Machines (VMs) (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing operations for a conversational AI application (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing operations for a generative AI application(Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing operations using a language model (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing one or more operations using a large language model (LLM) (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing one or more operations using a vision language model (VLM) (Xie, [0044] a vision language model maybe executed using a block-based image processor, [0063] the system is a computing system for executing multiple vision language models); a system implemented at least partially in a data center (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing hardware testing using simulation (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for performing one or more generative content operations using a language model (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a system for synthetic data generation (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); a collaborative content creation platform for 3D assets (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed); or a system implemented at least partially using cloud computing resources (Given that the claim requirements state the processor be comprised in “at least one of” the listed systems but not all of the listed system types, the examiner has referred to paragraphs [0044] and [0063] as meeting the requirements listed). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 3-4, 12 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Xie (US 20230394855 A1) in view of Petitpont (US 12148233 B1). Regarding claim 3 Xie fails to teach; The processor of claim 1, wherein the one or more circuits are further to: generate a first object list from the input image; generate a second object list from the first caption; and combine the first object list and the second object list to form the object list. However, Petitpont teaches; wherein the one or more circuits are further to: generate a first object list from the input image (Petitpont, Column 10 lines 44-52, the recognition engine may output a set of descriptive labels and bounding boxes corresponding to objects in the images (first list)); PNG media_image5.png 118 302 media_image5.png Greyscale (Petitpont, column 10, emphasis added) generate a second object list from the first caption (Petitpont, column 10, lines 22-33, the system generates a feature list from the caption (second list) including objects included in the caption); PNG media_image6.png 142 304 media_image6.png Greyscale (Petitpont, column 10, emphasis added ) and combine the first object list and the second object list to form the object list (Petitpont, Column 8, lines 32-42, labels produced by the object detection engine may be fused with the captioning labels to generate a combined label set) PNG media_image7.png 144 302 media_image7.png Greyscale (Petitpont, Column 8) The combination of Xie and Petitpont would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for adding the object label list merging of Petitpont to the captioning system of Xie is that merging labels based on general captioning with the labels from image-based object detection allows for more detailed image captions to be generated. (Petitpont, Columns 8-10) Regarding claim 4 the combination of Xie and Petitpont teaches; The processor of claim 3, wherein the one or more circuits are further to: provide the object list to one or more trained machine learning systems to generate the one or more bounding boxes (Petitpont, Column 10 line 8- 65 regions (bounding boxes) in the image may be determined from the features (object labels), for example a caption/object list which has the words man, woman, person, may be used to generate labeled regions, the labels are then used to generate a bounding box, which has a precise label on it). The combination of Xie and Petitpont would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the addition of the bounding box generation of Petitpont to the system of Xie is that this allows for detection of selected objects in the image, which can then be localized and more precisely labeled to generate a more detailed caption for the image. (Petitpont, Columns 8-10) Regarding claim 12, the combination of Xie and Petitpont teaches; The computer-implemented method of claim 11, further comprising: identifying the one or more identified objects in the image using a trained machine learning model (Petitpont, Column 10 lines 44-52, the recognition engine may output a set of descriptive labels and bounding boxes corresponding to objects in the images); generating, a second object list based on the one or more identified objects (Petitpont, column 10, lines 22-33, the system generates a feature list from the caption (second list) including objects identified in the caption); and generating the object list based on the first object list and the second object list (Petitpont, Column 8, lines 32-42, labels produced by the object detection engine may be fused with the captioning labels to generate a combined label set). The combination of Xie and Petitpont would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for adding the object label list merging of Petitpont to the captioning system of Xie is that merging labels based on general captioning with the labels from image-based object detection allows for more detailed image captions to be generated. (Petitpont, Columns 8-10) Regarding claim 19, the combination of Xie and Petitpont teaches; The system of claim 16, wherein the one or more processing units are further to provide the caption to a human-in-the-loop review engine responsive to determining a similarity metric between labeled objects in the input image is below a threshold (Petitpont column 15 line 45- column 16 line 20, when the recognition engine determines that the object label accuracy is lower than expected based on the known ground truth label, the labels may be sent to a user for review). The combination of Xie and Petitpont would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for adding the human review step of Petitpont to the system of Xie is that it allows for feedback to be provided to the model to improve training. (Petitpont columns 15-16) Claims 5 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Xie (US 20230394855 A1) in view of Suk (US 20120076408 A1). Regarding claim 5 Xie fails to teach; The processor of claim 1, wherein the one or more circuits are further to: identify a plurality of bounding boxes, of the one or more bounding boxes, having a common label; determine at least a portion of the plurality of bounding boxes are within a threshold distance; and combine the portion of the plurality of bounding boxes within a single bounding box. However, in the same field of endeavor, Suk teaches; identify a plurality of bounding boxes, of the one or more bounding boxes, having a common label (Suk, [0063] an input image is acquired and multiple regions (bounding boxes) are detected for the object, [0064] the region extractor determines that regions which are adjacent have the same determined features/characteristics (common label)); determine at least a portion of the plurality of bounding boxes are within a threshold distance (Suk, [0064] the regions must be adjacent to one another (threshold distance requirement)); and combine the portion of the plurality of bounding boxes within a single bounding box (Suk, [0064] the region extractor determines that regions which are adjacent have the same determined features/characteristics (common label), these adjacent regions are then combined). The combination of Xie and Suk would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the addition of the bounding box/region merging of Suk is that this allows for fast and accurate detection of objects by excluding regions which likely are mislabeled or do not contain the object, and narrows the regions which have feature labels which indicate it belongs to the same object. (Suk, [0010]-[0013]) Regarding claim 13, the combination of Xie and Suk teaches; The computer-implemented method of claim 10, further comprising: determining a plurality of bounding boxes, of the set of bounding boxes, have a common label (Suk, [0063] an input image is acquired and multiple regions (bounding boxes) are detected for the object, [0064] the region extractor determines that regions which are adjacent have the same determined features/characteristics (common label)); determining at least a portion of the plurality of bounding boxes are within a threshold distance (Suk, [0064] the regions must be adjacent to one another (threshold distance requirement)); and combining the portion of the plurality of bounding boxes within a single bounding box with the common label (Suk, [0064] the region extractor determines that regions which are adjacent have the same determined features/characteristics (common label), these adjacent regions are then combined). The combination of Xie and Suk would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the addition of the bounding box/region merging of Suk is that this allows for fast and accurate detection of objects by excluding regions which likely are mislabeled or do not contain the object, and narrows the regions which have feature labels which indicate it belongs to the same object. (Suk, [0010]-[0013]) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For a listing of analogous art as determined by the examiner, please see the attached PTO-892 Notice of References Cited form. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORDAN M ELLIOTT whose telephone number is (703)756-5463. The examiner can normally be reached M-F 8AM-5PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.M.E./Examiner, Art Unit 2666 /Molly Wilburn/Primary Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Nov 07, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705715
Tone Mapping for Preserving Contrast of Fine Features in an Image
3y 6m to grant Granted Aug 11, 2026
Patent 12682611
IMAGE ACQUISITION MODEL TRAINING METHOD AND APPARATUS, IMAGE DETECTION METHOD AND APPARATUS, AND DEVICE
3y 1m to grant Granted Jul 14, 2026
Patent 12573117
METHOD AND DEVICE FOR DEEP LEARNING-BASED PATCHWISE RECONSTRUCTION FROM CLINICAL CT SCAN DATA
3y 1m to grant Granted Mar 10, 2026
Patent 12475998
SYSTEMS AND METHODS OF ADAPTIVELY GENERATING FACIAL DEVICE SELECTIONS BASED ON VISUALLY DETERMINED ANATOMICAL DIMENSION DATA
2y 2m to grant Granted Nov 18, 2025
Patent 12450918
AUTOMATIC LANE MARKING EXTRACTION AND CLASSIFICATION FROM LIDAR SCANS
3y 1m to grant Granted Oct 21, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
41%
Grant Probability
15%
With Interview (-25.5%)
2y 12m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month