DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 21-22, 25-33, 35 and 37-44 are pending in the application. Claims 21, 25, 33, 35, 38 and 40 have been amended, claims 23-24, 34 and 36 have been canceled, and claims 41-44 have been added.
The amendment filed 6/25/26 overcomes objection to claim 38.
Response to Arguments
Applicant’s arguments, filed 6/25/26, with respect to 35 USC 101 rejection applied to claim 21 regarding abstract idea, in view of claim amendment, are persuasive. The 35 U.S.C. 101 rejection applied to originally filed claims 21-40 has been withdrawn.
Applicant’s arguments, filed 6/25/26, with respect to art rejection to claim(s) 21, especially “a selectable control”, have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 21 is rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 4 and 10 of U.S. Patent No. US 12039431 B1. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following reasons.
Listed in the following table is a limitation-to-limitation comparison of the examined claim 21 and the conflicting claims 1, 4 and 10.
Application being examined 18/746,326 (hereafter ‘326 application)
Conflicting patent US 12039431 B1 (hereafter ‘431 patent)
21. A method comprising:
receiving a user prompt comprising an area of emphasis in an image and a textual prompt (See ‘431 claim 10);
generating input data using the user prompt, the input data generated in a format usable by a machine learning model (See ‘431 claim 1 “by applying the input data to a multimodal machine learning model”), wherein generating the input data using the user prompt comprises tokenizing the user prompt, wherein tokenizing is patch-based tokenization or region- of-interest tokenization (See ‘431 claim 1 “generate a contextual prompt that indicates an area of emphasis in the image”, claim 4 “generating …token using the contextual prompt”);
generating a response to the user prompt by applying the input data to the machine learning model, the machine learning model being pre-trained and configured to condition the response to the image,
wherein the response comprises a prompt suggestion, wherein the prompt suggestion is conditioned on the user prompt and is displayed via a graphical user interface as a selectable control.
1. A method of interacting with a pre-trained multimodal machine learning model, the method comprising:
providing a graphical user interface configured to enable a user to interact with an image to generate a contextual prompt that indicates an area of emphasis in the image;
receiving the contextual prompt;
generating input data using the image and the contextual prompt;
generating a textual response to the image by applying the input data to a multimodal machine learning model configured to condition the textual response to the image on the contextual prompt; and
providing the textual response to the user,
wherein the textual response comprises a prompt suggestion and providing the textual response comprises displaying a selectable control in the graphical user interface configured to enable the user to select the prompt suggestion.
4. The method of claim 1, wherein: generating the input data using the image and the contextual prompt comprises:
generating a textual prompt or token using the contextual prompt; and generating the input data using the image and the textual prompt or token.
10. The method of claim 1, wherein: the method further includes receiving a textual prompt from the user; the input data is further generated using the textual prompt; and the multimodal machine learning model is configured to further condition the textual response to the image on the textual prompt.
Therefore, claims 1, 4 and 10 of ‘431 patent teaches every limitation as recited in examined claim 21 of the ‘326 application.
Claim 33 is rejected on the ground of nonstatutory double patenting as being unpatentable over claims 13 and 16 of U.S. Patent No. US 12039431 B1. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following reasons.
Listed in the following table is a limitation-to-limitation comparison of the examined claim 33 and the conflicting claims 13 and 16.
Application being examined 18/746,326 (hereafter ‘326 application)
Conflicting patent US 12039431 B1 (hereafter ‘431 patent)
33. A system comprising:
at least one processor; and
at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
receiving a user prompt comprising an area of emphasis in an image;
generating input data using the image and the user prompt, the input data generated in a format usable by a machine learning model (See ‘431 claim 13 “by applying the input data to the multimodal machine learning model”), wherein generating the input data using the image and the user prompt comprises tokenizing at least one of the image or the user prompt, and the tokenizing is patch-based tokenization or region-of-interest tokenization (See ‘431 claim 13 “generate a contextual prompt that indicates an area of emphasis in the image”, claim 16 “generating …token using the contextual prompt”);
generating a response to the user prompt by applying the input data to the machine learning model, the machine learning model being pre-trained and configured to condition the response to the image,
wherein the response comprises a prompt suggestion, wherein the prompt suggestion is conditioned on the user prompt and is displayed via a graphical user interface as a selectable control.
13. A system for interacting with a pre-trained multimodal machine learning model, comprising:
at least one processor; and
at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
providing a graphical user interface configured to enable a user to interact with an image to generate a contextual prompt that indicates an area of emphasis in the image;
receiving the contextual prompt;
generating input data using the image and the contextual prompt;
generating a textual response to the image by applying the input data to the multimodal machine learning model configured to condition the textual response to the image on the contextual prompt; and
providing the textual response to the user,
wherein: the textual response comprises a prompt suggestion, and providing the textual response comprises displaying a selectable control in the graphical user interface, the selectable control being programmed to enable the user to select the prompt suggestion.
16. The system of claim 13, wherein: generating the input data using the image and the contextual prompt comprises:
generating a textual prompt or token using the contextual prompt, the textual prompt or token indicating coordinates of a location in the image; and generating the input data using the image and the textual prompt or token.
Therefore, claims 13 and 16 of ‘431 patent teaches every limitation as recited in examined claim 33 of the ‘326 application.
Claim 40 is rejected on the ground of nonstatutory double patenting as being unpatentable over claims 13 and 16 of U.S. Patent No. US 12039431 B1. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following reasons.
Listed in the following table is a limitation-to-limitation comparison of the examined claim 40 and the conflicting claims 13 and 16.
Application being examined 18/746,326 (hereafter ‘326 application)
Conflicting patent US 12039431 B1 (hereafter ‘431 patent)
40. A machine learning system comprising:
at least one processor; and
at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
receiving a user prompt comprising an area of emphasis in an image;
generating at least one token using the user prompt, wherein generating the at least one token comprises patch-based tokenization or region-of-interest tokenization (See ‘431 claim 13 “generate a contextual prompt that indicates an area of emphasis in the image” and claim 16);
generating input data using the image and the at least one token (See claim 16), the input data generated in a format usable by a machine learning model (See ‘431 claim 13 “by applying the input data to the multimodal machine learning model”); and
generating a response to the user prompt by applying the input data to the machine learning model, the machine learning model being pre-trained and configured to condition the response to the image,
wherein the response comprises a prompt suggestion, wherein the prompt suggestion is conditioned on the user prompt and is displayed via a graphical user interface as a selectable control.
13. A system for interacting with a pre-trained multimodal machine learning model, comprising:
at least one processor; and
at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:
providing a graphical user interface configured to enable a user to interact with an image to generate a contextual prompt that indicates an area of emphasis in the image;
receiving the contextual prompt;
generating input data using the image and the contextual prompt;
generating a textual response to the image by applying the input data to the multimodal machine learning model configured to condition the textual response to the image on the contextual prompt; and
providing the textual response to the user,
wherein: the textual response comprises a prompt suggestion, and providing the textual response comprises displaying a selectable control in the graphical user interface, the selectable control being programmed to enable the user to select the prompt suggestion.
16. The system of claim 13, wherein: generating the input data using the image and the contextual prompt comprises: generating a textual prompt or token using the contextual prompt, the textual prompt or token indicating coordinates of a location in the image; and generating the input data using the image and the textual prompt or token.
Therefore, claims 13 and 16 of ‘431 patent teach every limitation as recited in examined claim 40 of the ‘326 application.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 21-22, 25-26, 28-33, 35, 37-38, 40-42 and 44 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (US Publication 2020/0380258A1, hereafter Yu), in view of Price et al. (US Publication 2019/0236394A1, hereafter Price) and Bean (US 20240320867 A1).
As per claim 21, Yu teaches a method (Abstract; FIG. 9) comprising:
receiving a user prompt comprising an area of emphasis in an image and a textual prompt (Yu discloses a data annotation system 102 (FIG. 1). The computing system 102 includes a crowd worker interface process 116 that is programmed to provide an interface between the machine-learning model 110 and the crowd workers 128 (via the work stations 126) (see FIG. 1-8; para. [0038] ln 1-4). A user prompt comprises an area of emphasis in an image (FIG. 2 the draw box command 212 may enable the crowd worker 128 to encompass an area on the figure to highlight a particular feature (para. [0045])) and a textual prompt (In FIG. 6, the crowd worker 128 placed a text label (“car”) on the vehicle (para. [0052]));
generating input data using the user prompt (para. [0038] “The scripted dialogs may include a particular request to the crowd worker 128 to provide input in a particular manner. For example, for annotating video images, the crowd worker interface 116 may request that the crowd worker 128 circle or highlight an area of the displayed image. In other examples, the crowd worker interface 116 may request a text input. In other examples, the crowd worker interface 116 may request the crowd worker 128 to point and click on areas of interest on the displayed image”; In FIG. 6, the crowd worker 128 placed a text label (“car”) on the vehicle, drew a boundary 612 around the car and highlighted another pedestrian with a second box 614 (para. [0054]). The displayed image and the added contextual prompt (label, 612, 614) constitutes the input data), the input data generated in a format usable by a machine learning model (Yu para. [0064] “For example, the bounding boxes may be represented as a set of image coordinates. The crowd worker interface 116 may convert the set of image coordinates to a box on the image frame at the appropriate location. Likewise, the crowd worker 128 may draw a box on the display screen. The crowd worker interface 116 may convert the box as displayed on the screen to coordinates that are understood by the machine-learning algorithm 110”);
generating a response to the user prompt by applying the input data to the machine learning model, the machine learning model being pre-trained and configured to condition the response to the image (FIG. 1 #110; para. [0036] “The machine-learning algorithm 110 may be configured to identify a feature in the raw source data 115 as a predetermined feature (e.g., pedestrian). The raw source data 115 may be derived from a variety of sources. For example, the raw source data 115 may be actual input data collected by a machine-learning system that is using the annotated dataset 114”; para. [0041] “During execution of the crowd-source task, the crowd worker interface 116 may execute the machine-learning algorithm 110 for a particular set of input data from the raw source dataset 115”), wherein the response comprises a prompt suggestion, wherein the prompt suggestion is conditioned on the user prompt and is displayed via a graphical user interface (para. [0042] “The crowd worker interface 116 may be configured to convert information from the machine-learning algorithm 110 to a natural language representation that non-experts can understand”; FIG. 3 #312; para. [0049] “FIG. 3 depicts an example of a second snapshot 300 of the display screen 201 that may represent another portion of the task execution. A second image frame 308 may be displayed on the display screen 201. In addition, features or data (e.g., annotations) generated by the machine-learning algorithm 110 may be overlaid on the image. For example, a bounding box 310 may be drawn to reflect a specific type of object recognized by the machine-learning algorithm 110. In addition, descriptive text 312 may be displayed proximate the bounding box 310. In the example, the bounding box 310 represents a pedestrian identified by the machine-learning algorithm 110 and the descriptive text 312 is a confidence level associated with the bounding box 310. In this example, the descriptive text 312 indicates a confidence level of 100%”; At least, the descriptive text 312 indicating a confidence level of 100% is considered a prompt suggestion. Further, Yu generates text suggestions, such as shown in FIG. 4 #406 “My confidence is low. Please pause …”).
Yu, although teaches converting patch-based or region-of-interest based input data into a machine-usable format, the limitation regarding generating the input data using the user prompt comprises tokenizing the user prompt, wherein tokenizing is patch-based tokenization or region-of-interest tokenization, is not apparently available.
Price in the same field of endeavor discloses systems and methods for selecting target objects within digital images utilizing a multi-modal object selection neural network trained to accommodate multiple input modalities (Abstract). The input modalities include a positive regional input modality (e.g., a positive or foreground click), a negative regional input modality (e.g., a negative or background click), boundary input modality (e.g., a first edge click and a second edge click), a language user input (e.g., words or phrases) etc. (See FIG. 1 and para. [0055]-[0057]). As shown in FIG. 1, the multi-modal object selection neural network 108 then analyzes the digital image 100, the regional user inputs 102a-102b corresponding to the regional input modality 102, the boundary user inputs 104a-104b corresponding to the boundary input modality 104, and the language user input 106a corresponding to the verbal input modality 106, and generates an object segmentation output 110.
Price specifically teaches tokenizing user prompt by generating distance maps before inputting the input data into the machine learning model. As shown in FIG. 2A, distance maps 210, 212, 215 (and/or additional map(s) 217) based on the digital image 200 and user input (i.e., the user inputs 204, 205, and 206) are generated. Further Price discloses the multi-modal selection system can generate the additional map(s) 217 based on language input (e.g., natural language expression). See para. [0061]-[0063], para. [0077].
Price further teaches the tokenizing is patch-based tokenization or region-of-interest tokenization (Price: As shown in FIG 2A, the multi-modal selection system tokenizes the user input by generating distance maps 210, 212, 215 (and/or additional map(s) 217) based on the digital image 200 and user input (i.e., the user inputs 204, 205, and 206). In particular, the multi-modal selection system generates the positive distance map 210 based on the user input corresponding to the positive regional input modality. Moreover, the multi-modal selection system generates the negative distance map 212 based on the user input 206 corresponding to the negative regional input modality. The multi-modal selection system generates the boundary distance map 215 based on the user input 205 corresponding to the boundary input modality (para. [0061]).Therefore the tokenizing is patch-based tokenization or region-of-interest tokenization. The system then generates an image/user interaction pair 226, which is a combination of user interaction data reflected in the distance maps 210, 212, and 215 and image data reflected in color channels 218-222, as input to a trained neural network to generate an object segmentation (FIG. 2A-2B; para.[0073])).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu to incorporate the teaching of Price to consider tokenizing the user prompt and the tokenizing is patch-based tokenization or region-of-interest tokenization. The motivation would be transforming various user inputs into tokens that can be analyzed by a machine learning model as suggested by Price (para. [0033]).
Yu in view of Price, however, does not further teach the prompt suggestion is displayed via a GUI as a selectable control.
Bean is evidenced that prompt suggestion can be displayed via a GUI as a selectable control (FIG. 3 top window shows prompt suggestion is displayed as selectable control which includes category options 320 that includes one or more radio buttons or checkboxes corresponding to a visual content class 330, a visual style class 340, a visual perspective class 350, a lighting class 360; para. [0051]-[0053]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu and Price to incorporate the teaching of Bean to display prompt suggestion via a GUI as a selectable control. The motivation would be to improve the overall user experience by reducing the time and effort associated with users identifying improvements or modifications to subsequent text prompts as recognized by Bean (para. [0027]).
As per claim 22, dependent upon claim 21, Yu in view of Price and Bean teaches the user prompt indicates at least one set of coordinates within the image (Yu para. [0064]: “Likewise, the crowd worker 128 may draw a box on the display screen. The crowd worker interface 116 may convert the box as displayed on the screen to coordinates that are understood by the machine-learning algorithm 110.”).
As per claim 25, dependent upon claim 21, Yu in view of Price and Bean teaches tokenizing the user prompt comprises:
tokenizing both the image and the textual prompt into separate sequences of tokens (Price [0078] “With regard to language input, the multi-modal selection system can generate the additional map(s) 217 by generating an embedding of a word or phrase. For example, the multi-modal selection system can utilize a recurrent neural network (RNN) to generate an embedding (e.g., vector representation) from the language input. The multi-modal selection system can concatenate the encoded vectors (and optionally, spatial coordinates) with the image features (e.g., the other maps 210, 212, 215, 218, 220, 222) to generate the image/user interaction pair 226”).
As per claim 26, dependent upon claim 25, Yu in view of Price and Bean teaches tokenizing both the image and the textual prompt into separate sequences of tokens comprises:
concatenating the tokenized textual prompt and tokenized image to form a singular tokenized input (Price [0078] “With regard to language input, the multi-modal selection system can generate the additional map(s) 217 by generating an embedding of a word or phrase. For example, the multi-modal selection system can utilize a recurrent neural network (RNN) to generate an embedding (e.g., vector representation) from the language input. The multi-modal selection system can concatenate the encoded vectors (and optionally, spatial coordinates) with the image features (e.g., the other maps 210, 212, 215, 218, 220, 222) to generate the image/user interaction pair 226”).
As per claim 28, dependent upon claim 21, Yu teaches drawing boundaries around objects in an image and using the drawn boundaries and other input, such as text “car”, as the input to the machine learning model (Yu FIG. 6 showing user-drawn boundary #612 and box #614, and input text “car” (para. [0054])). Note the user-drawn boundary #612 is a segmentation of the car in the image. Therefore the input data includes the image and a segmentation mask (para. [0054]). Yu, however, does not teach generating a segmentation mask by providing the user prompt to a segmentation model.
Price teaches a multi-modal selection system that utilizes an object selection neural network that analyzes multiple input modalities to generate object segmentation outputs corresponding to target objects portrayed in digital visual media (para. [0028]). As shown in FIG. 1, multiple input modalities 102, 104 and 106 are input into a multi-modal object selection neural network 108 and an object segmentation mask is generated (FIG. 1 #110; para. [0058]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu, Price and Bean as taught in claim 21, to further include the teaching of Price to generate a segmentation mask by providing user prompt to a segmentation model. Doing so would allow the segmentation model to suggest particular input modalities to a user for more active learning and accurate object segmentation (Price para. [0028]).
As per claim 29, dependent upon claim 21, Yu in view of Price and Bean teaches generating the input data using the user prompt comprises:
generating an updated image based on the user prompt (see below); and
generating the input data using the updated image (Yu: FIG. 6 a boundary 612 around the car and a box 614 are drawn in the displayed image, i.e., the image is updated. Therefore the input image data is changed in view of the user annotation).
As per claim 30, dependent upon claim 21, Yu in view of Price Bean teaches generating the input data using the user prompt comprises:
generating a prepended textual prompt which directs the machine learning model to provide a particular type of information regarding the area of emphasis in the image (Yu para. [0043]; FIG. 4 #406; FIG. 6 #606).
As per claim 31, dependent upon claim 21, Yu in view of Price and Bean further teaches:
providing a graphical user interface configured to enable a user to interact with the image to generate the user prompt (Yu FIG. 1 #116; FIG. 2-8).
As per claim 32, dependent upon claim 21, Yu in view of Price and Bean further teaches:
the user prompt indicates an object depicted in the image (Yu FIG. 3 #310 “pedestrian”; FIG. 6 “car”); and
the prompt suggestion is a textual response concerning the depicted object (Yu FIG. 3 #312 the prompt suggestion says “confidence 100%”, indicating the accuracy of recognition of the pedestrian is high; para. [0049]).
As per claim 33, an independent system claim, Yu teaches a system (Yu Abstract; FIG. 1) comprising:
at least one processor (Yu FIG. 1 #106 “CPU); and
at least one non-transitory computer readable medium (Yu FIG. 1 #108) containing instructions that, when executed by the at least one processor, cause the system to perform operations (Yu para. [0024] “During operation, the CPU 106 may execute stored program instructions that are retrieved from the memory unit 108. The stored program instructions may include software that controls operation of the CPU 106 to perform the operation described herein”) comprising:
receiving a user prompt comprising an area of emphasis in an image (Yu discloses a data annotation system 102 (FIG. 1). The computing system 102 includes a crowd worker interface process 116 that is programmed to provide an interface between the machine-learning model 110 and the crowd workers 128 (via the work stations 126) (Yu see FIG. 1-8; para. [0038] ln 1-4). A user prompt comprises an area of emphasis in an image (FIG. 2 the draw box command 212 may enable the crowd worker 128 to encompass an area on the figure to highlight a particular feature (para. [0045])) and a textual prompt (In FIG. 6, the crowd worker 128 placed a text label (“car”) on the vehicle (para. [0052]));
generating input data using the image and the user prompt (Yu para. [0038] “The scripted dialogs may include a particular request to the crowd worker 128 to provide input in a particular manner. For example, for annotating video images, the crowd worker interface 116 may request that the crowd worker 128 circle or highlight an area of the displayed image. In other examples, the crowd worker interface 116 may request a text input. In other examples, the crowd worker interface 116 may request the crowd worker 128 to point and click on areas of interest on the displayed image”; In FIG. 6, the crowd worker 128 placed a text label (“car”) on the vehicle, drew a boundary 612 around the car and highlighted another pedestrian with a second box 614 (para. [0054]). The displayed image and the added contextual prompt (label, 612, 614) constitutes the input data), the input data generated in a format usable by a machine learning model (Yu para. [0064] “For example, the bounding boxes may be represented as a set of image coordinates. The crowd worker interface 116 may convert the set of image coordinates to a box on the image frame at the appropriate location. Likewise, the crowd worker 128 may draw a box on the display screen. The crowd worker interface 116 may convert the box as displayed on the screen to coordinates that are understood by the machine-learning algorithm 110”),
generating a response to the user prompt by applying the input data to the machine learning model, the machine learning model being pre-trained and configured to condition the response to the image (Yu FIG. 1 #110; para. [0036] “The machine-learning algorithm 110 may be configured to identify a feature in the raw source data 115 as a predetermined feature (e.g., pedestrian). The raw source data 115 may be derived from a variety of sources. For example, the raw source data 115 may be actual input data collected by a machine-learning system that is using the annotated dataset 114”; para. [0041] “During execution of the crowd-source task, the crowd worker interface 116 may execute the machine-learning algorithm 110 for a particular set of input data from the raw source dataset 115”), wherein the response comprises a prompt suggestion, wherein the prompt suggestion is conditioned on the user prompt and is displayed via a graphical user interface (Yu para. [0042] “The crowd worker interface 116 may be configured to convert information from the machine-learning algorithm 110 to a natural language representation that non-experts can understand”; FIG. 3 #312; para. [0049] “FIG. 3 depicts an example of a second snapshot 300 of the display screen 201 that may represent another portion of the task execution. A second image frame 308 may be displayed on the display screen 201. In addition, features or data (e.g., annotations) generated by the machine-learning algorithm 110 may be overlaid on the image. For example, a bounding box 310 may be drawn to reflect a specific type of object recognized by the machine-learning algorithm 110. In addition, descriptive text 312 may be displayed proximate the bounding box 310. In the example, the bounding box 310 represents a pedestrian identified by the machine-learning algorithm 110 and the descriptive text 312 is a confidence level associated with the bounding box 310. In this example, the descriptive text 312 indicates a confidence level of 100%”; At least, the descriptive text 312 indicating a confidence level of 100% is considered a prompt suggestion. Further, Yu generates text suggestions, such as shown in FIG. 4 #406 “My confidence is low. Please pause …”).
Yu, although teaches converting patch-based or region-of-interest based input data into a machine-usable format, the limitation regarding generating the input data using the image and the user prompt comprises tokenizing at least one of the image or the user prompt and the tokenizing is patch-based tokenization or region-of-interest tokenization, is not apparently available.
Price in the same field of endeavor discloses systems and methods for selecting target objects within digital images utilizing a multi-modal object selection neural network trained to accommodate multiple input modalities (Abstract). The input modalities include a positive regional input modality (e.g., a positive or foreground click), a negative regional input modality (e.g., a negative or background click), boundary input modality (e.g., a first edge click and a second edge click), a language user input (e.g., words or phrases) etc. (See FIG. 1 and para. [0055]-[0057]). As shown in FIG. 1, the multi-modal object selection neural network 108 then analyzes the digital image 100, the regional user inputs 102a-102b corresponding to the regional input modality 102, the boundary user inputs 104a-104b corresponding to the boundary input modality 104, and the language user input 106a corresponding to the verbal input modality 106, and generates an object segmentation output 110.
Price specifically teaches tokening at least one of the image or the user prompt by generating distance maps before inputting the input data into the machine learning model. As shown in FIG. 2A, distance maps 210, 212, 215 (and/or additional map(s) 217) based on the digital image 200 and user input (i.e., the user inputs 204, 205, and 206) are generated. Further Price discloses the multi-modal selection system can generate the additional map(s) 217 based on language input (e.g., natural language expression). See para. [0061]-[0063], para. [0077].
Price further teaches the tokenizing is patch-based tokenization or region-of-interest tokenization (Price: As shown in FIG 2A, the multi-modal selection system tokenizes the user input by generating distance maps 210, 212, 215 (and/or additional map(s) 217) based on the digital image 200 and user input (i.e., the user inputs 204, 205, and 206). In particular, the multi-modal selection system generates the positive distance map 210 based on the user input corresponding to the positive regional input modality. Moreover, the multi-modal selection system generates the negative distance map 212 based on the user input 206 corresponding to the negative regional input modality. The multi-modal selection system generates the boundary distance map 215 based on the user input 205 corresponding to the boundary input modality (para. [0061]).Therefore the tokenizing is patch-based tokenization or region-of-interest tokenization. The system then generates an image/user interaction pair 226, which is a combination of user interaction data reflected in the distance maps 210, 212, and 215 and image data reflected in color channels 218-222, as input to a trained neural network to generate an object segmentation (FIG. 2A-2B; para.[0073])).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu to incorporate the teaching of Price to consider tokening the image or user prompt and the tokenizing is patch-based tokenization or region-of-interest tokenization. The motivation would be transforming various user inputs into tokens that can be analyzed by a machine learning model as suggested by Price (para. [0033]).
Yu in view of Price, however, does not further teach the prompt suggestion is displayed via a GUI as a selectable control.
Bean is evidenced that prompt suggestion can be displayed via a GUI as a selectable control (FIG. 3 top window shows prompt suggestion is displayed as selectable control which includes category options 320 that includes one or more radio buttons or checkboxes corresponding to a visual content class 330, a visual style class 340, a visual perspective class 350, a lighting class 360; para. [0051]-[0053]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu and Price to incorporate the teaching of Bean to display prompt suggestion via a GUI as a selectable control. The motivation would be improving the overall user experience by reducing the time and effort associated with users identifying improvements or modifications to subsequent text prompts as recognized by Bean (para. [0027]).
As per claim 35, dependent upon claim 33, Yu in view of Price and Bean teaches tokenizing at least one of the image or the user prompt indicates at least one set of coordinates within the image (Yu para. [0033] “The annotations may include descriptions that are associated with identified coordinates of the image frame”; para. [0064]: “Likewise, the crowd worker 128 may draw a box on the display screen. The crowd worker interface 116 may convert the box as displayed on the screen to coordinates that are understood by the machine-learning algorithm 110.”).
As per claim 37, dependent upon claim 33, Yu in view of Price and Bean teaches generating the input data using the image and the user prompt comprises:
tokenizing both the image and the user prompt into separate sequences of tokens (Price [0078] “With regard to language input, the multi-modal selection system can generate the additional map(s) 217 by generating an embedding of a word or phrase. For example, the multi-modal selection system can utilize a recurrent neural network (RNN) to generate an embedding (e.g., vector representation) from the language input. The multi-modal selection system can concatenate the encoded vectors (and optionally, spatial coordinates) with the image features (e.g., the other maps 210, 212, 215, 218, 220, 222) to generate the image/user interaction pair 226”).
As per claim 38, dependent upon claim 37, Yu in view of Price and Bean teaches generating the input data using the image and the user prompt comprises:
concatenating the tokenized user prompt and tokenized image to form a singular tokenized input (Price [0078] “With regard to language input, the multi-modal selection system can generate the additional map(s) 217 by generating an embedding of a word or phrase. For example, the multi-modal selection system can utilize a recurrent neural network (RNN) to generate an embedding (e.g., vector representation) from the language input. The multi-modal selection system can concatenate the encoded vectors (and optionally, spatial coordinates) with the image features (e.g., the other maps 210, 212, 215, 218, 220, 222) to generate the image/user interaction pair 226”).
As per claim 40, an independent claim, Yu teaches a machine learning system (Abstract; FIG. 1) comprising:
at least one processor (FIG. 1 #106 “CPU); and
at least one non-transitory computer readable medium (FIG. 1 #108) containing instructions that, when executed by the at least one processor, cause the system to perform operations (para. [0024] “During operation, the CPU 106 may execute stored program instructions that are retrieved from the memory unit 108. The stored program instructions may include software that controls operation of the CPU 106 to perform the operation described herein”) comprising:
receiving a user prompt comprising an area of emphasis in an image (Yu discloses a data annotation system 102 (FIG. 1). The computing system 102 includes a crowd worker interface process 116 that is programmed to provide an interface between the machine-learning model 110 and the crowd workers 128 (via the work stations 126) (see FIG. 1-8; para. [0038] ln 1-4). A user prompt comprises an area of emphasis in an image (FIG. 2 the draw box command 212 may enable the crowd worker 128 to encompass an area on the figure to highlight a particular feature (para. [0045])) and a textual prompt (In FIG. 6, the crowd worker 128 placed a text label (“car”) on the vehicle (para. [0052]));
generating input data using the image and the user prompt (Yu para. [0038] “The scripted dialogs may include a particular request to the crowd worker 128 to provide input in a particular manner. For example, for annotating video images, the crowd worker interface 116 may request that the crowd worker 128 circle or highlight an area of the displayed image. In other examples, the crowd worker interface 116 may request a text input. In other examples, the crowd worker interface 116 may request the crowd worker 128 to point and click on areas of interest on the displayed image”; In FIG. 6, the crowd worker 128 placed a text label (“car”) on the vehicle, drew a boundary 612 around the car and highlighted another pedestrian with a second box 614 (para. [0054]). The displayed image and the added contextual prompt (label, 612, 614) constitutes the input data), the input data generated in a format usable by a machine learning model (Yu para. [0064] “For example, the bounding boxes may be represented as a set of image coordinates. The crowd worker interface 116 may convert the set of image coordinates to a box on the image frame at the appropriate location. Likewise, the crowd worker 128 may draw a box on the display screen. The crowd worker interface 116 may convert the box as displayed on the screen to coordinates that are understood by the machine-learning algorithm 110”); and
generating a response to the user prompt by applying the input data to the machine learning model, the machine learning model being pre-trained and configured to condition the response to the image (FIG. 1 #110; para. [0036] “The machine-learning algorithm 110 may be configured to identify a feature in the raw source data 115 as a predetermined feature (e.g., pedestrian). The raw source data 115 may be derived from a variety of sources. For example, the raw source data 115 may be actual input data collected by a machine-learning system that is using the annotated dataset 114”; para. [0041] “During execution of the crowd-source task, the crowd worker interface 116 may execute the machine-learning algorithm 110 for a particular set of input data from the raw source dataset 115”), wherein the response comprises a prompt suggestion, wherein the prompt suggestion is conditioned on the user prompt and is displayed via a graphical user interface (Yu para. [0042] “The crowd worker interface 116 may be configured to convert information from the machine-learning algorithm 110 to a natural language representation that non-experts can understand”; FIG. 3 #312; para. [0049] “FIG. 3 depicts an example of a second snapshot 300 of the display screen 201 that may represent another portion of the task execution. A second image frame 308 may be displayed on the display screen 201. In addition, features or data (e.g., annotations) generated by the machine-learning algorithm 110 may be overlaid on the image. For example, a bounding box 310 may be drawn to reflect a specific type of object recognized by the machine-learning algorithm 110. In addition, descriptive text 312 may be displayed proximate the bounding box 310. In the example, the bounding box 310 represents a pedestrian identified by the machine-learning algorithm 110 and the descriptive text 312 is a confidence level associated with the bounding box 310. In this example, the descriptive text 312 indicates a confidence level of 100%”; At least, the descriptive text 312 indicating a confidence level of 100% is considered a prompt suggestion. Further, Yu generates text suggestions, such as shown in FIG. 4 #406 “My confidence is low. Please pause …”).
Yu does not teach generating at least one token using the user prompt, wherein generating the at least one token comprises patch-based tokenization or region-of-interest tokenization.
Price in the same field of endeavor discloses systems and methods for selecting target objects within digital images utilizing a multi-modal object selection neural network trained to accommodate multiple input modalities (Abstract). The input modalities include a positive regional input modality (e.g., a positive or foreground click), a negative regional input modality (e.g., a negative or background click), boundary input modality (e.g., a first edge click and a second edge click), a language user input (e.g., words or phrases) etc. (See FIG. 1 and para. [0055]-[0057]). As shown in FIG. 1, the multi-modal object selection neural network 108 then analyzes the digital image 100, the regional user inputs 102a-102b corresponding to the regional input modality 102, the boundary user inputs 104a-104b corresponding to the boundary input modality 104, and the language user input 106a corresponding to the verbal input modality 106, and generates an object segmentation output 110.
Price specifically teaches tokening user prompt by generating distance maps before inputting the input data into the machine learning model. As shown in FIG. 2A, distance maps 210, 212, 215 (and/or additional map(s) 217) based on the digital image 200 and user input (i.e., the user inputs 204, 205, and 206) are generated. Further Price discloses the multi-modal selection system can generate the additional map(s) 217 based on language input (e.g., natural language expression). See para. [0061]-[0063], para. [0077].
Price further teaches the tokenizing is patch-based tokenization or region-of-interest tokenization (Price: As shown in FIG 2A, the multi-modal selection system tokenizes the user input by generating distance maps 210, 212, 215 (and/or additional map(s) 217) based on the digital image 200 and user input (i.e., the user inputs 204, 205, and 206). In particular, the multi-modal selection system generates the positive distance map 210 based on the user input corresponding to the positive regional input modality. Moreover, the multi-modal selection system generates the negative distance map 212 based on the user input 206 corresponding to the negative regional input modality. The multi-modal selection system generates the boundary distance map 215 based on the user input 205 corresponding to the boundary input modality (para. [0061]).Therefore the tokenizing is patch-based tokenization or region-of-interest tokenization. The system then generates an image/user interaction pair 226, which is a combination of user interaction data reflected in the distance maps 210, 212, and 215 and image data reflected in color channels 218-222, as input to a trained neural network to generate an object segmentation (FIG. 2A-2B; para.[0073])).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu to incorporate the teaching of Price to consider tokening the user prompt and the tokenizing is patch-based tokenization or region-of-interest tokenization. The motivation would be transforming various user inputs into tokens that can be analyzed by a machine learning model as suggested by Price (para. [0033]).
Yu in view of Price, however, does not further teach the prompt suggestion is displayed via a GUI as a selectable control.
Bean is evidenced that prompt suggestion can be displayed via a GUI as a selectable control (FIG. 3 top window shows prompt suggestion is displayed as selectable control which includes category options 320 that includes one or more radio buttons or checkboxes corresponding to a visual content class 330, a visual style class 340, a visual perspective class 350, a lighting class 360; para. [0051]-[0053]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu and Price to incorporate the teaching of Bean to display prompt suggestion via a GUI as a selectable control. The motivation would be improving the overall user experience by reducing the time and effort associated with users identifying improvements or modifications to subsequent text prompts as recognized by Bean (para. [0027]).
As per claim 41, dependent upon claim 21, Yu in view of Price and Bean teaches:
wherein the graphical user interface includes an annotation tool; and
the area of emphasis in the image comprises an annotation generated using the annotation tool (Yu FIG. 2 #212 “draw box”, #214 “draw line”; FIG. 5 #512; FIG. 6 #612).
As per claim 42, dependent upon claim 41, Yu in view of Price and Bean teaches:
wherein the annotation tool includes a loupe, a marker, or a segmentation tool (Yu FIG. 6 #612 being a drawn boundary, regarded as a segmentation; Price FIG. 1 102a, 102b, 104a and 104b being generated by a marker (see para. [0055])).
As per claim 44, dependent upon claim 33, Yu in view of Price and Bean teaches:
wherein the graphical user interface includes an annotation tool; and
the area of emphasis in the image comprises an annotation generated using the annotation tool (Yu FIG. 2 #212 “draw box”, #214 “draw line”; FIG. 5 #512; FIG. 6 #612).
Claims 27 and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (US Publication 2020/0380258A1, hereafter Yu), in view of Price et al. (US Publication 2019/0236394A1, hereafter Price) and Bean (US 20240320867 A1), and further in view of Liu et al. (Liu C, Lin Z, Shen X, Yang J, Lu X, Yuille A. Recurrent Multimodal Interaction for Referring Image Segmentation. In2017 IEEE International Conference on Computer Vision (ICCV) 2017 Oct 22 (pp. 1280-1289). IEEE. Hereafter Liu).
As per claim 27, dependent upon claim 26, Yu in view of Price and Bean teaches embedding the tokenized textual input into a vector space using a convolution neural network (Price [0078] “With regard to language input, the multi-modal selection system can generate the additional map(s) 217 by generating an embedding of a word or phrase. For example, the multi-modal selection system can utilize a recurrent neural network (RNN) to generate an embedding (e.g., vector representation) from the language input”). Yu in view of Price and Bean, however, does not teach embedding the singular tokenized input into a vector space using at least one of: a convolution neural network , a linear projection, or a graph neural network.
Liu in an analogous field discloses tokenizing image input and language input separately, concatenating the two tokenized inputs into a singular input and embedding the singular input into a vector space using a neural network (Fig. 3; page 1282 left col. (see below picture)).
PNG
media_image1.png
802
536
media_image1.png
Greyscale
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to modify the teaching of Yu, Price and Bean to incorporate the teaching of Liu to embed the singular tokenized input into a vector space using a convolution neural network. Doing so would allow joint modeling of two modalities for the image segmentation as suggested by Liu (Abstract).
Regarding claim 39, dependent upon claim 38, claim 39 recites a system with elements corresponding to the step recited in claim 27. Therefore, the recited elements of claim 39 are mapped to Yu in view of Price, Bean and Liu in the same manner as the corresponding step in its corresponding method claim, claim 27. Additionally, the rationale and motivation to combine Yu, Price, Bean and Liu presented in rejection of claim 27 apply to this claim.
Claim(s) 43 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (US Publication 2020/0380258A1, hereafter Yu), in view of Price et al. (US Publication 2019/0236394A1, hereafter Price) and Bean (US 20240320867 A1), as applied above to claim 41, and further in view of Bhatt (US Publication 2015/0135125 A1).
As per claim 43, dependent upon claim 41, Yu in view of Price and Bean does not teach the recited limitations.
Bhatt discloses bubble loupes that can be displayed over a portion of a display screen to magnify a region of the display screen. Users may select a region of an image to be displayed as a magnified view on the display screen as a bubble loupe, which automatically resizes and repositions as size and position of the image changes (Abstract; FIG. 1A-1B; para. [0044]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of this application, to combine the teachings of Bhatt with Yu and Price and Bean to consider resizing a selected area. In this way, bubble loupes may be used to view selected regions of the image at higher or lower levels of magnification simultaneously with the remainder of the image displayed on the display screen as mentioned by Bhatt (Abstract).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XUEMEI G CHEN whose telephone number is (571)270-3480. The examiner can normally be reached Monday-Friday 9am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M Villecco can be reached on (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XUEMEI G CHEN/Primary Examiner, Art Unit 2661