DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 19 is rejected under 35 U.S.C. § 103 as being unpatentable over Turgutlu et al., US 2023/0274478 A1 (“Turgutlu”), in view of Wang et al., “Deep Convolutional Priors for Indoor Scene Synthesis,” ACM Transactions on Graphics, Vol. 37, No. 4, Article 70 (2018) (“Wang”).
Regarding claim 19:
Turgutlu teaches: a computer-readable storage medium for storing computer-readable instructions, a processing system executing the computer-readable instructions to perform operations (Turgutlu explains that the disclosed functionality may be implemented by a processor executing instructions stored in memory and identifies suitable non-transitory computer readable storage media, including random-access memory, read-only memory, EEPROM, optical storage, and magnetic storage. See Turgutlu ¶¶ 101–114 and Fig. 8).
receiving an input image (Turgutlu describes receiving an image depicting one or more existing objects, detecting class and position information for those objects, and generating a token sequence representing the detected objects and their respective bounding boxes. See Turgutlu ¶¶ 42, 50–51, 93, 96, and 117–119; Figs. 1, 4, 7, and 9.)
producing result information based on the input image that specifies an object to be placed in the input image and a location at which to place the object in the input image (Turgutlu represents an object using a class token and placement tokens corresponding to an x coordinate, a y coordinate, a width, and a height. The disclosed sequence encoder predicts a class-token value and associated placement-token values for an additional object to be inserted into the image. Turgutlu further explains that the system may predict the most likely class, location, and scale of an additional object. See Turgutlu ¶¶ 44, 52–53, 58–63, and 67–77; Figs. 5–6);
the producing including (Turgutlu teaches class conditioned object placement. Turgutlu explains that an object class may be provided by a user, such as through a request to insert a chair. When the class is supplied, the class token is included in the input sequence and the sequence encoder predicts the corresponding x-coordinate, y-coordinate, width, and height values for placement of the object. See Turgutlu ¶¶ 42–44, 48, 52–54, 76–88, and 119–120; Figs. 1, 3, 5–6, and 9);
More particularly, Turgutlu describes an input sequence having a known class token and masked placement tokens. The sequence encoder iteratively predicts the x-coordinate, y-coordinate, width, and height values. The predicted placement-token values identify the location and scale at which the specified object is to be inserted. See Turgutlu ¶¶ 76–88 and Fig. 6.
and a third mode in which the machine-trained model identifies both the object and the location (Turgutlu teaches that the object class need not be supplied by the user. When an insertion request does not identify a particular object class, the system predicts a likely object class based on the input image and thereafter predicts corresponding position and scale information. Turgutlu describes inserting mask tokens corresponding to the class, x-coordinate, y-coordinate, width, and height fields, predicting the class-token value, and then iteratively predicting the placement-token values. See Turgutlu ¶¶ 44, 52, 63–66, and 74–77; Figs. 1 and 5–6);
the producing being performed independently of an image of the object (Turgutlu teaches generating the object-class and placement information from the layout of the input image, the class and bounding-box tokens representing existing scene objects, and an optional textual query or supplied semantic class. The sequence encoder predicts the semantic class and placement information before a particular foreground-object image or object mask is selected for compositing. Turgutlu describes first identifying the predicted class-token value and corresponding semantic category and thereafter selecting an object mask associated with that semantic category. Thus, the information specifying the object category and placement is produced without requiring an image of the candidate object as an input to the class-and-placement prediction operation. See Turgutlu ¶¶ 42–44, 50–53, 58–63, 72–77, and 88; Figs. 1 and 5–6);
synthesizing an output image based on the result information, the output image including the object placed at the location (Turgutlu teaches inserting the selected additional object into the input image using the predicted position and scale information to produce a composite image. Turgutlu describes placing one or more objects at predicted bounding box locations, adjusting object scale to fit the existing scene, and applying alpha composition to generate the output image. See Turgutlu ¶¶ 45–48, 54, 93, 96, 112, and 121; Figs. 1–3 and 7–9);
Turgutlu therefore teaches the limitations of claim 19 except that Turgutlu does not expressly teach the claimed a first mode in which a machine-trained model identifies the object based on input that specifies the location.
However, in the same field of endeavor, Wang teaches a first mode in which a machine-trained model identifies the object based on input that specifies the location (Wang teaches a machine trained CategoryLocation convolutional network for selecting an object category based on a specified location in an existing scene. Wang represents the current scene using an image based scene representation (V(S)). Given a specified location ((x,y)), Wang generates a spatial attention mask centered at the specified location and supplies the scene representation and attention mask to the trained network. The network produces: pcat(c|S,x,y) (See Wang section 5.1), which represents the probability that object category (c) is appropriate for placement in scene (S) at the specified location ((x,y)). Wang therefore teaches identifying an object based on input specifying the location at which the object is to be placed. See Wang section 5.2 and Figs. 5–6. Wang further teaches evaluating object category probabilities over scene locations and using those probabilities to select an object category and location during iterative scene synthesis. See Wang section 5.2).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Turgutlu’s image compositing system to include Wang’s location conditioned object selection mode. Turgutlu and Wang are analogous art because both references concern machine-trained systems that determine a contextually appropriate object and spatial placement for supplementing an existing image or scene. A person having ordinary skill in the art would have had reasons to incorporate Wang’s location conditioned object selection technique into Turgutlu’s image compositing system to support a location first image editing workflow. Such a workflow would permit a user to designate a region or location in an input image and request that the system recommend an object suitable for placement at that location.
Allowable Subject Matter
Claims 1-18 are allowed.
The following is an examiner’s statement of reasons for allowance:
Regarding claim 1:
Claim 1 is considered allowable because the prior art of record does not teach or adequately suggest the claimed training process considered as an integrated whole.
In particular, claim 1 requires a machine trained model having weights produced by a training process that includes: removing objects from original images; using the model to predict the removed objects based on their locations in the original images; using the model to predict the locations of the removed objects based on the objects; and adjusting the model weights to improve subsequent predictions of both the removed objects and their locations.
The closest prior art considered includes
Turgutlu et al., US 2023/0274478 A1 (“Turgutlu”);
Kelkar et al., US 2024/0020954 A1 (“Kelkar”);
Dvornik et al., “Modeling Visual Context is Key to Augmenting Object Detection Datasets” (“Dvornik”);
Wang et al., “Deep Convolutional Priors for Indoor Scene Synthesis” (“Wang”); and
Ritchie et al., “Fast and Flexible Indoor Scene Synthesis via Deep Convolutional Generative Models” (“Ritchie”).
Turgutlu teaches a machine trained layout model that represents an object using a class token and placement tokens corresponding to an x-coordinate, a y-coordinate, a width, and a height. Turgutlu further teaches predicting object-class and placement information and masking tokens associated with an object during training. See Turgutlu ¶¶ 58–63, 67–88, and 122–127; Figs. 5–6 and 10–12. Turgutlu, however, does not teach removing the depicted object from the original image and using the resulting object-removed image to perform both of the reciprocal prediction tasks recited in claim 1. Masking an abstract class and placement token tuple does not establish the claimed use of object removal from original images as part of the recited reciprocal training process.
Kelkar teaches removing a foreground object from a training image, optionally inpainting the vacated region, and training a machine learning model using the resulting modified training image. Kelkar therefore demonstrates that object removal may be performed as part of model training. Kelkar’s training objective, however, is directed to producing an object agnostic representation of the background scene. Kelkar does not teach predicting the removed object based on its location, predicting the removed object’s original location based on the object, or adjusting model weights to improve both reciprocal predictions.
Dvornik teaches masking the pixels within an annotated object bounding box and training a visual context classifier to predict an object category associated with the known bounding box location. See Dvornik § 3.2 and Figs. 2–3. Dvornik therefore teaches an object from location prediction task using an object-removed image. Dvornik does not teach training the model to perform the reciprocal task of predicting the original location of a removed object based on the object. Dvornik also does not teach adjusting the model weights based on errors from both reciprocal prediction tasks.
Wang teaches creating partial scenes by removing objects and training a CategoryLocation network to predict the category of an object associated with a specified location. See Wang §§ 4–5.2 and Figs. 5–6. Wang further evaluates object category scores across candidate locations during scene synthesis. Wang does not, however, clearly teach a second training task in which the removed object is provided to the model, the model predicts the object’s original removed location, and the model weights are adjusted based on a comparison between the predicted location and the original location.
Ritchie teaches a category conditioned location module that receives an object category and predicts a spatial distribution of locations for missing objects. Ritchie trains that location module using the locations of objects removed from training scenes. See Ritchie § 3.2. Ritchie, however, employs separate decision modules and does not teach producing the weights of the claimed model through a training process that includes both the object from location task and the reciprocal location from object task recited in claim 1.
Regarding claim 16:
Claim 16 is considered allowable for similar reasons to those of claim 1 and because it more particularly recites the loss generation and weight adjustment operations of the training process.
Claim 20 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WASSIM MAHROUKA whose telephone number is (571)272-2945. The examiner can normally be reached Monday-Thursday 8:00-5:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen Koziol can be reached at (408) 918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WASSIM MAHROUKA/Primary Examiner, Art Unit 2665