DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of claims: claims 1-20 are pending below.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 10/2/2024 and 2/3/2025 was filed and considered. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 5-9, 11-12 and 15-19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by HARARY et al (US 2025/0005727).
Claim 1, similarly claim 11:
HARARY et al (US 2025/0005727) anticipated the following subject matter:
A computer-implemented method for object detection, comprising:
generating a negative description for an input image based on a positive description of the input image using a language model; generating a negative image based on the input image and the negative description by replacing a portion of the input image that is described by the positive description with content that is described by the negative description using a generative image model (abstract, figures 2-7, with supporting paragraphs 0004-0006, 0021, 0039-0062); and
training an object detection model with the input image, the positive description, the negative description, and the negative image (paragraph 0020 detail detection model may be developed by using a pre-trained vision-language model to transform training texts for an object (including positive and negative textual descriptions of the object) and a query image (input image) into a shared embedding space).
Regarding claim 11, HARARY et al teaches system in figures 1-5.
Claim 2, similarly claim 12:
The method of claim 1, wherein training the object detection model includes using the negative description with the input image and using the positive description with the negative image as negative examples (figures 2-4).
Claim 5, similarly claim 15:
The method of claim 1, further comprising fine-tuning the language model to generate negative descriptions based on a set of positive-negative description pairs (0066-0068, where 0066 detail pairing of data (negative 510 to normal 502), where 0068 further detail train model 555 with refine/tune the model 555 to achieve optimal performance).
Claim 6, similarly claim 16:
The method of claim 5, further comprising generating the positive-negative description pairs using a second language model (0039 detail use language model such as Cyclic Contrastive Language -Image Pretraining (CyCLIP) model as well as additional Bootstrapping Language -Image Pre-training (BLIP) model for unified vision- language understanding and generation.).
Claim 7, similarly claim 17:
The method of claim 6, wherein generating the negative description includes changing a word of the positive description (0051 detail changing wording on object for example including good, normal, standard...etc).
Claim 8, similarly claim 18:
The method of claim 6, wherein generating the negative description includes re-combining noun phases of the positive description (0051-0052, specially 0052-0054 detail combined reference vector combine with concept such as “normal object” to contains information from both images and texts).
Claim 9, similarly claim 19:
The method of claim 6, wherein generating the negative description includes extracting features that describe a difference between positive descriptions and corresponding negative descriptions (0061-0063 detail comparison module 452 (difference) between vectors (feature) between positive and negative).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3-4 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over HARARY et al (US 2025/0005727) in view of COHEN et al (2018/0322339).
Claim 3, similarly claim 13:
HARARY et al teaches all the subject matter above, but not the following:
The method of claim 1, wherein generating the negative image includes identifying a bounding box associated with the positive description and replacing content within the bounding box.
COHEN et al (2018/0322339) teaches: The method of claim 1, wherein generating the negative image includes identifying a bounding box associated with the positive description and replacing content within the bounding box (0065 detail boxing of text and figures and/or tables further to replacing such as text replace with figure with positive and negative data).
HARARY et al and COHEN et al are both in the field of image analysis, especially in object detection by means of language model with positive and negative description such that the combine outcome is predictable.
Therefore it would have been obvious to one having ordinary skill before the effective filing date to modify HARARY et al by COHEN et al using description language model (0003) such replacement using machine learning algorithms, and may therefore may be trained to self-improved or self-adjust over a number of iterations to provide improved page segmentation results as compared to prior solutions as disclosed by COHEN et al in paragraph 0023.
Claim 4, similarly claim 13:
COHEN et al
The method of claim 3, wherein replacing content includes using inpainting, conditioning, and text-to-image diffusion based on the negative description (0065 detail replacing with different element with positive or negative data with feedback to improve results, where 0074 detail segmentation (diffusion) as well as fusion of pixels between text, figure or table).
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over HARARY et al (US 2025/0005727) in view of Tang et al (2022/0172369).
Claim 10, similarly claim 20:
HARARY et al teaches all the subject matter above, but not the following:
The method of claim 1, further comprising employing the trained object detection model to identify an object within a driving scene and to perform a driving action responsive to an output of the trained object detection model.
Tang et al (2022/0172369) teaches the following subject matter: The method of claim 1, further comprising employing the trained object detection model to identify an object within a driving scene and to perform a driving action responsive to an output of the trained object detection model (figure 1 step 410, figure 6 and paragraph 0005-0009, paragraph 0003-0004; 0038-0039 and 0045).
HARARY et al and Tang et al are both in the field of image analysis, especially in object detection by means of language model with positive and negative description such that the combine outcome is predictable.
Therefore it would have been obvious to one having ordinary skill before the effective filing date to modify HARARY et al by Tang et al provide the system with the enhanced capacity of using the well know technology of using the input image, the positive description, the negative description and the negative image to training an object detection model.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Xu (US 2021/0133501) teaches METHOD AND APPARATUS FOR GENERATING VEHICLE DAMAGE IMAGE ON THE BASIS OF GAN NETWORK - system obtains a real vehicle image, generates an intermediate image based on the real vehicle image by labeling a target box on the real vehicle image and removing a portion of the real vehicle image within the target box, and generates the vehicle damage image based on the intermediate image by inputting the intermediate image into a machine-learning model, which outputs the vehicle damage image by filling a local image indicating vehicle damage into the target box of the intermediate image (abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TSUNG-YIN TSAI whose telephone number is (571)270-1671. The examiner can normally be reached 7am-4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at (571) 272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TSUNG YIN TSAI/Primary Examiner, Art Unit 2656