Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 4-12, and 14-20 are rejected under 35 U.S.C. 103 as being unpatentable over Tripathi (10909349) in view of Maschmeyer (20240161258) in further view of Pandya (11803955).
Regarding claim 1, Tripathi teaches A method, comprising: providing, by the one or more processors, the training data image as input to at least one computer vision neural network to configure the at least one computer vision neural network (col. 11 lines 31-50 and col. 12 lines 52-65, synthetic images train object detector).
Tripathi does not teach receiving a prompt identifying characteristics of the training data image, generating that image using a generative AI model, or preconfiguring that model using images.
Maschmeyer teaches receiving, by one or more processors, a prompt identifying one or more characteristics of a training data image (pars. 49, 51 and 69, receives natural language image prompt).
generating, by the one or more processors using at least one generative artificial intelligence (AI) model, the training data image based on the one or more characteristics (pars. 3, 51 and 68-72, text to image model generates output).
wherein the at least one generative AI model is pre-configured using images (pars. 44, 54, 56, 86 and 90, reference images customize diffusion model).
It would have been obvious prior to the effective filing date of the invention to one of ordinary skill in the art to include in Tripathi the prompt and reference image configured generative model taught by Maschmeyer. The reason is to produce diverse photorealistic training images from natural language scene requirements.
Tripathi and Maschmeyer do not teach that the image characteristics correspond to workplace risk criteria involving objects, that the reference images depict equipment in building or workplace scenes, or that the configured computer vision neural network performs risk condition detection.
Pandya teaches wherein the one or more characteristics correspond to a risk criteria associated with one or more objects in at least one of a building or a workplace to be represented in the training data image (col. 1 lines 35-55 and col. 2 lines 15-35, workplace objects define safety risk).
of items of equipment in one or more scenes of one or more of an example building or an example workplace (col. 7 lines 15-48 and col. 17 lines 20-34, workplace scenes contain equipment).
to perform computer vision-based risk condition detection of received image data (col. 18 lines 37-47 and col. 20 lines 5-25, trained model detects unsafe conditions).
It would have been obvious prior to the effective filing date of the invention to one of ordinary skill in the art to include in Tripathi and Maschmeyer the workplace scene characteristics and risk detection task taught by Pandya. The reason is to train the detector on safety arrangements of workers and equipment.
Regarding claim 2, see Pandya, col. 3 lines 30-45 and col. 17 lines 20-34, worker-equipment geometry indicates safety.
Regarding claim 4, see Pandya, col. 2 lines 37-49 and col. 3 lines 30-45, scene model captures object relationships.
Regarding claim 5, Maschmeyer teaches wherein receiving the prompt comprises: generating, by the one or more processors, a query associated with the one or more characteristics (pars. 80-81, suggests prompt modification to user).
receiving, by the one or more processors, a response to the query, wherein the prompt is based at least in part on the response (par. 81, receives user-modified text prompt).
Regarding claim 6, see Maschmeyer teaches generating, by the one or more processors using the at least one generative AI model, an additional training data image, wherein the additional training data image is a variant of the training data image (pars. 71-72 and 82-84; claims 1 and 10, varies prompt seed and parameters).
Tripathi teaches configuring the at least one computer vision neural network further based on the additional training data image (col. 13 lines 8-30 and col. 19 lines 9-27, additional images retrain detector).
Regarding claim 7, see Tripathi col. 12 lines 40-51 and col. 13 lines 40-55, stores known image annotations.
Regarding claim 8, see Tripathi col. 2 lines 10-18 and col. 13 lines 47-55, bounding box identifies object pixels.
Regarding claims 9-10 and 12, see the rejection of claim 1. These claims are broader than claim 1.
Regarding claim 11, see the rejection of claim 2.
Regarding claim 14, see the rejection of claim 4.
Regarding claim 15, see the rejection of claim 5.
Regarding claim 16, see Maschmeyer pars. 71-72 and 82-84, varies prompt seed and parameters.
Regarding clam 17, see the rejection of claim 7.
Regarding claim 18, see the rejection of claim 8.
Regarding claim 19, Maschmeyer teaches receiving, by one or more processors, using a conversational interface, an input indicative of one or more characteristics of a person and an object to represent in a scene in a synthetic image (pars. 49, 69 and 80-81, exchanges prompt suggestions with user).
Pandya teaches the scene comprising at least one of an example building or an example workplace, the one or more characteristics corresponding to one or more risk criteria associated with the person and the object (col. 7 lines 15-48 and col. 17 lines 20-47, worker-object scene defines risk).
Maschmeyer teaches providing, by the one or more processors, the input to a generative artificial intelligence (GAI) model to cause the GAI model to generate the synthetic image according to the one or more characteristics (pars. 49, 51 and 68-72; claims 14-16, prompt causes generated image).
Tripathi teaches training, by the one or more processors, a computer vision model using the synthetic image and the one or more characteristics to configure the computer vision model (col. 12 line 52 to col. 13 line 8 and col. 13 lines 40-60, annotations train object detector).
Pandya teaches to perform risk condition detection relating to image data captured of at least one of a building or a workplace (col. 18 lines 17-47 and col. 20 lines 5-25, detects unsafe workplace conditions).
Regarding claim 20, see the rejection of claim 6.
Claim 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Tripathi (10909349) in view of Maschmeyer (20240161258) in view of Pandya (11803955) in further view of Lin (20220036127).
Regarding claim 3, Maschmeyer teaches generating, by the one or more processors using the first generative AI model, a first image (pars. 3, 51 and 68-72, text to image model generates image).
Lin teaches modifying, by the second generative AI model, the first image to generate the training data image (pars. 6, 45, 62 and 100, trained decoder generates edited image.
It would have been obvious prior to the effective filing date of the invention to one of ordinary skill in the art to include in Tripathi, Maschmeyer, and Pandya the second semantic image generation model as taught by Lin. The reason is to modify selected visual attributes of a first generated.
Regarding claim 13, see the rejection of claim 3.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Kumar (20230244996) teaches construction site worker safety examples combined with local backgrounds for machine learning retraining.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HADI AKHAVANNIK whose telephone number is (571)272-8622. The examiner can normally be reached 9 AM - 5 PM Monday to Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HADI AKHAVANNIK/Primary Examiner, Art Unit 2676