Prosecution Insights
Last updated: August 16, 2026
Application No. 19/013,214

IMAGE EDITING METHOD AND APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM

Non-Final OA §101§102§103
Filed
Jan 08, 2025
Priority
Jan 25, 2024 — CN 202410107990.2
Examiner
MCDOWELL, JR, MAURICE L
Art Unit
Tech Center
Assignee
Lemon Inc.
OA Round
1 (Non-Final)
87%
Grant Probability
Favorable
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
809 granted / 934 resolved
+26.6% vs TC avg
Moderate +13% lift
Without
With
+13.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
16 currently pending
Career history
944
Total Applications
across all art units

Statute-Specific Performance

§101
17.3%
-22.7% vs TC avg
§103
50.7%
+10.7% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
8.0%
-32.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 934 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification Title of the Invention The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: IMAGE EDITING METHOD AND APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM TO EDIT AN IMAGE BASED ON AN EDITING INSTRUCTION AND EDITING POSITION TO OBTAIN A TARGET IMAGE Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-7 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because claim 1 is directed to an image editing method comprising the steps of receiving, generating and editing which are nothing more than software instructions. Software instructions are non-statutory under 35 U.S.C. 101. Claims 2-7 depend from claim 1 and comprise additional steps, for example claim 2 comprises the steps of obtaining, extracting, generating, determining and constructing, therefore claims 2-7 have the same problem as claim 1 and are rejected under the same rationale. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1, 8 and 15 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by FU et. al., “Guiding Instruction-based Image Editing via Multimodal Large Language Models,” pgs. 1-14, Published: 16 Jan 2024, Last Modified: 05 Mar 2024ICLR 2024, Guiding Instruction-based Image Editing via Multimodal Large Language Models | OpenReview Regarding claim 1, FU teaches: 1. An image editing method, comprising: receiving an image to be edited and an editing theme (FU fig. 1, (the bottom left images of cheetah and her cubs clearly shows input image (i.e., received image to be edited and editing theme (i.e., “add contrast to simulate more light”); generating, by using a preset vision-language model, an editing instruction and an editing position corresponding to the editing instruction based on the image to be edited and the editing theme (FU, fig. 1, and caption; the examiner interprets the MLLM-Guided Image Editing (MGIE) as the preset vision-language model; the MGIE (Expressive Instruction): “The light enhances the detail of the mother cheetah and her cubs on the rock hillside” is interpreted as the editing instruction and editing position corresponding to the editing instruction which are clearly based on the image to be edited and the editing theme as discussed supra); and editing the image to be edited based on the editing instruction and the editing position, to obtain a target image (FU, fig. 1 and caption; The right image of the cheetah and her cubs are interpreted as the target image). Claim 8 is analogous to claim 1 and is therefore rejected using the same rationale. Claim 8 further requires a different preamble and two additional limitations, that are also taught by FU. 8. An electronic device, comprising: one or more processors; and a storage apparatus configured to store one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to (FU, pg. 6, Implementation Details. see last line, “All experiments are conducted in PyTorch on 8 A100 GPUs,” POSITA would recognize that conducting experiments in PyTorch on 8 A100 GPUs comprise an electronic device comprising one or more processors; and a storage apparatus configured to store one or more programs, wherein the one or more programs, are executed by the one or more processors) Claim 15 is analogous to claim 1 and is therefore rejected using the same rationale. Claim 15 further requires a different preamble, that is also taught by FU. 15. A non-transitory storage medium containing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to perform steps (FU, pg. 6, Implementation Details. see last line, “All experiments are conducted in PyTorch on 8 A100 GPUs,” POSITA would recognize that conducting experiments in PyTorch on 8 A100 GPUs comprise a non-transitory storage medium containing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to perform steps). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 6, 13 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over FU in view of SOMMERLADE (US2019/0051057A1). Regarding claim 6, FU doesn’t teach however the analogous prior art SOMMERLADE teaches: 6. The method according to claim 1, wherein after the generating an editing instruction and an editing position corresponding to the editing instruction, the method further comprises: receiving a selection operation for a target editing instruction, to determine the target editing instruction from at least two editing instructions (SOMMERLADE: par. 156). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine receiving a selection operation for a target editing instruction, to determine the target editing instruction from at least two editing instructions as shown in SOMMERLADE with FU for the benefit of implementing processing techniques with minimal demands on computer hardware and/or power such that they provide results at or near input data frame rate or user feedback requirements [SOMMERLADE, par. 4 lines 5-10]. Claim 13 is analogous to claim 6 and is therefore rejected using the same rationale. Claim 20 is analogous to claim 6 and is therefore rejected using the same rationale. Claim(s) 7 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over FU in view of WANG (CN111708597A). Regarding claim 7, FU doesn’t teach however the analogous prior art WANG teaches: 7. The method according to claim 1, wherein after the generating an editing instruction and an editing position corresponding to the editing instruction, the method further comprises: receiving an adjustment operation for the editing position, to adjust the editing position (WANG: pg. 7 lines 30-33). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine receiving an adjustment operation for the editing position, to adjust the editing position as shown in WANG with FU for the benefit of reducing the information amount needed to be sent when processing the cooperative information, which is good for improving the stability of the information cooperative system [WANG, abstract lines 10-12]. Claim 14 is analogous to claim 7 and is therefore rejected using the same rationale. Allowable Subject Matter Claims 2-5 would be objected to (except for the 101 rejection), claims 9-12 and 16-19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claims 2-5, 9-12 and 16-19 the prior art doesn’t teach: 2. The method according to claim 1, wherein the preset vision-language model comprises a pre-trained model adjusted based on an instruction dataset; and wherein a construction process of the instruction dataset comprises: obtaining a sample image and a sample editing theme; extracting an object and an object position from the sample image; generating a global image description and a local object description based on the sample image, the object, and the object position; determining a sample target object, a sample associated object, and a sample editing instruction based on the global image description, the local object description, the sample editing theme, and a preset list of theme-associated objects; wherein the sample target object is contained in the sample image, the sample associated object is contained in the list of theme-associated objects, and the sample editing instruction is used to describe an editing operation performed on the sample target object based on the sample associated object; and constructing the instruction dataset based on the sample image, the sample editing theme, the sample editing instruction, the sample associated object, the sample target object, and an object position of the sample target object. 3. The method according to claim 2, wherein after the determination of the sample editing instruction, the method further comprises: editing the sample image according to the sample editing instruction, to obtain a sample target image; and filtering the sample editing instruction based on a similarity between the sample target image and the sample editing theme. 4. The method according to claim 1, wherein the generating, by using a preset vision-language model, an editing instruction and an editing position corresponding to the editing instruction based on the image to be edited and the editing theme comprises: performing, by using the preset vision-language model, feature extraction on the image to be edited, to obtain an implicit image feature; generating a token sequence of the editing instruction based on the image to be edited and the editing theme, wherein the token sequence comprises a spatial token; and decoding the token sequence into the editing instruction, and generating the editing position corresponding to the editing instruction based on the implicit image feature and the spatial token. 5. The method according to claim 4, wherein the generating the editing position corresponding to the editing instruction based on the implicit image feature and the spatial token comprises: performing a cross-attention calculation on the implicit image feature and the spatial token, and predicting the editing position corresponding to the editing instruction based on a result of the calculation. 9. The electronic device according to claim 8, wherein the preset vision-language model comprises a pre-trained model adjusted based on an instruction dataset; and wherein the one or more programs for a construction process of the instruction dataset further comprise one or more programs which, when executed by the one or more processors, cause the one or more processors to: obtain a sample image and a sample editing theme; extract an object and an object position from the sample image; generate a global image description and a local object description based on the sample image, the object, and the object position; determine a sample target object, a sample associated object, and a sample editing instruction based on the global image description, the local object description, the sample editing theme, and a preset list of theme-associated objects; wherein the sample target object is contained in the sample image, the sample associated object is contained in the list of theme-associated objects, and the sample editing instruction is used to describe an editing operation performed on the sample target object based on the sample associated object; and construct the instruction dataset based on the sample image, the sample editing theme, the sample editing instruction, the sample associated object, the sample target object, and an object position of the sample target object. 10. The electronic device according to claim 9, wherein after the determination of the sample editing instruction, the one or more programs further cause the one or more processors to: edit the sample image according to the sample editing instruction, to obtain a sample target image; and filter the sample editing instruction based on a similarity between the sample target image and the sample editing theme. 11. The electronic device according to claim 8, wherein the one or more programs for the generating, by using a preset vision-language model, an editing instruction and an editing position corresponding to the editing instruction based on the image to be edited and the editing theme further comprise one or more programs which, when executed by the one or more processors, cause the one or more processors to: perform, by using the preset vision-language model, feature extraction on the image to be edited, to obtain an implicit image feature; generate a token sequence of the editing instruction based on the image to be edited and the editing theme, wherein the token sequence comprises a spatial token; and decode the token sequence into the editing instruction, and generate the editing position corresponding to the editing instruction based on the implicit image feature and the spatial token. 12. The electronic device according to claim 11, wherein the one or more programs for the generating the editing position corresponding to the editing instruction based on the implicit image feature and the spatial token further comprise one or more programs which, when executed by the one or more processors, cause the one or more processors to: perform a cross-attention calculation on the implicit image feature and the spatial token, and predict the editing position corresponding to the editing instruction based on a result of the calculation. 16. The non-transitory storage medium according to claim 15, wherein the preset vision- language model comprises a pre-trained model adjusted based on an instruction dataset; and wherein the computer-executable instructions used for a construction process of the instruction dataset further comprise computer-executable instructions which, when executed by the computer processor, are used to: obtain a sample image and a sample editing theme; extract an object and an object position from the sample image; generate a global image description and a local object description based on the sample image, the object, and the object position; determine a sample target object, a sample associated object, and a sample editing instruction based on the global image description, the local object description, the sample editing theme, and a preset list of theme-associated objects; wherein the sample target object is contained in the sample image, the sample associated object is contained in the list of theme-associated objects, and the sample editing instruction is used to describe an editing operation performed on the sample target object based on the sample associated object; and construct the instruction dataset based on the sample image, the sample editing theme, the sample editing instruction, the sample associated object, the sample target object, and an object position of the sample target object. 17. The non-transitory storage medium according to claim 16, wherein after the determination of the sample editing instruction, the computer-executable instructions are further used to: edit the sample image according to the sample editing instruction, to obtain a sample target image; and filter the sample editing instruction based on a similarity between the sample target image and the sample editing theme. 18. The non-transitory storage medium according to claim 15, wherein the computer- executable instructions used for the generating, by using a preset vision-language model, an editing instruction and an editing position corresponding to the editing instruction based on the image to be edited and the editing theme further comprise computer-executable instructions which, when executed by the computer processor, are used to: perform, by using the preset vision-language model, feature extraction on the image to be edited, to obtain an implicit image feature; generate a token sequence of the editing instruction based on the image to be edited and the editing theme, wherein the token sequence comprises a spatial token; and decode the token sequence into the editing instruction, and generate the editing position corresponding to the editing instruction based on the implicit image feature and the spatial token. 19. The non-transitory storage medium according to claim 18, wherein the computer- executable instructions used for the generating the editing position corresponding to the editing instruction based on the implicit image feature and the spatial token further comprise computer- executable instructions which, when executed by the computer processor, are used to: perform a cross-attention calculation on the implicit image feature and the spatial token, and predict the editing position corresponding to the editing instruction based on a result of the calculation. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. WU (US2025/0200283A1) discloses approaches presented herein provide for the use of language models to generate tokenized descriptions of physical environments. In at least one embodiment, sensor and/or observational data can be obtained for an environment and used to generate a set of perception data. The perception data can be analyzed, along with approximate positional data within the environment, to identify a set of aligned map data. The aligned map data and perception data can be provided as input to a trained language model, which can be trained to correlate and/or fuse the information to generate a single, consistent representation of the environment. The language model can output a tokenized description of the environment, which can be in a domain-specific language, that is a compact but robust textual description of the environment; NARAYANA (US2025/0078361A1) discloses methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for enabling artificial intelligence to generate new images based on contextual data and to generate digital components based on the images. In one aspect, a method includes receiving one or more queries from a client device of a user. A digital component is selected based on the one or more queries. A customized digital component is generated by obtaining an image of an object corresponding to the selected digital component and generating, using a language model, an image editing prompt for editing the image based on digital component data related to the digital component and query data including the one or more queries and contextual data. The image and the image editing prompt are provided to an image editing model. An edited image is received and used to generate the customized digital component; LIU (US2024/0404145A1) discloses a method, an electronic device, and a computer program product for generating images. The method includes acquiring a descriptive text for describing image content of a target image, determining position prior information, and generating the target image based on the descriptive text and position prior information. According to the method of the embodiments of the present disclosure, a target image that can be used for rare data simulation in rare scenarios can be generated by means of an input descriptive text. In addition, the method allows for position perception editing and operation, and can control, based on position prior information, a direction and a position of an object generated in the target image, thereby effectively and diversely generating images. Moreover, the method provided in the present disclosure is based on object types in each subdivided image block, thus making position perception more accurate; ZHU (CN111832581B) relates to artificial intelligence field, the invention claims a lung feature recognition method, device, computer device and storage medium, the method comprises: obtaining the data to be identified including the lung image to be identified and the lung text to be identified; performing lung image feature extraction by lung image recognition model; generating lung image feature vector and image recognition result; at the same time, performing lung text feature extraction by lung text recognition model; generating lung text feature vector and text recognition result; using attention mechanism to fuse lung image feature vector and lung text feature vector by lung fusion identification model, and extracting image text fusion feature to identify, obtaining the fusion identification result; The lung feature recognition result is obtained by voting. The invention realizes accurately identifying the lung characteristic and improves the identification accuracy and reliability. The invention is suitable for the field of intelligent medical treatment and so on; it can further promote the construction of intelligent city; LIU (WO2021/190115A1) discloses a method and apparatus for searching for a target. A specific implementation mode of the method comprises: obtaining at least one image and a description text of a specified object; extracting image features of the image and text features of the description text by using a pre-trained cross-media feature extraction network; and matching the image features and the text features to determine an image that contains the specified object. Thus, features are extracted by using cross-media features, and the image features and the text features are projected to an image and text common feature space for feature matching, thereby achieving cross-media target search. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAURICE L MCDOWELL, JR whose telephone number is (571)270-3707. The examiner can normally be reached Mon-Fri: 2pm-10pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said A. Broome can be reached at 571-272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MAURICE L. MCDOWELL, JR/Primary Examiner, Art Unit 2612
Read full office action

Prosecution Timeline

Jan 08, 2025
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705427
PROVIDING DIVERSE VISUAL CONTENTS BASED ON PROMPTS
2y 9m to grant Granted Aug 11, 2026
Patent 12682546
RENDERING SYSTEM AND AUTOMATED DRIVING VERIFICATION SYSTEM
3y 10m to grant Granted Jul 14, 2026
Patent 12682563
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND NON-TRANSITORY COMPUTER-READABLE MEDIUM
2y 3m to grant Granted Jul 14, 2026
Patent 12675953
METHOD AND APPARATUS OF HAIR PROCESSING FOR VIRTUAL OBJECT, DEVICE, STORAGE MEDIUM, AND PROGRAM PRODUCT
1y 11m to grant Granted Jul 07, 2026
Patent 12626822
DEEP LEARNING FOR AUTOMATED SMILE DESIGN
2y 3m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
87%
Grant Probability
99%
With Interview (+13.0%)
2y 11m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 934 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month