Prosecution Insights
Last updated: October 01, 2026
Application No. 18/713,602

IMAGE CROPPING METHOD AND APPARATUS, MODEL TRAINING METHOD AND APPARATUS, ELECTRONIC DEVICE, AND MEDIUM

Final Rejection §103
Filed
May 24, 2024
Priority
Nov 24, 2021 — CN 202111407110.6 +1 more
Examiner
CROCKETT, JOSHUA BRIGHAM
Art Unit
2661
Tech Center
2600 — Communications
Assignee
Beijing Bytedance Network Technology Co., Ltd.
OA Round
2 (Final)
84%
Grant Probability
Favorable
3-4
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
31 granted / 37 resolved
+21.8% vs TC avg
Strong +18% interview lift
Without
With
+17.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
13 currently pending
Career history
53
Total Applications
across all art units

Statute-Specific Performance

§101
8.9%
-31.1% vs TC avg
§103
45.1%
+5.1% vs TC avg
§102
9.8%
-30.2% vs TC avg
§112
35.3%
-4.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 37 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in the instant application. Information Disclosure Statement The information disclosure statement (IDS) submitted on 3 August 2026 was received and the information disclosure statement has been considered by the examiner. Response to Arguments Claims 1, 2, 6, 11-13, 17, and 20 are amended. Claims 3, 14, and 21 are canceled. Claims 1-2, 4-6, 11-13, 15-17, 20, and 22 are pending in this action. Applicant’s arguments, see pg. 9-11, filed 10 July 2026, with respect to the rejection(s) of claim(s) 1-6, 11-17, and 20-22 under 35 U.S.C. 103 have been fully considered and are persuasive. Specifically, the applicant argues that Yuan et al. (CN113159028A; hereafter, Yuan) and Lu et al. (CN109146892A; hereafter, Lu) do not disclose or reasonably suggest identifying a bounding box covering all pixels belonging to a semantic class of an object. The applicant further argues that while Lu uses the term “semantics” in their disclosure, they do not disclose identifying pixels belonging to a semantic class and instead relies on a saliency probability map. The examiner agrees. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Csurka (US 20090208118 A1). Csurka discloses: determining, based on position coordinates of pixels in the first segmented image ([0082] and Fig. 3C, relevant region 44 is identified as a region of interest which is understood as a segmentation of the first image) that belong to a semantic class of a first object ([0058] semantic values are determined for pixels based on classes. [0112] examples of classes include statues, castles, and people, therefore, classes are understood as a semantic class) , a bounding box of the first object in the first image ([0077] and Fig. 3C, subpart 40 is generated around the region 44 which is understood as a bounding box), wherein the bounding box encloses all pixels of the first object ([0077] and Fig. 3C, the bounding box 40 "includes all" of the region 44 which is understood as encloses all pixels of the first object); generating a plurality of first candidate boxes within the bounding box ([0093] and Fig. 3C, "In the above description, it is assumed that a given ROI 44 leads to one single crop, by pre-selecting one of the criteria for each step. In other embodiments, several of the above criteria may be employed to generate a set of potential crops, one of which is then selected based on further selection criteria." This is understood as determining a plurality of potential crops 44 within the region 40 which is generating candidate boxes within the bounding box), The complete rejection, including motivations to combine, is included below in the section “Claim Rejections - 35 USC § 103”. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 4, 11-13, 15, 20, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 20210392278 A1; hereafter, Zhang) in view of Csurka (US 20090208118 A1). Regarding claim 1, Zhang discloses: An image cropping method, comprising: segmenting an image to be cropped to obtain a first segmented image ([0036] a mask of an object is identified. A mask is understood as a segmentation as it differentiates the object from the rest of the image. See Fig. 2 and [0056] how masks 218A and 218B are segmentations of the dogs), and determining a bounding box of the first object in the first image ([0037] a bounding box is generated around the object based on the segmentation); generating a plurality of first candidate boxes ([0039] candidate cropping boxes are generated based on the hotspot and object mask. [0038] the hotspot and object mask in subsequent frames are based on the bounding box. Therefore, the candidate boxes are based on the bounding boxes) and selecting a first target box from the plurality of first candidate boxes ([0039] a candidate crop box is selected) based on a score of a first feature map corresponding to each first candidate box ([0039] the selection is based on a composition score. [0055] composition scores may be determined by a variety of ways which consider the location of features, such as the hotspot map and the object mask map, in the crop candidate. For example, at least the "rule of thirds technique" and the "triangle composition technique" consider the location of features, such as the center of gravity, in the candidate crop box which may be understood as a feature map because they map the location of the feature in the candidate crop box. The resulting composition score is therefore a score of a first feature map); and using an image located within the first target box in the image to be cropped as a cropping result ([0039] and [0055] the selected candidate cropping box is output which is understood as using an image in the box as a cropping result). Zhang does not disclose expressly determining a bounding box based on pixels belonging to a semantic class of a first object and that the plurality of first candidate boxes are generated within the bounding box. Csurka discloses: determining, based on position coordinates of pixels in the first segmented image ([0082] and Fig. 3C, relevant region 44 is identified as a region of interest which is understood as a segmentation of the first image) that belong to a semantic class of a first object ([0058] semantic values are determined for pixels based on classes. [0112] examples of classes include statues, castles, and people, therefore, classes are understood as a semantic class) , a bounding box of the first object in the first image ([0077] and Fig. 3C, subpart 40 is generated around the region 44 which is understood as a bounding box), wherein the bounding box encloses all pixels of the first object ([0077] and Fig. 3C, the bounding box 40 "includes all" of the region 44 which is understood as encloses all pixels of the first object); generating a plurality of first candidate boxes within the bounding box ([0093] and Fig. 3C, "In the above description, it is assumed that a given ROI 44 leads to one single crop, by pre-selecting one of the criteria for each step. In other embodiments, several of the above criteria may be employed to generate a set of potential crops, one of which is then selected based on further selection criteria." This is understood as determining a plurality of potential crops 44 within the region 40 which is generating candidate boxes within the bounding box), Zhang and Csurka are combinable because they are from the same field of endeavor of saliency aware image cropping (Zhang, [0014]; Csurka, [0015]). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention to combine the bounding box based on semantic class and the candidate boxes of Csurka with the invention of Zhang. The motivation for doing so would have been "to identify the best crop" (Csurka, [0084]). Therefore, it would have been obvious to combine Csurka with Zhang to obtain the invention as specified in claim 1. Regarding claim 2, Zhang in view of Csurka discloses the subject matter of claim 1. Zhang further discloses: The method according to claim 1, wherein the segmenting an image to be cropped to obtain a first segmented image comprises: performing feature extraction on the image to be cropped to obtain a second feature map ([0036] "The gaze-to-mask model 114 may then encode a concatenation (e.g., combination or integration) of the image and the hotspot map to extract features from the image and the hotspot map." Therefore, features are extracted from the image which is understood as a second feature map), and performing feature reconstruction on the second feature map using a segmentation model, to obtain the first segmented image ([0036] "the encoded concatenation is then decoded using residual refinement blocks, up-sampling at each block, to produce the object mask." The decoding is understood as a reconstruction which produces the object mask, i.e. segmented image); and before the selecting a first target box from the first candidate boxes based on a score of a first feature map corresponding to the first candidate box, the method further comprises: determining the first feature map based on the second feature map and the first candidate boxes ([0055] the identified first feature map of claim 1 is based on the hotspot map and the object mask map and the location of those maps in the candidate crop box. [0036] the object mask map is based on extracted features understood as the second feature map. Therefore, the first feature map is understood to be based on the second feature map, by being based on the object mask map and hotspot map, and the candidate crop box). Regarding claim 4, Zhang in view of Csurka discloses the subject matter of claim 1. Zhang further discloses: The method according to claim 1, wherein the generating a plurality of first candidate boxes within the bounding box comprises at least one of: generating, based on an input crop ratio, a plurality of first candidate boxes conforming to the crop ratio ([0039] "For example, a user may define a size (e.g., 50% of original) and an aspect ratio (e.g., 16:9) for a cropping box, and a machine learning module may generate a set of candidate crop boxes" The aspect ratio is understood as an input crop ratio); and generating, based on an input cropping precision, a number of first candidate boxes corresponding to the cropping precision ([0039] "For example, a user may define a size (e.g., 50% of original) and an aspect ratio (e.g., 16:9) for a cropping box, and a machine learning module may generate a set of candidate crop boxes" The defined size is understood as an input cropping precision because the defined size of the crop determines the possible number of candidate crop boxes to generate. For example, a system would be able to generate fewer boxes at a larger size before becoming redundant and would be able to generate a greater number of boxes at a smaller size. The number of boxes generated in the image is understood as a crop precision. Therefore, the input size requirement is understood as a precision) Zhang does not disclose expressly that the candidate boxes are generated withing the bounding box. Csurka discloses: generating a plurality of first candidate boxes within the bounding box ([0093] and Fig. 3C, "In the above description, it is assumed that a given ROI 44 leads to one single crop, by pre-selecting one of the criteria for each step. In other embodiments, several of the above criteria may be employed to generate a set of potential crops, one of which is then selected based on further selection criteria." This is understood as determining a plurality of potential crops 44 within the region 40 which is generating candidate boxes within the bounding box), It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention to combine the bounding box contained candidate boxes of Csurka with the invention of Zhang. The motivation for doing so would have been "to identify the best crop" (Csurka, [0084]). Therefore, it would have been obvious to combine Csurka with Zhang to obtain the invention as specified in claim 4. Regarding claim 11, claim 11 recites a system with elements corresponding to the steps recited in claim 1. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 1. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 1, apply to this claim. Finally, Zhang discloses: An electronic device, comprising: at least one processor ([0072] the device includes a processor); and a storage apparatus configured to store at least one program, wherein the at least one program, when executed by the at least one processor, causes the at least one processor to: ([0074] the device includes a storage device storing instructions for performing the method) Regarding claim 12, claim 12 recites a system with elements corresponding to the steps recited in claim 1. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 1. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 1, apply to this claim. Finally, Zhang discloses: A non-transitory storage medium comprising computer-executable instructions, wherein the computer-executable instructions ([0074] the device includes a storage device storing instructions for performing the method) when executed by a computer processor, are used to perform ([0072] the device includes a processor): Regarding claim 13, claim 13 recites a system with elements corresponding to the steps recited in claim 2. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 2. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 2, apply to this claim. Regarding claim 15, claim 15 recites a system with elements corresponding to the steps recited in claim 4. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 4. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 4, apply to this claim. Regarding claim 20, claim 20 recites a system with elements corresponding to the steps recited in claim 2. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 2. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 2, apply to this claim. Regarding claim 22, claim 22 recites a system with elements corresponding to the steps recited in claim 4. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 4. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 4, apply to this claim. Claims 5, 6, 16, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 20210392278 A1; hereafter, Zhang) in view of Csurka (US 20090208118 A1) in further view of Machefer et al. (US 20220391615 A1; hereafter, Machefer). Regarding claim 5, Zhang in view of Csurka discloses the subject matter of claim 1. Zhang further discloses: such that the scoring model outputs the score of each first feature map ([0055] the crop suggestion module may be understood as a scoring model as it outputs a candidate score for each of the crop candidates, which is understood as a score for the first feature map as shown in the teaching of claim 1). Zhang in view of Csurka does not disclose expressly to input feature maps into the scoring model in batches based on throughput. Machefer discloses: The method according to claim 1, wherein after the generating a plurality of first candidate boxes, the method further comprises: inputting, based on a single throughput of a scoring model (the examiner interprets throughput as the amount of input and output performed by a model. [0078] the number of the ROIs, which may be understood as bounding boxes, determines the size of a batch. Therefore, the batch is based on the amount of input, i.e. throughput), the plurality of first feature maps respectively corresponding to the plurality of first candidate boxes into the scoring model in batches ([0078] "the input of the convolution layers is a collection of same size squared feature maps, and the size of the batch is the number of ROIs." Therefore, the feature maps are input in batches), Machefer is combinable with Zhang in view of Csurka because it solves a similar problem of performing object analysis using feature maps (Machefer, [0015]). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the batch input into a model of Machefer with the invention of Zhang in view of Csurka. The motivation would have been that doing so is the use of a known technique, the batching of Machefer, to improve a similar device, the scoring model of Zhang in view of Csurka, in the same way, performing analysis by the model according to the capability of the model. The scoring model of Zhang in view of Csurka contains a base device because the scoring model is described by a term known in the art to describe a "base" neural network, "The crop suggestion module 220, using a neural network (e.g., DCNN−“Deep Convolutional Neural Network”)," (Zhang, [0055]). Machefer contains a comparable device, "An FPN is a feature extractor that takes a single-scale image of an arbitrary size as input, and outputs proportionally sized feature maps at multiple levels, in a fully convolutional fashion" (Machefer, [0059]), that has been improved in the same way as the claimed invention by the use of inputting batches into the model (see the application of Machefer [0078] above). A person of ordinary skill in the art could have applied the improvement of Machefer in the same way to the base device of Zhang in view of Csurka because both devices are neural network models which receive input data and, regardless of the inner workings of the models, the input processes would be understood as similar, i.e. preparing data size and entering data. The results would have been predictable to a person of ordinary skill in the art because a person of ordinary skill in the art would have experience working with neural network models and would understand the throughput of a model and how the improvement of Machefer takes the throughput into account when running the model. Therefore, it would have been obvious to combine Machefer with Zhang in view of Csurka to obtain the invention as specified in claim 5. Regarding claim 6, Zhang in view of Csurka in further view of Machefer discloses the subject matter of claim 5. Zhang in view of Csurka does not disclose expressly to resize the feature maps to a preset size. Machefer discloses: The method according to claim 5, wherein before the inputting the plurality of first feature maps respectively corresponding to the plurality of first candidate boxes into the scoring model in batches, the method further comprises: resizing the first feature maps corresponding to the first candidate boxes to a preset size ([0078] "pool all feature maps from FPN lying within the ROIs by discretising them into a set of fixed square pooled size bins". The "fixed square pooled size" is understood as a preset size. Discretising the feature maps is understood as resizing). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention to combine the resizing of Machefer with the invention of Zhang in view of Csurka. The motivation would have been that doing so is the use of a known technique, the resizing of Machefer, to improve a similar device, the scoring model of Zhang in view of Csurka, in the same way, performing analysis by the model according to the capability of the model. The scoring model of Zhang in view of Csurka contains a base device because the scoring model is described by a term known in the art to describe a "base" neural network, "The crop suggestion module 220, using a neural network (e.g., DCNN−“Deep Convolutional Neural Network”)," (Zhang, [0055]). Machefer contains a comparable device, "An FPN is a feature extractor that takes a single-scale image of an arbitrary size as input, and outputs proportionally sized feature maps at multiple levels, in a fully convolutional fashion" (Machefer, [0059]), that has been improved in the same way as the claimed invention by the use of resizing input into the model (see the application of Machefer [0078] above). A person of ordinary skill in the art could have applied the improvement of Machefer in the same way to the base device of Zhang in view of Csurka because both devices are neural network models which receive input data and, regardless of the inner workings of the models, the input processes would be understood as similar, i.e. preparing data size and entering data. The results would have been predictable to a person of ordinary skill in the art because a person of ordinary skill in the art would understand that neural network models require specific input data size for the convolutional layers to function and would understand that the resizing of Machefer would produce the predictable result of a neural network model operating as designed. Therefore, it would have been obvious to combine Machefer with Zhang in view of Csurka to obtain the invention as specified in claim 6. Regarding claim 16, claim 16 recites a system with elements corresponding to the steps recited in claim 5. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 5. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 5, apply to this claim. Regarding claim 17, claim 17 recites a system with elements corresponding to the steps recited in claim 6. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim, claim 6. Additionally, the rationale and motivation to combine Zhang in view of Csurka, presented in rejection of claim 6, apply to this claim. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Martin et al. (US 20190361522 A1) discloses a bounding box generated around an object belonging to a semantic class. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSHUA B CROCKETT whose telephone number is (571)270-7989. The examiner can normally be reached Monday-Thursday 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JOSHUA B. CROCKETT/Examiner, Art Unit 2661 /AARON W CARTER/Primary Examiner, Art Unit 2661
Read full office action

Prosecution Timeline

May 24, 2024
Application Filed
Apr 10, 2026
Non-Final Rejection mailed — §103
Jul 10, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749309
DYNAMIC SERVICE OF IMAGE METADATA
3y 11m to grant Granted Sep 29, 2026
Patent 12737949
SYSTEMS AND METHODS OF ACCELERATED DYNAMIC IMAGING IN PET
4y 3m to grant Granted Sep 15, 2026
Patent 12725238
METHODS, APPARATUSES AND COMPUTER PROGRAM PRODUCTS FOR DEPALLETIZING MIXED OBJECTS
3y 11m to grant Granted Sep 01, 2026
Patent 12693653
MACHINE LEARNING BASED CYCLE TIME TRACKING AND REPORTING FOR VEHICLES
2y 6m to grant Granted Jul 28, 2026
Patent 12693424
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING SYSTEM AND IMAGE PROCESSING METHOD
2y 10m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+17.6%)
3y 2m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 37 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month