Prosecution Insights
Last updated: October 02, 2026
Application No. 18/562,784

ADAPTIVE OBJECT DETECTION

Non-Final OA §103
Filed
Nov 20, 2023
Priority
Jun 30, 2021 — nonprovisional of PCTCN2021103872
Examiner
YAO, JULIA ZHI-YI
Art Unit
2666
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Non-Final)
63%
Grant Probability
Moderate
2-3
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
53 granted / 84 resolved
+1.1% vs TC avg
Strong +48% interview lift
Without
With
+48.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
23 currently pending
Career history
107
Total Applications
across all art units

Statute-Specific Performance

§101
6.2%
-33.8% vs TC avg
§103
55.0%
+15.0% vs TC avg
§102
9.9%
-30.1% vs TC avg
§112
26.3%
-13.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 84 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendments In the amendments received on May 28th, 2026, claims 1, 3-4, 8, and 10-12 are amended. Accordingly, claims 1-15 are currently pending for examination in the Application No. 18/562,784 filed November 20th, 2023. Applicant’s amendments filed May 28th, 2026, to the Claims have overcome each and every objection, 35 U.S.C. § 112 (a) and (b) rejections, and claim interpretation previously set forth in the Non-Final Office Action mailed April 22nd, 2026. Accordingly, the objection(s), 35 U.S.C. § 112 (a) and (b) rejection(s), and claim interpretation(s) are withdrawn in response to the remarks and amendments filed. Examiner warmly thanks Applicant for considering the suggested amendments to be made to the disclosure. Response to Arguments Applicant’s arguments filed May 28th, 2026, regarding the rejection(s) of independent claim(s) have been fully considered but are moot because the arguments do not apply to the new combination of the references being used in the current rejection below. Priority (Previously Presented) Acknowledgment is made of applicant’s status as a U.S. National Stage Filing under 35 U.S.C. § 371 of International Application No. PCT/CN2021/103872, filed on June 30th, 2021. Information Disclosure Statement The information disclosure statement (IDS) submitted on July 30th, 2026, is in compliance with the provisions of 37 CFR 1.97. Accordingly, the IDS is being considered and attached by the examiner. Claim Objections Claim 7 is objected to because of the following informalities: The examiner respectfully suggests amending the phrase “on image to be captured by the camera” in claim 8 to recite “on an image to be captured by the camera” or similar to prevent confusion regarding antecedent basis of the “image” recited in the phrase. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-4, 6-11, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Yamasaki (US 2022/0262031 A1; previously cited in Non-Final Action mailed April 22nd, 2026) in view of Chaterji et al. (Chaterji; US 2022/0327826 A1; previously cited in Non-Final Action mailed April 22nd, 2026). Regarding claim 1, Yamasaki discloses a computer-implemented method, comprising: obtaining object distribution information associated with a set of historical images captured by a camera, the object distribution information indicating a size distribution of detected objects in the set of historical images (para(s). [0016], [0037], and [0041], recite(s) [0016] “There are increasing cases where an image captured by a monitoring camera or an image captured by the monitoring camera and then stored in a storage device is analyzed and utilized. …” [0037] “As illustrated in FIG. 3, the setting unit 203 sets a plurality of partial regions 301 a, 301 b, and 301 c in an image 300 input from the imaging apparatus 110. In the present exemplary embodiment, based on information regarding the size and the position of a person appearing at each of a plurality of different points on the image 300, the setting unit 203 sets the plurality of partial regions 301 a, 301 b, and 301 c in the image 300. The size of each partial region in the image captured by the imaging apparatus 110 depends on the size of the particular object appearing in the captured image. …” [0041] “Then, based on the size information f(x,y) regarding the size of the person at any position on the image, the setting unit 203 sets the plurality of partial regions in the image (the predetermined region in the input image). …” , where the “size information” is object distribution information associated with a set of historical images captured by a camera (e.g., images “captured by [a] monitoring camera and then stored in a storage device”) indicating a size distribution (i.e., “size and the position of a person appearing at each of a plurality of different points on the image”) of detected objects (e.g., “particular object[s]” and/or people) captured in the set of historical images) generating at least one detection plan based on the object distribution information(para(s). [0030] and [0047], recite(s) [0030] “…In the regression-based estimation method, using a regressor (a trained recognition model) to which a small image of a certain fixed size s is input and from which the number of particular objects present in the small image is output, the number of particular objects in each of the plurality of partial regions in the predetermined region in the input image is estimated. When the regressor is trained, many small images of the fixed size s in which the position of a particular object is known are prepared, and the regressor is trained in advance on these target small images as training data based on a machine learning technique. At this time, to improve the accuracy of estimating the number of particular objects, it is desirable that the ratio between the size (the fixed size s) of each small image as the training data and the size of the particular object present in the small image should be approximately constant. …” [0047] “First, in step S 401 , based on a trained model that estimates the presence of a particular object in an input image, the estimation unit 204 acquires likelihood information including the positions and the likelihoods of particular objects. That is, the estimation unit 204 executes the above estimation process for estimating particular objects (people or heads) and estimating the number of particular objects. The estimation unit 204 inputs each of a plurality of partial regions (first partial regions) obtained by dividing the input image to the trained model, thereby acquiring the likelihoods of particular objects included in each of the partial regions.” , where training a “regressor” or “recognition model” to detect the “presence of a particular object in an input image” in “each of the partial regions” using “training data” comprising of “small images” capturing “the size of the particular object” is generating at least one detection plan indicating one predetermined object detection model (i.e., “trained model” and/or “trained recognition model”) based on the object distribution information (e.g., “size of the particular object present in the small image”) is to be applied to at least one sub-image (i.e., “partial regions” and/or “small images of the fixed size s”) in a target image (i.e., “input image”) to be captured by the camera (e.g., “monitoring camera”)); and providing the at least one detection plan for object detection on the target image (para(s). [0030] and [0047]—see citations in preceding limitation immediately above—, where “using a regressor (a trained model) to which a small image of a certain fixed size s is input” is providing the at least one detection plan for object detection (i.e., “presence” and/or “recognition”) on the target image). Where Yamasaki does not specifically disclose obtaining performance metrics associated with a set of predetermined object detection models; generating at least one detection plan based on the object distribution information and the performance metrics, the at least one detection plan indicating which one of the predetermined object detection models in the set of predetermined object detection models is to be applied to at least one sub-image in a target image to be captured by the camera; Chaterji teaches in the same field of endeavor of providing at least one object detection plan based on object distribution information for object detection obtaining performance metrics associated with a set of predetermined object detection models (abstract and para(s). [0015-0016], [0021], and [0032], recite(s) [abstract] “…The system may approximate, based on the content features and the utilization metric, latency metrics, for a plurality of execution configuration sets, respectively. The system may also approximate, based on the content features, accuracy metrics for the execution configuration sets, respectively. The system may select the optimized execution configuration set in response to satisfaction of a performance criterion. The system may perform object detection and object tracking based on the optimized execution configuration set.” [0015] “…the system manages to keep a latency below the requirement with increased level of contention while achieving a better accuracy. To this end, the system may use a model with multiple approximation parameters that are dynamically tuned at runtime to stay on the Pareto optimal frontier (of the latency-accuracy curve in this case). We refer to the execution branch with a particular configuration set an approximation branch (AB).” [0016] “…Third, the system and methods described herein further consider how the video content influences both accuracy and latency. The system and methods described herein leverages video characteristics such as the object motion (fast vs. slow) and the sizes and the number of objects, to better predict the accuracy and latency of the ABs, and to select the best AB with reduced latency and increased accuracy. Additional benefits, efficiencies, and improvements over existing market solutions are made evident in the systems and methods described herein.” [0021] “…The object detector may perform object detection based on an object detection model. For example, the object detection model may include, for example, a deep neural network (DNN). There are various non-limiting examples of DNN's for the object detection…” [0032] “…The scheduler 104 may perform the decision-making at runtime on which AB (aka execution configuration set) should be used to run the inference on the input video frames. …” , where “latency” and “accuracy” are performance metrics and the “plurality of execution configuration sets” (i.e., “ABs”) of an “object detection model” are a set of predetermined object detection models); generating at least one detection plan based on the object distribution information and the performance metrics, the at least one detection plan indicating which one of the predetermined object detection models in the set of predetermined object detection models is to be applied… (para(s). [0016]—see citation in preceding limitation immediately above—, where selecting a “best” object detection model “configuration set” is generating at least one detection plan (e.g., a “best” object detection model “configuration set”) based on at least object distribution information (e.g., “sizes and the number of objects”) and performance metrics (e.g., “accuracy and latency”) indicating which one of the predetermined object models (e.g., the model with the “best” “configuration set”) in the set of predetermined object detection models (e.g., other object detection model “configuration set[s]”) is to be applied (i.e., “deploy[ed]”)). Since each of Yamasaki and Chaterji discloses generating at least one detection plan based on at least object distribution information indicating a size distribution of detected objects as detailed above, it would have been obvious to one of ordinary skill in the art before the effective filing date of the presently filed invention to modify the system of Yamasaki to incorporate obtaining performance metrics associated with a set of predetermined object detection models and generating at least one detection plan based on the object distribution information and the performance metrics, the at least one detection plan indicating which one of the predetermined object detection models in the set of predetermined object detection models is to be applied to at least one sub-image in a target image to be captured by the camera, to improve the at least one object detection plan by providing further performance metrics for determining a best predetermined object detection model from a set of predetermined object detection models (e.g., object detection model configuration sets) is to be applied to the at least one sub-image in a target image in resource constrained devices as taught by Chaterji above. Regarding claim 2, Yamasaki in view of Chaterji discloses the method of Claim 1, wherein Chaterji further teaches a performance metric associated with an object detection model comprises at least one of: a latency metric, indicating an estimated latency for processing a batch of images by the object detection model; or an accuracy metric, indicating an estimated accuracy for detecting objects on a particular size level by the object detection model (abstract and para(s). [0015-0016], [0021], and [0032]—see citations in claim 1 above—, where “accuracy and latency”) are performance metrics for processing a batch of images by the object detection model (e.g., video) and estimating accuracy for detecting objects on a particular size level by the object detection model (e.g., “accuracy and latency” leveraging characteristics such as “sizes and the number of objects”)). Regarding claim 3, Yamasaki in view of Chaterji discloses the method of Claim 1, wherein generating at least one detection plan comprises: generating a plurality of partition modes based on input sizes of the set of predetermined object detection models, a partition mode indicating partitioning an image view of the camera into a set of regions (Yamasaki; para(s). [0016], [0037], and [0041]—see citations in claim 1 limitations “obtaining object distribution information…” and “generating at least one detection plan…” above—, where the “plurality of partial regions” depending on “the size of the particular object appearing in the captured image” are a plurality of partition modes based on input sizes of the set of predetermined object detection models (e.g., “small image of a certain fixed size s is input”) indicating partitioning an image view of the camera into a set of regions as depicted in Fig. 3 below: PNG media_image1.png 755 1248 media_image1.png Greyscale ); generating a plurality of candidate detection plans based on the plurality of partition modes by assigning an object detection model to each region in the set of regions (Yamasaki; para(s). [0030] and [0047]—see citation in claim 1 limitation “generating at least one detection plan…” above—, where para(s). [0030] further recite(s): [0030] “…When the regressor is trained, many small images of the fixed size s in which the position of a particular object is known are prepared, and the regressor is trained in advance on these target small images as training data based on a machine learning technique. At this time, to improve the accuracy of estimating the number of particular objects, it is desirable that the ratio between the size (the fixed size s) of each small image as the training data and the size of the particular object present in the small image should be approximately constant. Then, the estimation unit 204 generates a small image by resizing an image of each of the plurality of partial regions set in the predetermined region in the input image to the fixed size s and inputs the generated small image to the regressor, thereby obtaining “the position and the likelihood (the estimated value) of the particular object in the partial region” as an output from the regressor. …” , where “resizing an image of each of the plurality of partial regions” to be input into the trained object detection model (i.e., “regressor”) is generating a plurality of candidate detection plans (e.g., a trained “regressor” to be applied to “each of the plurality of partial regions”) based on the plurality of partition modes (e.g., “partial regions”) by assigning an object detection model (e.g., a trained “regressor” taking in a “fixed size s” small image as input) to each region in the set of regions (i.e., “each of the plurality of partial regions”)); and determining the at least one detection plan based on estimated performances of the plurality of candidate detection plans (Chaterji; abstract and para(s). [0015-0016], [0021], and [0032]—see citations in claim 1 above—, where determining the “best” object detection model “configuration set” is determining the at least one detection plan (i.e., the “best” object detection model “configuration set”) based on estimated performances (e.g., “accuracy and latency”) of the plurality of candidate detection plans (e.g., the object detection model “configuration sets”)), the estimate performances being determined based on object distribution information associated with each of the set of regions and a performance metric associated with the assigned object detection model (Chaterji; abstract and para(s). [0015-0016], [0021], and [0032]—see citations in claim 1 above—, where selecting the “best” object detection model “configuration set” is the estimate performances (e.g., “latency” and “accuracy”) of the plurality of object detection plans (e.g., object detection model “configuration sets”) being determined based on object distribution information (e.g., “sizes and the number of objects”) associated with each set of regions (i.e., each region is associated with the “size” of the object as disclosed in Yamasaki above) and a performance metric (e.g., “latency” or “accuracy”) associated with the assigned object detection model (i.e., the “best” object detection model “configuration set”)). Regarding claim 4, Yamasaki in view of Chaterji discloses the method of Claim 3, wherein Chaterji further teaches determining the at least one detection plan comprises: obtaining a latency threshold for object detection on the target image (abstract and para(s). [0015-0016], [0021], and [0032]—see citations in claim 1 above—, where the “latency below a requirement” is a latency threshold for object detection); and selecting the at least one detection plan from the plurality of candidate detection plans based on a comparison of an estimated latency of a candidate detection plan and the latency threshold (abstract and para(s). [0015-0016], [0021], and [0032]—see citations in claim 1 above—, where selecting the “best” object detection model “configuration set” is selecting at least one detection plan from a plurality of candidate detection plans (e.g., “plurality of execution configuration sets”) based on a comparison of an estimated latency of a candidate detection plan (i.e., latency of the “best” object detection model “configuration set”) and the latency threshold (i.e., a “latency below a requirement”)). Regarding claim 6, Yamasaki in view of Chaterji discloses the method of Claim 3, wherein Yamasaki further teaches updating the first candidate detection plan comprises: for a first region of the plurality of regions, extending a boundary of the first region by a distance, the distance being determined based on an estimated size of a potential object to be located at the boundary of the first region (para(s). [0037] and [0041] further recite(s): [0037] “…The size of each partial region in the image captured by the imaging apparatus 110 depends on the size of the particular object appearing in the captured image. In the example of FIG. 3, a human body (a person) is taken as an example of the particular object, and the size of each partial region depends on the size of the human body appearing in a different size according to the position in the captured image. For example, in a case where the imaging apparatus 110 is placed at a position where the imaging apparatus 110 looks down on a horizontal surface such as a floor surface, i.e., in a case where the imaging apparatus 110 is installed such that the optical axis of the lens is directed below the horizontal axis, the human body appears large in a lower portion of the captured image, and the human body appears small in an upper portion of the captured image. Thus, among the plurality of partial regions 301a, 301b, and 301c in the image 300, the sizes of the partial regions 301a disposed in a lower portion of the image 300 are larger than the sizes of the partial regions 301b and 301c disposed in portions above the partial regions 301a. Similarly, the sizes of the partial regions 301b disposed in a middle portion of the image 300 are larger than the sizes of the partial regions 301c disposed in a portion above the partial regions 301b. …” [0041] “Then, based on the size information f(x,y) regarding the size of the person at any position on the image, the setting unit 203 sets the plurality of partial regions in the image (the predetermined region in the input image). …” , where setting the “plurality of partial regions in the image” is dependent on the “size of the particular object appearing in the captured image” including making a region at least “larger” than other partial regions is at least extending (i.e., making “larger”) a boundary of a first region by a distance (e.g., “size”) determined based on an estimated size of a potential object to be located at the boundary of the first region (i.e., “size of the particular object appearing in the captured image”)). Regarding claim 7, Yamasaki in view of Chaterji discloses the method of Claim 1, wherein Yamasaki further discloses the method of claim 1 further comprising: obtaining updated object distribution information associated with a new set of historical images captured by a camera (para(s). [0030], [0037], and [0041]—see citations in claims 3 and 6 above—, where para(s). [0018] further recite(s): [0018] “The imaging apparatus 110 is an apparatus that captures an image of an object. In the present exemplary embodiment, a monitoring camera is taken as an example of the imaging apparatus 110, and a particular object such as a human body or a head appearing in an image captured by the monitoring camera is a target object of estimation and measurement (hereinafter, a “measurement target object”) according to the present exemplary embodiment. The imaging apparatus 110 transmits data on an image acquired by capturing the image, information regarding the image capturing date and time when the image is captured, and identification information that is information identifying the imaging apparatus 110, in association with each other to an external apparatus such as the information processing apparatus 100 or the recording apparatus 120 via the network 140. Hereinafter, the data on the image captured by the imaging apparatus 110 will be referred to simply as an “image”, and the image capturing date and time of the image and the identification information will be referred to as “related information”. …” , where “resizing an image of each of the plurality of partial regions” including making a region “larger” is at least updating object distribution information (i.e., “sizes” of “particular object[s]” in a “captured image”) associated with a new set of historical images (i.e., “captured image[s]” other than an initial set of “captured image[s]” obtained from at least a “monitoring camera” are at least a new set of historical images)); generating at least one updated detection plan based on the updated object distribution information (para(s). [0030]—see citation in claim 3 limitation “generating a plurality of candidate detection plans…” above—, where “resizing an image of each of the plurality of partial regions” is generating at least one updated detection plan based on the updated object distribution information (e.g., a “resize[d]” partial region image based on the size of a particular object in the partial region image)); and providing the at least one updated detection plan for object detection on image to be captured by the camera (para(s). [0030]—see citation in claim 3 limitation “generating at plurality of candidate detection plans…” above—, where using the trained “regressor” after “resizing” is providing the at least one detection plan for object detection on image to be captured by the camera (e.g., an input image)). Regarding claim 8, Yamasaki discloses a computer-implemented method, comprising: partitioning a target image captured by a camera into at least one sub-image (para(s). [0016], [0037], and [0041]—see citations in claim 1 limitation “obtaining object distribution information…” above—, where partitioning an “image… from the imaging apparatus” into a “plurality of partial regions” is partitioning a target image captured by a camera (e.g., “monitoring camera”) into at least one sub-image (i.e., a “partial region”)); detecting objects in the at least one sub-image according to a target detection plan, the target detection plan indicating(para(s). [0030] and [0047]—see citation in claim 1 limitation “generating at least one detection plan…” above—, where training a “regressor” or “recognition model” to detect the “presence of a particular object in an input image” in “each of the partial regions” using “training data” comprising of “small images” capturing “the size of the particular object” is detecting (i.e., “presence” and/or “recognition”) objects in the at least one sub-image (e.g., a “partial region”) according to a target detection plan indicating one predetermined object detection model (i.e., “trained model” and/or “trained recognition model”) is to be applied to at least one sub-image (i.e., “partial regions” and/or “small images of the fixed size s”)); and determining objects in the target image based on the detected objects in the at least one sub- image (para(s). [0030] and [0047]—see citations in claim 1 limitation “generating at least one detection plan…” above—, where “using a regressor (a trained model) to which a small image of a certain fixed size s is input” to recognize and/or detect the “presence of a particular object in an input image” in “each of the partial regions” is determining objects in the target image based on the detected objects in the at least one sub-image (e.g., a “partial region”)). Where Yamasaki does not specifically disclose …the target detection plan indicating which one of the predetermined object detection models in the set of predetermined object detection models is to be applied to at least one sub-image; Chaterji teaches in the same field of endeavor of providing at least one object detection plan based on object distribution information and performance metrics for object detection …the target detection plan indicating which one of the predetermined object detection models in the set of predetermined object detection models is to be applied to at least one sub-image (abstract and para(s). [0015-0016], [0021], and [0032]—see citations in similar limitation in claim 1 above—, where selecting a “best” object detection model “configuration set” is a target detection plan indicating which one of the predetermined object models (e.g., the model with the “best” “configuration set”) is to be applied (i.e., “deploy[ed]”)). Claim 8 recites similar limitations to claim 1 and is rejected for similar rationale and reasoning (see the rejection of claim 1 above). Regarding claim 9, Yamasaki in view of Chaterji discloses the method of Claim 8, wherein Chaterji further teaches the target detection plan is generated based on object distribution information associated with a set of historical images captured by the camera and performance metrics associated with the set of predetermined object detection models (abstract and para(s). [0015-0016], [0021], and [0032]—see similar limitation in claim 1 above— where selecting a “best” object detection model “configuration set” is generating the target detection plan (e.g., “best” object detection model “configuration set”) based on at least object distribution information (e.g., “sizes and the number of objects”) associated with a set of historical images (e.g., captured “video”) and performance metrics (e.g., “accuracy and latency”) associated with the set of predetermined object models (e.g., other object detection model “configuration set[s]”)). Regarding claim 10, Yamasaki in view of Chaterji discloses the method of Claim 8, wherein Chaterji further teaches the method of Claim 8 further comprising: obtaining a plurality of detection plans (abstract and para(s). [0015-0016], [0021], and [0032]—see similar limitation in claim 1 above— where the “plurality of execution configuration sets” of an “object detection model” are a plurality of detection plans); and selecting the target plan from the plurality of detection plans based on a latency threshold for object detection on the target image (abstract and para(s). [0015-0016], [0021], and [0032]—see citations in similar claim 4 limitation “selecting the at least one detection plan…” above—, where selecting a “best” “configuration set” of an “object detection model” comprises of determining a “latency below a requirement” or the “best” configuration set (i.e., “best AB”) with at least “reduced latency” is selecting the target plan from the plurality of detection plans (i.e., “configuration sets”) based on a latency threshold (e.g., “latency below a requirement” or best configuration set with “reduced latency”) for object detection (e.g., an “object detection model”)). Regarding claim 11, Yamasaki in view of Chaterji discloses the method of Claim 10, wherein Chaterji further teaches selecting the target plan from the plurality of detection plans based on a latency threshold for object detection on the target image comprises: obtaining a historical latency for object detection on a historical image captured by the camera, the object detection on the historical image being performed according to a first detection plan of the plurality of detection plans (para(s). [0068] and [0071], recite(s) [0068] “The scheduler may forecast latency metrics for execution configuration sets ( 306 ). The latency metric may measure the end-to-end latency of the object detection for detecting the objects in a video frame and averaged across all the frames of the video, which essentially maps to the entire length of the video. Typically, this will be in milliseconds for latency-sensitive applications, and more specifically in the realm of 33 msec to 50 msec to support 20-30 frames/sec. Alternatively or in addition, the latency metric may be expressed as a percentile, such as a p50, p75, p99 etc.” [0071] “The scheduler may select an execution configuration from the domain of execution configuration 310 . The accuracy and latency metrics associated with the selected execution configuration may satisfy a performance criterion provided to the scheduler. For example, the scholar may receive, via user input or some other source, the performance criterion. The performance criterion may have a rule that compares the accuracy and/or latency metrics to predefined threshold values or evaluates the metrics under predefined logic to provide an indication of acceptance, such as a Boolean value or the like. If the criterion is satisfied, then the execution configuration is selected for the multi-branch object detector.” , where the “latency metric” determined for an object detection model includes “detecting the objects in a video frame” is a historical latency for object detection on a historical image (e.g., a “video frame”), wherein the object detection on the historical image is performed according to a first detection plan (e.g., an object detection model “configuration set”, such as a “best” object detection model “configuration set”) of the plurality of detection plans (e.g., the object detection model “configuration set[s]”)); and selecting the target detection plan from the plurality of detection plans based on a difference between the historical latency and the latency threshold (para(s). [0068] and [0071]—see citations in the preceding limitation immediately above—, where selecting the “best” object detection model “configuration set” by “compar[ing] the accuracy and/or latency metrics to predefined threshold values” is selecting a detection plan from amongst a plurality of detection plans (e.g., other “configuration[s]”) based on at least a difference between the historical latency (i.e., “latency metric of the object detection for detecting the objects in a video frame…” of an object detection model) and the latency threshold (i.e., latency metric “predefined threshold values”)). Regarding claim 15, the claim recites similar limitations to claim 1 but in the form of an electronic device. Therefore, claim 15 is rejected for similar rationale and reasoning as claim 1 (see the analysis for claim 1 above). Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Yamasaki in view of Chaterji as applied to claim 3 above, and further in view of Bhatia et al. (Bhatia; US 2021/0090257 A1). Regarding claim 5, Yamasaki in view of Chaterji discloses the method of Claim 3, wherein Yamasaki further discloses the method of Claim 3 further comprising: updating the at least one detection plan, comprising: for a first detection plan of the at least one detection plan indicating that the image view is to be partitioned into a plurality of regions, updating the first detection plan by adjusting sizes of the plurality of regions(para(s). [0030]—see citation in claim 3 limitation “generating a plurality of candidate detection plans…” above—, where “resizing an image of each of the plurality of partial regions” is updating a first detection plan by adjusting sizes (i.e., “resizing”) of the plurality of regions (i.e., “plurality of partial regions”)). Where Yamasaki in view of Chaterji does not specifically disclose …adjusting sizes of the plurality of regions such that each pair of neighboring regions among the plurality of regions are partially overlapped; Bhatia teaches in the same field of endeavor of image windows in object detection …adjusting sizes of the plurality of regions such that each pair of neighboring regions among the plurality of regions are partially overlapped (para(s). [0083], recite(s) [0083] “In the sliding window approach, a window “slides” over an image (e.g., the representation image and/or the first/second slices and/or the sequential images) for dividing it into patches. Every slide, the window is moved a specific number of pixels to the side, which is also called “the stride”. The stride may be such that subsequent image patches may comprise some overlap with the previous image patches. Optionally, the size of the window is adjusted based on the size of the region of interest, e.g., such that the window has approximately the same size as the region of interest.” , where “the size of the window is adjusted based on the size of the region of interest” such that subsequent image windows (i.e., “patches”) “comprise some overlap with previous image patches”) is adjusting sizes of a plurality of regions (e.g., “patches” or “size of the window”) such that each pair of neighboring regions among the plurality of regions are partially overlapped). It would have been obvious to one of ordinary skill in the art before the effective filing date of the presently filed invention to modify the system of Yamasaki in view of Chaterji to incorporate adjusting the sizes of the plurality of regions when updating the detection plan such that that each pair of neighboring regions among the plurality of regions are partially overlapped to improve capturing individual objects in an overall image for object detection processing as taught by Bhatia (para(s). [0084], recite(s) [0084] “…The use of an overlapping sliding window method enables to more readily capture individual structures in the medical image volume. This is because it is usually unlikely that patterns in an image volume conform to a fixed grid of image patches. This might have the consequences that patterns may be cut by the image patches adversely affecting the calculation result (i.e., the degree of similarity). By using an overlapping sliding window approach, a form of a mean calculation is applied across a plurality of image patches partially mitigating such finite size effects. Moreover, the resolution of the degrees of similarity with respect to the medical image volume is enhanced.” , where “more readily captur[ing] individual structures in the medical image volume” is improving capturing of individual objects in an overall image). Claims 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Yamasaki in view of Chaterji as applied to claim 8 above, and further in view of Habibian et al. (Habibian; US 2022/0159278 A1). Regarding claim 12, Yamasaki in view of Chaterji discloses the method of Claim 8, wherein Yamasaki further teaches the at least one image comprises a plurality of sub-images (para(s). [0016], [0037], and [0041]—see citations in claim 1 limitation “obtaining object distribution information…” above—, where the “plurality of partial regions” are a plurality of sub-images), and wherein detecting objects in the at least one sub-image according to a target detection plan comprises: determining a first set of sub-images from the plurality of sub-images based on object detection results of a plurality of historical images obtained according to the target detection plan (para(s). [0016], [0037], and [0041]—see citations in claim 1 limitation “obtaining object distribution information…” above—and [0030] and [0047]—see citations in claim 1 limitation “generating at least one detection plan…” above—, where the “plurality of partial regions” of a first “size” are a first set of sub-images based on object detection results of a plurality of historical images (e.g., “captured image[s]”) obtained according to the target detection plan (e.g., training a “regressor or “recognition model” using “training data” comprising of “small images” of a particular “fixed size s”)); and detecting objects in the plurality of sub-images(para(s). [0030] and [0047]—see citations in claim 1 limitation “generating at least one detection plan…” above—, where performing “recognition” and/or detecting “presence” of objects (e.g., “particular object[s]” and/or people) in “each of the partial regions” is detecting objects in the plurality of sub-images). Where Yamasaki in view of Chaterji does not specifically disclose and detecting objects in the plurality of sub-images including skipping object detection on the first set of sub-images; Habibian teaches in the same field of endeavor of sub-images in an image for object detection and detecting objects in the plurality of sub-images including skipping object detection on the first set of sub-images (para(s). [0060] and [0064], recite(s) [0060] “As described, artificial neural networks (ANN) are useful for processing videos. However, videos may have large amounts of redundancy across video frames. As such, ANN-based video processing systems may perform the same convolution operation multiple times. Such processing is time consuming and results in significant energy consumption. Accordingly, aspects of the present disclosure are directed to avoiding processing or skipping performance of convolution operations of redundant portions of consecutive video frames. The skip convolution may be applied to video processing systems including convolutional neural networks. The neural network may learn to skip the processing of residual frames or a portion of frames within the video stream.” [0064] “…To improve processing (e.g., object detection or human pose estimation), a residual or difference between the frames 502 , 504 may be computed. The residual frame may put a strong prior on the informative regions to be processed versus non-informative regions that may be skipped. As shown in FIG. 5, a residual frame 506 may be determined by computing the difference between the input features for frame 502 xt and frame 504 xt-1 . A video processing system, including a CNN (e.g., deep convolutional network 350 ), may be configured to skip processing of residual portions or portions of the frame 502 that are determined to be the same as the previous frame 504 . That is, based on the residual rt , convolution operations for one or more regions or portions of frame 502 may be avoided. …” , where the “one or more regions or portions” of a current frame “may be avoided” when the “one or more regions or portions” of the current frame is “determined to be the same as the previous frame” is a first set of sub-images skipped in object detection (i.e., “skip the processing of… a portion of frames within the video stream”)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the presently filed invention to modify the system of Yamasaki in view of Chaterji to incorporate skipping object detection on the first set of sub-images in the detection of objects in the plurality of sub-images to improve processing in the predetermined object detection models that are part of the at least one detection plan on detecting objects in video streams by reducing processing of redundant consecutive image frames in video streams as taught by Habibian above. Regarding claim 13, Yamasaki, as modified by Chaterji and Habibian, discloses the method of Claim 12, wherein Habibian further teaches a first set of sub-images comprise at least one of: a first sub-image corresponding to a first region, wherein no object is detected from sub-images corresponding to the first region of the plurality of historical images, or a second sub-image corresponding to a second region, wherein object detection on a sub-image corresponding to the second region of a previous historical image is skipped (para(s). [0060] and [0064]—see citations in claim 12 above—, where the avoided “one or more regions or portions” of a current frame is at least a second sub-image skipped in object detection; wherein the second sub-image corresponds to a second region of a previous historical image (i.e., the “one or more regions or portions” of “the previous frame” similar to the “one or more regions or portions” of the current frame)). Regarding claim 14, Yamasaki, as modified by Chaterji and Habibian, discloses the method of Claim 13, wherein Habibian further teaches object detection on sub-images corresponding to the second region is skipped for a plurality of consecutive historical images (para(s). [0060] and [0064]—see citations in claim 12 above—, where “skipping performance of convolution operations of redundant portions of consecutive video frames” is skipping object detection on sub-images corresponding to the second region for a plurality of consecutive historical images (i.e., “video frames”)), and wherein a number of the plurality of consecutive historical images is less than a threshold number (para(s). [0060] and [0064]—see citations in claim 12 above—, where the number of the plurality of consecutive historical images is less than the number of total frames within the video stream). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JULIA Z YAO whose telephone number is (571)272-2870. The examiner can normally be reached Monday - Friday (8:30AM - 5PM). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571)270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.Z.Y./Examiner, Art Unit 2666 /MING Y HON/Primary Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Nov 20, 2023
Application Filed
Apr 22, 2026
Non-Final Rejection mailed — §103
May 05, 2026
Examiner Interview Summary
May 05, 2026
Applicant Interview (Telephonic)
May 28, 2026
Response Filed
Aug 20, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738027
IMAGE PROCESSING APPARATUS AND IMAGE PROCESSING METHOD FOR DETEERMINING THE PRESENCE OR ABSENCE OF ABNORMALITIES IN PARTIAL IMAGES.
4y 6m to grant Granted Sep 15, 2026
Patent 12694384
REUSABLE BAG RECOGNITION
2y 2m to grant Granted Jul 28, 2026
Patent 12682651
APPARATUS AND METHOD FOR MODIFYING GROUND TRUTH FOR CHECKING ACCURACY OF MACHINE LEARNING MODEL
4y 4m to grant Granted Jul 14, 2026
Patent 12656334
METHOD AND DEVICE FOR DETERMINING RED BLOOD CELLS DEFORMABILITY
4y 1m to grant Granted Jun 16, 2026
Patent 12657677
METHOD FOR INSPECTING THE SIDE WALL OF AN OBJECT
3y 9m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
63%
Grant Probability
99%
With Interview (+48.3%)
3y 3m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 84 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month