Prosecution Insights
Last updated: August 30, 2026
Application No. 18/952,789

AI-ENABLED MACHINE VISION TASK AWARE ADAPTIVE INFERENCE

Non-Final OA §103
Filed
Nov 19, 2024
Examiner
WILLIAMS, REBECCA COLETTE
Art Unit
2677
Tech Center
2600 — Communications
Assignee
InterDigital Inc.
OA Round
1 (Non-Final)
36%
Grant Probability
At Risk
1-2
OA Rounds
1y 5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants only 36% of cases
36%
Career Allowance Rate
4 granted / 11 resolved
-25.6% vs TC avg
Strong +70% interview lift
Without
With
+70.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
20 currently pending
Career history
37
Total Applications
across all art units

Statute-Specific Performance

§101
10.5%
-29.5% vs TC avg
§103
59.4%
+19.4% vs TC avg
§102
15.0%
-25.0% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 11 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The Information Disclosure Statement filed 11/19/2024 has been considered by examiner. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7, 9-17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zheng (US 10366053 B1) in view of Lee (WO 2023239043 A1). With respect to claim 1, Zheng teaches a wireless transmit/receive unit (WTRU) (see figure 36) comprising: a processor (see figure 36) configured to: transmit a request to a plurality of AI-enabled machine vision (AI-MV) servers to perform inference on an AI-MV task (“In at least some embodiments, a server that implements one or more of the components of a machine learning service (including control-plane components such as API request handlers, input record handlers, recipe validators and recipe run-time managers, plan generators, job schedulers, artifact repositories, and the like, as well as data plane components such as MLS servers) may include a general-purpose computer system that includes or is configured to access one or more computer-accessible media.” Page 70 col 59 lines 12-20 And “The MLS programmatic interfaces may enable users to submit respective requests for several related tasks of a given machine learning workflow, such as tasks for extracting records from data sources, generating statistics on the records, feature processing, model training, prediction, and so on.” Page 43 col 5 lines 54-59 and figures 1 and 4); receive configuration information for a plurality of inference modes (“In some embodiments, a pool of compute servers and/or storage servers may be pre-configured for the MLS, and the resources for a given job may be selected from such a pool. In other embodiments, the resources may be selected from a pool assigned to the client on whose behalf the job is to be executed—e.g., the client may acquire resources from a computing service of the provider network prior to submitting API requests, and may provide an indication of the acquired resources to the MLS for job scheduling. If client-provided code (e.g., code that has not necessarily been thoroughly tested by the MLS, and/or is not included in the MLS's libraries) is being used for a given job, in some embodiments the client may be required to acquire the resources to be used for the job, so that any side effects of running the client-provided code may be restricted to the client's own resources instead of potentially affecting other clients.” Page 44 col 8 lines 29-44), wherein the configuration information comprises network resource availability for executing the AI-MV task for each of the plurality of inference modes (“The term “MLS control plane” may be used herein to refer to a collection of hardware and/or software entities that are responsible for implementing various types of machine learning functionality on behalf of clients of the MLS, and for administrative tasks not necessarily visible to external MLS clients, such as ensuring that an adequate set of resources is provisioned to meet client demands” page 43 col 5 line 19-25) and processing capabilities of the plurality of AI-MV servers (“The MLS may be responsible for ensuring that the dependencies of a given job have been met before the corresponding operations are initiated. The MLS may also be responsible in such embodiments for generating a processing plan for each job, identifying the appropriate set of resources (e.g., CPUs/cores, storage or memory) for the plan, scheduling the execution of the plan, gathering results, providing/saving the results in an appropriate destination, and at least in some cases for providing status updates or responses to the requesting clients.” Page 43 col 6 lines 25-34); select an inference mode from the plurality of inference modes based on the configuration information (“A number of choices may be available with respect to the manner in which the operations corresponding to a given job are mapped to MLS servers. For example, it may be possible to partition the work required for a given job among many different servers to achieve better performance. As part of developing the processing plan for a job, the MLS may select a workload distribution strategy for the job in some embodiments. The parameters determined for workload distribution in various embodiments may differ based on the nature of the job. Such factors may include, for example, (a) determining a number of passes of processing, (b) determining a parallelization level (e.g., the number of “mappers” and “reducers” in the case of a job that is to be implemented using the Map-Reduce technique), (c) determining a convergence criterion to be used to terminate the job, (d) determining a target durability level for intermediate data produced during the job, or (e) determining a resource capacity limit for the job (e.g., a maximum number of servers that can be assigned to the job based on the number of servers available in MLS server pools, or on the client's budget limit)” page 44 col 8 lines 5-25), network conditions detected by the WTRU (“A number of choices may be available with respect to the manner in which the operations corresponding to a given job are mapped to MLS servers. For example, it may be possible to partition the work required for a given job among many different servers to achieve better performance. As part of developing the processing plan for a job, the MLS may select a workload distribution strategy for the job in some embodiments. The parameters determined for workload distribution in various embodiments may differ based on the nature of the job. Such factors may include, for example, (a) determining a number of passes of processing, (b) determining a parallelization level (e.g., the number of “mappers” and “reducers” in the case of a job that is to be implemented using the Map-Reduce technique), (c) determining a convergence criterion to be used to terminate the job, (d) determining a target durability level for intermediate data produced during the job, or (e) determining a resource capacity limit for the job (e.g., a maximum number of servers that can be assigned to the job based on the number of servers available in MLS server pools, or on the client's budget limit)” page 44 col 8 lines 5-25), and resource availability at the WTRU (“A number of choices may be available with respect to the manner in which the operations corresponding to a given job are mapped to MLS servers. For example, it may be possible to partition the work required for a given job among many different servers to achieve better performance. As part of developing the processing plan for a job, the MLS may select a workload distribution strategy for the job in some embodiments. The parameters determined for workload distribution in various embodiments may differ based on the nature of the job. Such factors may include, for example, (a) determining a number of passes of processing, (b) determining a parallelization level (e.g., the number of “mappers” and “reducers” in the case of a job that is to be implemented using the Map-Reduce technique), (c) determining a convergence criterion to be used to terminate the job, (d) determining a target durability level for intermediate data produced during the job, or (e) determining a resource capacity limit for the job (e.g., a maximum number of servers that can be assigned to the job based on the number of servers available in MLS server pools, or on the client's budget limit)” page 44 col 8 lines 5-25), network conditions detected by the WTRU (“A number of choices may be available with respect to the manner in which the operations corresponding to a given job are mapped to MLS servers. For example, it may be possible to partition the work required for a given job among many different servers to achieve better performance. As part of developing the processing plan for a job, the MLS may select a workload distribution strategy for the job in some embodiments. The parameters determined for workload distribution in various embodiments may differ based on the nature of the job. Such factors may include, for example, (a) determining a number of passes of processing, (b) determining a parallelization level (e.g., the number of “mappers” and “reducers” in the case of a job that is to be implemented using the Map-Reduce technique), (c) determining a convergence criterion to be used to terminate the job, (d) determining a target durability level for intermediate data produced during the job, or (e) determining a resource capacity limit for the job (e.g., a maximum number of servers that can be assigned to the job based on the number of servers available in MLS server pools, or on the client's budget limit)” page 44 col 8 lines 5-25); transmit an indication of the selected inference mode to the plurality of AI-MV servers (see figure 1); transmit data related to the AI-MV task to the plurality of AI-MV servers for processing in accordance with the selected inference mode (see figure 1); receive an inference result from at least one AI-MV server of the plurality of AI-MV servers generated in accordance with the selected inference mode (“Results of some jobs may be stored as MLS artifacts within repository 120 in some embodiments, as indicated by arrow 147” page 45 col 9 lines 33-35); and perform a follow-on action based on the inference result (“In some embodiments, partial dependencies among tasks may be supported—e.g., in a sequence of tasks (T1, T2, T3), T2 may depend on partial completion of T1, and T2 may therefore be scheduled before T1 completes. For example, T1 may comprise two phases or passes P1 and P2 of statistics calculations, and T2 may be able to proceed as soon as phase P1 is completed, without waiting for phase P2 to complete. Partial results of T1 (e.g., at least some statistics computed during phase P1) may be provided to the requesting client as soon as they become available in some cases, instead of waiting for the entire task to be completed.” Page 43 col 6 lines 34-47). Zheng does not explicitly teach an AI-MV task associated with a machine vision application. Lee teaches an AI-MV task associated with a machine vision application (“In one embodiment, the processor 350 may transmit at least part of the information related to the object to an external electronic device through the communication module 310. For example, when an electronic device operates as an IoT device or hub device on an IoT network, the processor 350 provides at least one image and the type of object detected within the at least one image to the IoT service provider. Transfers may be made to electronic devices (e.g., electronic devices of users whose accounts are registered with the server).” Page 18 paragraph 5). Lee is analogous art in the same field of endeavor as the claimed invention. Lee is directed to a transmitting an AI-MV task to a larger system (“In one embodiment, the processor 350 may transmit at least part of the information related to the object to an external electronic device through the communication module 310. For example, when an electronic device operates as an IoT device or hub device on an IoT network, the processor 350 provides at least one image and the type of object detected within the at least one image to the IoT service provider. Transfers may be made to electronic devices (e.g., electronic devices of users whose accounts are registered with the server).” Page 18 paragraph 5). A person of ordinary skill in the art, before the effective filing date of the claimed invention would have found it obvious that incorporating the AI-MV task of Lee into the job based system of Zheng would be effective, as seen in the very similar process depicted in Lee (“In one embodiment, the processor 350 may transmit at least part of the information related to the object to an external electronic device through the communication module 310. For example, when an electronic device operates as an IoT device or hub device on an IoT network, the processor 350 provides at least one image and the type of object detected within the at least one image to the IoT service provider. Transfers may be made to electronic devices (e.g., electronic devices of users whose accounts are registered with the server).” Page 18 paragraph 5); with the expectation that doing so would lead to improved performance ( see Zheng “A number of choices may be available with respect to the manner in which the operations corresponding to a given job are mapped to MLS servers. For example, it may be possible to partition the work required for a given job among many different servers to achieve better performance…” page 44 col 8 lines 5-25), especially given the multi-step AI-MV task as presented in Lee (“A method for detecting an object and an electronic device supporting the same according to an embodiment include, using an artificial intelligence model, from a frame including at least one area corresponding to at least one area of interest obtained from an image, By performing an operation to obtain related information, an operation to detect an object can be performed more quickly and more accurately.” Page 4 paragraph 2). With respect to claim 2, Zheng and Lee teach the WTRU of claim 1. Lee further teaches wherein, to perform the follow-on action, the processor is configured to detect an object (“A method for detecting an object and an electronic device supporting the same according to an embodiment include, using an artificial intelligence model, from a frame including at least one area corresponding to at least one area of interest obtained from an image, By performing an operation to obtain related information, an operation to detect an object can be performed more quickly and more accurately.” Page 4 paragraph 2) or track an object (“In one embodiment, when an object is detected in at least one image as a result of the inference operation,the processor 350 tracks the region of interest in the acquired at least one image without performing some of the operations for detecting the object. (tracking) operations can be performed.” Page 18 paragraph 3) With respect to claim 3, Zheng and Lee teach the WTRU of claim 1. Lee further teaches wherein the inference result comprises information indicating detections (“In one embodiment, when an object is detected in at least one image as a result of the inference operation,the processor 350 tracks the region of interest in the acquired at least one image without performing some of the operations for detecting the object. (tracking) operations can be performed.” Page 18 paragraph 3, ROI as environment) or identifications of objects detected in an environment surrounding the WTRU (“In one embodiment, when an object is detected in at least one image as a result of the inference operation,the processor 350 tracks the region of interest in the acquired at least one image without performing some of the operations for detecting the object. (tracking) operations can be performed.” Page 18 paragraph 3, ROI as environment). With respect to claim 4, Zheng and Lee teach the WTRU of claim 1. Zheng further teaches wherein the network conditions detected by the WTRU comprise: bandwidth availability detected at the WTRU during communication with the plurality of AI-MV servers (A number of choices may be available with respect to the manner in which the operations corresponding to a given job are mapped to MLS servers. For example, it may be possible to partition the work required for a given job among many different servers to achieve better performance. As part of developing the processing plan for a job, the MLS may select a workload distribution strategy for the job in some embodiments. The parameters determined for workload distribution in various embodiments may differ based on the nature of the job. Such factors may include, for example … determining a resource capacity limit for the job (e.g., a maximum number of servers that can be assigned to the job based on the number of servers available in MLS server pools, or on the client's budget limit)” page 44 col 8 lines 5-25). With respect to claim 5, Zheng and Lee teach the WTRU of claim 1. Zheng further teaches wherein the data related to the AI-MV task comprises metadata associated with the input data (“A client request 111 may indicate one or more parameters that may be used by the MLS to perform the operations, such as a data source definition 150, a feature processing transformation recipe 152, or parameters 154 to be used for a particular machine learning algorithm. In some embodiments, artifacts respectively representing the parameters may also be stored in repository 120. Some machine learning workflows, which may correspond to a sequence of API requests from a client 164, may include the extraction and cleansing of input data records from raw data repositories 130 (e.g., repositories indicated in data source definitions 150) by input record handlers 160 of the MLS, as indicated by arrow 114. This first portion of the workflow may be initiated in response to a particular API invocation from a client 164, and may be executed using a first set of resources from pool 185.” Page 45 col 9 lines 49-64), or task-specific parameters identifying the AI-MV task and associated sub-tasks (“A client request 111 may indicate one or more parameters that may be used by the MLS to perform the operations, such as a data source definition 150, a feature processing transformation recipe 152, or parameters 154 to be used for a particular machine learning algorithm. In some embodiments, artifacts respectively representing the parameters may also be stored in repository 120. Some machine learning workflows, which may correspond to a sequence of API requests from a client 164, may include the extraction and cleansing of input data records from raw data repositories 130 (e.g., repositories indicated in data source definitions 150) by input record handlers 160 of the MLS, as indicated by arrow 114. This first portion of the workflow may be initiated in response to a particular API invocation from a client 164, and may be executed using a first set of resources from pool 185.” Page 45 col 9 lines 49-64). With respect to claim 6, Zheng and Lee teach the WTRU of claim 1. Zheng further teaches wherein the processor is configured to determine whether to execute the AI-MV task locally at the WTRU (“In the depicted embodiment, a client 164 of the MLS may submit a model execution request 812 to the MLS control plane 180 via a programmatic interface 861. The model execution request may specify the execution mode (batch, online or local), the input data to be used for the model run (which may be produced using a specified data source or recipe in some cases), the type of output (e.g., a prediction or an evaluation) that is desired, and/or optional parameters (such as desired model quality targets, minimum input record group sizes to be used for online predictions, and so on). In response the MLS may generate a plan for model execution and select the appropriate resources to implement the plan.” Page 50 col 20 lines 30-33), remotely at one or more AI-MV servers of the plurality of AI-MV servers (“In the depicted embodiment, a client 164 of the MLS may submit a model execution request 812 to the MLS control plane 180 via a programmatic interface 861. The model execution request may specify the execution mode (batch, online or local), the input data to be used for the model run (which may be produced using a specified data source or recipe in some cases), the type of output (e.g., a prediction or an evaluation) that is desired, and/or optional parameters (such as desired model quality targets, minimum input record group sizes to be used for online predictions, and so on). In response the MLS may generate a plan for model execution and select the appropriate resources to implement the plan.” Page 50 col 20 lines 30-33), or in a split manner across the WTRU and the one or more AI-MV servers of the plurality of AI-MV servers to select the inference mode from the plurality of inference modes, wherein the plurality of inference modes comprises a local inference mode, a remote inference mode, and a split inference mode ( “In at least some embodiments, a job object may be generated upon receiving the execution request 812 as described earlier, indicating any dependencies on other jobs (such as the execution of a recipe for feature processing), and the job may be placed in a queue. For batch mode 865, for example, one or more servers may be identified to run the model. “ page 50 col 20 lines 32-28 and “In the depicted embodiment, a client 164 of the MLS may submit a model execution request 812 to the MLS control plane 180 via a programmatic interface 861. The model execution request may specify the execution mode (batch, online or local), the input data to be used for the model run (which may be produced using a specified data source or recipe in some cases), the type of output (e.g., a prediction or an evaluation) that is desired, and/or optional parameters (such as desired model quality targets, minimum input record group sizes to be used for online predictions, and so on). In response the MLS may generate a plan for model execution and select the appropriate resources to implement the plan.” Page 50 col 20 lines 30-33). With respect to claim 7, Zheng and Lee teach the WTRU of claim 1. Zheng further teaches wherein the configuration information comprises information regarding available codecs at the WTRU or within the network (“In at least some embodiments, the input data reaching the MLS may be encrypted or compressed, and the MLS input data handling machinery may have to perform decryption or decompression before the input data records can be used for machine learning tasks. In some embodiments in which encryption is used, MLS clients may have to provide decryption metadata (e.g., keys, passwords, or other credentials) to the MLS to allow the MLS to decrypt data records. Similarly, an indication of the compression technique used may be provided by the clients in some implementations to enable the MLS to decompress the input data records appropriately. The output produced by the input record handlers may be fed to feature processors 162 (as indicated by arrow 115), where a set of transformation operations may be performed in accordance with recipes 152 using another set of resources from pool 185.” Page 45 col 10 lines 9-24); and wherein the processor is configured to select the inference mode based on codec performance criteria (see figure 1 and “In at least some embodiments, the input data reaching the MLS may be encrypted or compressed, and the MLS input data handling machinery may have to perform decryption or decompression before the input data records can be used for machine learning tasks. In some embodiments in which encryption is used, MLS clients may have to provide decryption metadata (e.g., keys, passwords, or other credentials) to the MLS to allow the MLS to decrypt data records. Similarly, an indication of the compression technique used may be provided by the clients in some implementations to enable the MLS to decompress the input data records appropriately. The output produced by the input record handlers may be fed to feature processors 162 (as indicated by arrow 115), where a set of transformation operations may be performed in accordance with recipes 152 using another set of resources from pool 185.” Page 45 col 10 lines 9-24 and “A number of choices may be available with respect to the manner in which the operations corresponding to a given job are mapped to MLS servers. For example, it may be possible to partition the work required for a given job among many different servers to achieve better performance. As part of developing the processing plan for a job, the MLS may select a workload distribution strategy for the job in some embodiments. The parameters determined for workload distribution in various embodiments may differ based on the nature of the job. Such factors may include, for example…determining a target durability level for intermediate data produced during the job, or (e) determining a resource capacity limit for the job (e.g., a maximum number of servers that can be assigned to the job based on the number of servers available in MLS server pools, or on the client's budget limit)” page 44 col 8 lines 5-25 ). With respect to claim 9, Zheng and Lee teach the WTRU of claim 1. Zheng further teaches wherein the processor is configured to communicate with one or more AI-MV servers of the plurality of AI-MV servers to offload portions of the AI-MV task for processing in a split inference mode (“In at least some embodiments, a job object may be generated upon receiving the execution request 812 as described earlier, indicating any dependencies on other jobs (such as the execution of a recipe for feature processing), and the job may be placed in a queue. For batch mode 865, for example, one or more servers may be identified to run the model. “ page 50 col 20 lines 32-28). With respect to claim 10, Zheng and Lee teach the WTRU of claim 9. Zheng further teaches wherein the processor is configured to transmit a first portion of the AI-MV task to a first AI-MV server of the plurality of AI-MV servers (“In at least some embodiments, a job object may be generated upon receiving the execution request 812 as described earlier, indicating any dependencies on other jobs (such as the execution of a recipe for feature processing), and the job may be placed in a queue. For batch mode 865, for example, one or more servers may be identified to run the model. “ page 50 col 20 lines 32-28) and receive an inference result that corresponds to the first portion of the AI-MV task processed by the first AI-MV server of the plurality of AI-MV servers (“In some embodiments, partial dependencies among tasks may be supported—e.g., in a sequence of tasks (T1, T2, T3), T2 may depend on partial completion of T1, and T2 may therefore be scheduled before T1 completes. For example, T1 may comprise two phases or passes P1 and P2 of statistics calculations, and T2 may be able to proceed as soon as phase P1 is completed, without waiting for phase P2 to complete. Partial results of T1 (e.g., at least some statistics computed during phase P1) may be provided to the requesting client as soon as they become available in some cases, instead of waiting for the entire task to be completed.” Page 43 col 6 lines 34-47). With respect to claim 11, Zheng and Lee teach all claim limitations in consideration of claim 1, due to the substantial similarities between claims 1 and 11, with claim 11 being directed towards the method that the claim 1 processor is configured to perform. With respect to claim 12, Zheng and Lee teach the method of claim 11, and all additional limitations in consideration of claim 2, due to the substantial similarities between claims 2 and 12, with claim 12 being directed towards the method that the claim 1 processor is configured to perform in claim 2. With respect to claim 13, Zheng and Lee teach the method of claim 11, and all additional limitations in consideration of claim 3, due to the substantial similarities between claims 3 and 13, with claim 13 being directed towards the method that the claim 1 processor is configured to perform in claim 3. With respect to claim 14, Zheng and Lee teach the method of claim 11, and all additional limitations in consideration of claim 4, due to the substantial similarities between claims 4 and 14, with claim 14 being directed towards the method that the claim 1 processor is configured to perform in claim 4. With respect to claim 15, Zheng and Lee teach the method of claim 11, and all additional limitations in consideration of claim 5, due to the substantial similarities between claims 5 and 15, with claim 15 being directed towards the method that the claim 1 processor is configured to perform in claim 5. With respect to claim 16, Zheng and Lee teach the method of claim 11, and all additional limitations in consideration of claim 6, due to the substantial similarities between claims 6 and 16, with claim 16 being directed towards the method that the claim 1 processor is configured to perform in claim 6. With respect to claim 17, Zheng and Lee teach the method of claim 11, and all additional limitations in consideration of claim 7, due to the substantial similarities between claims 7 and 17, with claim 17 being directed towards the method that the claim 1 processor is configured to perform in claim 7. With respect to claim 20, Zheng and Lee teach the method of claim 11. Zheng further teaches it comprising transmitting a first portion of the AI-MV task to a first AI-MV server of the plurality of AI-MV server (“In at least some embodiments, a job object may be generated upon receiving the execution request 812 as described earlier, indicating any dependencies on other jobs (such as the execution of a recipe for feature processing), and the job may be placed in a queue. For batch mode 865, for example, one or more servers may be identified to run the model. “ page 50 col 20 lines 32-28) and receiving an inference result that corresponds to the first portion of the AI-MV task processed by the first AI-MV server of the plurality of AI-MV servers (“In some embodiments, partial dependencies among tasks may be supported—e.g., in a sequence of tasks (T1, T2, T3), T2 may depend on partial completion of T1, and T2 may therefore be scheduled before T1 completes. For example, T1 may comprise two phases or passes P1 and P2 of statistics calculations, and T2 may be able to proceed as soon as phase P1 is completed, without waiting for phase P2 to complete. Partial results of T1 (e.g., at least some statistics computed during phase P1) may be provided to the requesting client as soon as they become available in some cases, instead of waiting for the entire task to be completed.” Page 43 col 6 lines 34-47). Claims 8, 18, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Zheng and Lee as applied to claims 1 and 11 and further in view of Kesavan (WO 2024102145 A1). With respect to claim 8, Zheng and Lee teach the WTRU of claim 1. Zheng further teaches select the inference mode based on a confidence level achieved in a previous MV task result generated by the MV task AI model (see figure 1 element 122), however Zheng and Lee do not teach wherein the processor is configured to: select the inference mode based on energy consumption levels observed in prior inference operations. Kesavan teaches wherein the processor is configured to: select the inference mode based on energy consumption levels observed in prior inference operations (“In other embodiments, the ML/NN model may assign certain higher or lower weights to certain servers/nodes to achieve improved probability with respect to network traffic and/or power requirements of those servers/nodes. Here, the output of the ML/NN model can be identification of the recommended or suggested target server cluster system and/or a target server/node within a server cluster system that is best suited to execute and/or run a particular application, task, job, program, or operation, such as any one or more of servers/nodes 222, 224, 226.” See page 30 paragraph 0066 lines 7-14). Kesavan is analogous art in the same field of endeavor as the claimed invention. Kesavan is directed towards network resource allocation (“The present disclosure described herein relates to a method and system for a dynamic resource management and allocation for cluster networks.” Page 1 paragraph 0001). A person of ordinary skill in the art before the effective filing date of the claimed invention, would have found it obvious to combine the system of Zheng and Lee with Kesavan by utilizing Kesavan’s server selection scheme inside the server selection system of Zheng, with the expectation that doing so would lead to better resource allocation and improved energy savings (see Kesavan “…in order to better allocate network resources and improve energy savings and efficiency within a cluster network system.” Page 2 Paragraph 0006) With respect to claim 18, Zheng and Lee teach the method of claim 11, and in view of Kesavan teach all additional limitations in consideration of claim 8, due to the substantial similarities of claims 18 and 8, with claim 18 being directed towards the method that the claim 1 processor is configured to perform in claim 8. With respect to claim 19, Zheng, Lee and Kesavan teach the method of claim 18. Zheng teaches it further comprising communicating with one or more AI-MV servers of the plurality of AI-MV servers to offload portions of the AI-MV task for processing in a split inference mode (“In at least some embodiments, a job object may be generated upon receiving the execution request 812 as described earlier, indicating any dependencies on other jobs (such as the execution of a recipe for feature processing), and the job may be placed in a queue. For batch mode 865, for example, one or more servers may be identified to run the model. “ page 50 col 20 lines 32-28). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to REBECCA C WILLIAMS whose telephone number is (571)272-7074. The examiner can normally be reached M-F 7:30am - 4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew W Bee can be reached at (571)270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /REBECCA COLETTE WILLIAMS/Examiner, Art Unit 2677 /ANDREW W BEE/Supervisory Patent Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Nov 19, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718365
INTELLIGENT PLAN OPTIMIZATION METHOD AND SYSTEM
2y 4m to grant Granted Aug 25, 2026
Patent 12705526
Image processing using photonic quantum computing
3y 7m to grant Granted Aug 11, 2026
Patent 12633080
SYSTEMS AND METHODS FOR INSPECTION OF GAS PLUME USING OBJECT DETECTION AND SEGMENTATION MODELS
1y 5m to grant Granted May 19, 2026
Patent 12626335
IMAGE PROCESSING METHOD, ELECTRONIC DEVICE, AND STORAGE MEDIUM
2y 10m to grant Granted May 12, 2026
Patent 12620212
Locked-Model Multimodal Contrastive Tuning
3y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
36%
Grant Probability
99%
With Interview (+70.0%)
3y 3m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 11 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month