Prosecution Insights
Last updated: August 17, 2026
Application No. 17/684,680

RUNTIME-SPECIFIC PARTITIONING OF MACHINE LEARNING MODELS

Non-Final OA §103§112
Filed
Mar 02, 2022
Examiner
LIN, HSING CHUN
Art Unit
2195
Tech Center
2100 — Computer Architecture & Software
Assignee
Adobe Inc.
OA Round
3 (Non-Final)
61%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
71 granted / 117 resolved
+5.7% vs TC avg
Strong +81% interview lift
Without
With
+81.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
21 currently pending
Career history
152
Total Applications
across all art units

Statute-Specific Performance

§101
15.3%
-24.7% vs TC avg
§103
37.3%
-2.7% vs TC avg
§102
7.0%
-33.0% vs TC avg
§112
34.4%
-5.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 117 resolved cases

Office Action

§103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending in this application. Response to Arguments Applicant’s arguments regarding the rejections of claims 1-20 under 35 U.S.C. 112b have been fully considered and are persuasive. The rejections have been withdrawn. However, new 35 U.S.C. 112b rejections are applied to claims 1-20 based on the amendments. Applicant's arguments regarding the 35 U.S.C. 103 rejections of claims 1-20 have been fully considered but they are either unpersuasive or moot in light of the references being applied in the current rejection. Regarding the 35 U.S.C. 103 rejections, the applicant argues the following in the remarks: Ki and Ng fail to teach the plurality of partitions configured so that the machine learning model releases the GPU during an inference to process the document. Ki and Ng fail to teach sharing the backbone partition based on the runtime requirements using any of the GPU, a CPU, or a tensor processing unit. Examiner has thoroughly considered Applicant’s arguments, but respectfully finds them unpersuasive for at least the following reasons: As to point (a), the examiner finds this point moot in light of the references being applied in the current rejection. As to point (b), the examiner respectfully disagrees. Applicant argues that the use of an 8-bit backbone that can be shared using the processing unit that is most appropriate based on runtime requirements improves performance and that this optimization is not taught by Ki or Ng. However, this is not what the claims recite. The claims do not recite that a most appropriate processing unit based on runtime requirements is selected. Ng teaches the recited limitation because it recites in [0091] “In some embodiments, scene-segmentation module 112 can be implemented by various fast CNN-based semantic segmentation models. In one embodiment, scene-segmentation module 112 can be implemented based on a DeepLabV3+ model; [0093] Modifying the original DeepLabV3+ model by using a fast MobileNetV2 network (described in “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” Sandler et al., arXiv:1801.04381) as the backbone/feature extractor the modified DeepLabV3+ model to speed up and simplify the original DeepLabV3+ model”, [0094] “Quantizing the network parameters and running the network inference in 8-bit integer precision instead of the existing 32-bit floating-point precision to reduce the memory usage and the frequency of memory access”, and [0096] “The above-described network modifications/improvements can significantly speed up the execution of the disclosed scene-segmentation model. For example, the runtime of the disclosed scene segmentation model on Hi3559A CPU can be reduced from about 43 seconds to ˜2 seconds when the above modifications are implemented. In some embodiments, the disclosed scene-segmentation module 112 is only executed during the booting-up phase of embedded fall-detection system 100 or distributed fall-detection system 200 when the system is being calibrated, or when there is no motion in the input video images 104 for some time. As a result, the execution speed of the disclosed scene-segmentation module 112 is sufficient fast to allow room layout information 126 to be generated for an input image before the generation of action labels 124 for that input image”. Ng discloses that an 8-bit backbone partition is shared with a CPU and that the runtime is sufficiently fast. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. As per claims 1, 8, and 15 (line numbers refer to claim 1): Lines 13-15 recite “executing each of the plurality of partitions of the machine learning model to process the document using a runtime environment corresponding to a selected operating system accelerator from among a plurality of operating system accelerators” but this is not supported by the specification. The specification recites in [0020] “Splitting the model into smaller partitions provides for more efficient memory use. Certain blocks of operators use less memory when run on CPU, a GPU, or other delegates of the operating system”. The partitions are not executed on a selected operating system accelerator from among a plurality of operating system accelerators. Claims 2-7, 9-14, and 16-20 are dependent claims of claims 1, 8, and 15, and fail to resolve the deficiencies of claims 1, 8, and 15, so they are rejected for the same reasons. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. As per claims 1, 8, and 15 (line numbers refer to claim 1): Lines 2-3 recite “accessing a machine learning model configured for processing a document, at least in part by using a GPU”, lines 5-6 recite “the plurality of partitions configured so that the machine learning model releases the GPU during an inference to process the document”, and lines 11-12 recite “sharing the backbone partition based on the runtime requirements using any of the GPU, a CPU, or a tensor processing unit”. Therefore, it is unclear why the GPU is utilized, released, and can be utilized again for the backbone partition. Lines 5-6 recite “the plurality of partitions configured so that the machine learning model releases the GPU during an inference to process the document” and lines 13-15 recite “executing each of the plurality of partitions of the machine learning model to process the document using a runtime environment corresponding to a selected operating system accelerator”. Therefore, it is unclear why the GPU is released while processing the document if the document is processed with a selected operating system accelerator. Claims 2-7, 9-14, and 16-20 are dependent claims of claims 1, 8, and 15, and fail to resolve the deficiencies of claims 1, 8, and 15, so they are rejected for the same reasons. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 6-9, 13-16, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ki et al. (US 20220113915 A1 hereinafter Ki), in view of Ng et al. (US 20200211154 A1 hereinafter Ng), and further in view of Wang et al. (US 20210158151 A1 hereinafter Wang). Ki and Ng were cited in a prior office action. As per claim 1, Ki teaches the invention substantially as claimed including a method comprising: accessing a machine learning model configured for processing a document, at least in part by using a GPU ([0003] Some data processing workloads, such as machine learning workloads, may involve the use of models; [0028] In some embodiments, a workflow and/or a model such as a graph, a machine learning model, a neural network; [0034] Workload partitioning may involve splitting a workload (e,g., data and model) across multiple machines (e.g., accelerator devices); [0044] The accelerator 206 may be implemented with any type of device that may include one or more processing resources suitable for an accelerator, for example, a graphics processing unit (GPU); [0040] For example, in some embodiments, the device 200 may be used for ML training and/or inference (e.g., DL and/or DNNs), speech recognition, language processing, image recognition); partitioning the machine learning model into a plurality of partitions of the machine learning model ([0028] In some embodiments, a workflow and/or a model such as a graph, a machine learning model, a neural network, and/or the like, may be partitioned between multiple accelerator devices; [0052] a model (e.g., a graph, an ML model, and/or the like) may be partitioned into portions that may each be assigned to a virtual accelerator to implement model parallelism); characterizing each of the plurality of partitions of the machine learning model with respect to runtime requirements; executing each of the plurality of partitions of the machine learning model to process the document using a runtime environment corresponding to a selected operating system accelerator from among a plurality of operating system accelerators and the runtime requirements of a respective partition from among the plurality of partitions ([0028] In some embodiments, a workflow and/or a model such as a graph, a machine learning model, a neural network, and/or the like, may be partitioned between multiple accelerator devices and/or virtual accelerators in accordance with example embodiments of the disclosure. For example, a host may partition a model between virtual accelerators based on the memory requirements and/or compute times of the portions of the model, as well as the memory resources and/or cores of the virtual accelerators. In some embodiments, based on the partitioning, the host may generate a clustered graph with data groups to be executed by the virtual accelerators and scheduled by a memory manager; [0003] Some data processing workloads, such as machine learning workloads, may involve the use of models; [0127] At operation 887 (Part 3.), the device (e.g., the device implementing the virtual NPUs) may extract one or more operational parameters from the graph partitions provided by the host such as memory usage, task dependencies, timing information for one or more graph partitions, and/or the like for use at runtime; [0040] For example, in some embodiments, the device 200 may be used for ML training and/or inference (e.g., DL and/or DNNs), speech recognition, language processing, image recognition; [0044] The accelerator 206 may be implemented with any type of device that may include one or more processing resources suitable for an accelerator, for example, a graphics processing unit (GPU), a neural processing unit (NPU), tensor processing unit (TPU), an accelerator based on combinational logic, sequential logic, one or more timers, counters, registers, state machines, complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), central processing units (CPUs) such as complex instruction set computer (CISC) processors such as x86 processors and/or reduced instruction set computer (RISC) processors such as ARM processors and/or the like, or any combination thereof.). Ki fails to teach the plurality of partitions configured so that the machine learning model releases the GPU during an inference to process the document, and wherein the plurality of partitions includes a backbone partition optimized for 8-bit processing; sharing the backbone partition based on the runtime requirements using any of the GPU, a CPU, or a tensor processing unit; rendering output using the machine learning model. However, Ng teaches wherein the plurality of partitions includes a backbone partition optimized for 8-bit processing; sharing the backbone partition based on the runtime requirements using any of the GPU, a CPU, or a tensor processing unit (Fig. 1; [0073] In some embodiments, to allow bottom-up pose-estimation models to run in real-time with optimized performance on embedded systems/devices such as embedded fall-detection system 100, the proposed pose-estimation module 106 implements a bottom-up pose-estimation framework with a number of improvements to the existing framework. Some of these modifications/improvements include: [0074] Replacing the commonly used complex VGG16 network (described in “Very Deep Convolutional Networks for Large-Scale Image Recognition,” Simonyan et al., arXiv:1409.1556) with a faster VGG16×4 network (described in “Channel Pruning for Accelerating Very Deep Neural Networks,” He et al., ICCV 2017 and “AMC: AutoML for Model Compression and Acceleration on Mobile Devices,” He et al., ECCV 2018) as the backbone/feature extractor, which has an inference speed 4× faster than the VGG16 network. Note that the term “backbone” herein refers to the neural network which receives an input image and extracts image features for use in subsequent deep-learning tasks such as classification, regression, and segmentation. This speed-up is largely due to performing channel pruning, i.e., reducing the width of the feature map, which in turn shrinks the network into a thinner one; [0077] Quantizing the network parameters and run the network inference in 8-bit integer precision instead of the typical 32-bit floating-point precision. This modification not only reduces the memory usage and the frequency of memory access; [0091] In some embodiments, scene-segmentation module 112 can be implemented by various fast CNN-based semantic segmentation models. In one embodiment, scene-segmentation module 112 can be implemented based on a DeepLabV3+ model; [0093] Modifying the original DeepLabV3+ model by using a fast MobileNetV2 network (described in “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” Sandler et al., arXiv:1801.04381) as the backbone/feature extractor the modified DeepLabV3+ model to speed up and simplify the original DeepLabV3+ model; [0094] Quantizing the network parameters and running the network inference in 8-bit integer precision instead of the existing 32-bit floating-point precision to reduce the memory usage and the frequency of memory access; [0096] The above-described network modifications/improvements can significantly speed up the execution of the disclosed scene-segmentation model. For example, the runtime of the disclosed scene segmentation model on Hi3559A CPU can be reduced from about 43 seconds to ˜2 seconds when the above modifications are implemented. In some embodiments, the disclosed scene-segmentation module 112 is only executed during the booting-up phase of embedded fall-detection system 100 or distributed fall-detection system 200 when the system is being calibrated, or when there is no motion in the input video images 104 for some time. As a result, the execution speed of the disclosed scene-segmentation module 112 is sufficient fast to allow room layout information 126 to be generated for an input image before the generation of action labels 124 for that input image; [0131] First, the proposed face-recognition model can train a lightweight ResNet-18 network (described in “Deep Residual Learning for Image Recognition,” He et al., CVPR 2016) on the MS1M-refine-v2 dataset. Second, the proposed face-recognition model can be configured to quantize the neural network model and run the inference using 8-bit integer precision; [0160] In some embodiments, neural network accelerators 1012 can include any type of microprocessor designed as hardware acceleration for executing AI-based and deep-learning-based programs and models, and in particular various deep learning neural networks such as various CNN and RNN frameworks mentioned in this disclosure. Neural network accelerators 1012 can perform the intended functions of each of the described deep-learning-based modules within the disclosed embedded fall-detection system 100, i.e., pose-estimation module 106, action-recognition module 108, fall-detection module 110, scene-segmentation module 112, face-detection module 116, face-recognition module 118, and the ADL statistics module; [0154] FIG. 10 illustrates an exemplary hardware environment 1000 for the disclosed embedded fall-detection system 100 of FIG. 1 in accordance with some embodiments described herein. Note that hardware environment 1000 can be used to implement each of the one or multiple embedded fall-detection vision sensors 202-1, 202-2, . . . , and 202-N within distributed fall-detection system 200. As can be seen in FIG. 10, hardware environment 1000 can include a bus 1002, one or more processors 1004; [0156] Processors 1004 can include any type of processor, including, but not limited to, one or more central processing units (CPUs), one or more microprocessors, one or more graphic processing units (GPUs), one or more tensor processing units (TPUs); [0166] In some embodiments, to take full advantage of the available processing power of hardware environment 1000, a customized task scheduler can be designed to utilize multiple hardware resources such as ARM CPU and NNIE accelerator in parallel to achieve a maximum processing throughput; [0167] Each task scheduler 1100 can instantiate an arbitrary number of workers to complete the same task in parallel, e.g., three CPU workers: CPU_Worker0, CPU_Worker1, and CPU_Worker2, and two NNIE workers: NNIE_Worker0 and NNIE_Worker0. Furthermore, each worker in task scheduler 1100 can use a different hardware resource (i.e., either the CPU or the NNIE accelerator) offered by hardware environment 1000. In some embodiments, input scheduler 1102 can be configured to receive raw video images as input 1106, and schedule the set of workers to perform the following two streams of tasks on the input video images: (1) the pose-estimation tasks followed by the action-recognition tasks and the fall-detection tasks, and subsequently generating fall detection output including fall alarms, sanitized video clips and/or ADLs as output 1108; and (2) face-detection tasks followed by face-recognition tasks, and subsequently generating person-IDs as output 1108); rendering output using the machine learning model ([0054] Embedded fall-detection system 100 can generate fall-detection output 140 including fall alarms/notifications 140-1 and sanitized video clips 140-2 when human falls are detected; [0052] Note that various embodiments of the disclosed embedded fall-detection system are based on implementing various deep-learning-based fast neural networks; [0062] In some embodiments, after receiving the fall-detection output from server 204, mobile app 212 can play the received sanitized video clip on one or more mobile devices 206-210). It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined Ki with the teachings of Ng to reduce memory usage (see Ng [0094] Quantizing the network parameters and running the network inference in 8-bit integer precision instead of the existing 32-bit floating-point precision to reduce the memory usage and the frequency of memory access). Ki and Ng fail to teach the plurality of partitions configured so that the machine learning model releases the GPU during an inference to process the document. However, Wang teaches the plurality of partitions configured so that the machine learning model releases the GPU during an inference to process the document (Fig. 4; [0031] FIG. 4 shows a graph partition of a model 400 according to one embodiment of the present invention, wherein the model 400 may be an unknown model for the platform 110. As shown in FIG. 4, the model 400 comprises operations 402-420, and the H2O engine 122 can partition the operations into several graphs so that the CPU 132, the GPU 134 and the DLA 138 are used to execute the model 400. Specifically, the CPU 132 receives the input data and sequentially performs the operations 402 and 404 to output two results to the GPU 134 and the DLA 138, respectively. Then, the GPU 134 performs the operations 406 - 414 based on the result generated by the CPU 132 in Step 404, and the DLA 138 performs the operation 416 based on the other result generated by the CPU 132 is Step 404. Then, the CPU 132 executes the operations 418 and 420 based on the results generated by the GPU 134 and DLA 138 in Steps 412, 414 and 416, respectively, to generate the output data; [0013] the model 110 may generally be referred to as an inference model and can include any of a variety of models arranged to generate output data (e.g., inference) from input data; [0014] model 110 comprises an operation that greatly reduces the image resolution such as the ratio of image resolution reduction is greater than a threshold (e.g., the image resolution is reduced from 640*480 to 32*24)). It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined Ki and Ng with the teachings of Wang to execute a model efficiently (see Wang [0019] partition the model 110 to generate two or more graphs for two or more processing circuits. Therefore, the model 110 can be executed efficiently.). As per claim 2, Ki, Ng, and Wang teach the method of claim 1. Ki teaches wherein the runtime requirements comprise at least one of a low-memory complex operator requirement, a high-memory optimized operator requirement, a specialized post-processing requirement, GPU-efficient runtime requirements, or CPU-efficient runtime requirements ([0086] However, in some embodiments, the system illustrated in FIG. 4 may be used to implement a large DL model (e.g., a large DNN model); [0038] In some embodiments, the second type memory in the second memory tier 210 may be implemented with one or more types of memory that may provide relatively high capacity; [0028] a host may partition a model between virtual accelerators based on the memory requirements and/or compute times of the portions of the model, as well as the memory resources and/or cores of the virtual accelerators.). As per claim 6, Ki, Ng, and Wang teach the method of claim 1. Ng teaches wherein the processing of the document comprises detecting portions of the document to apply a plurality of rendering resources to the document ([0159] Bus 1002 is also coupled to camera system 1010. Camera system 1010 is configured to capture a sequence of video images at predetermined resolutions and couple the captured video images to various components within hardware environment 1000 via bus 1002, such as to memory 1006 for buffering and to processors 1004 and neural network accelerators 1012 for various deep-learning and neural network-based operations; [0054] Embedded fall-detection system 100 can generate fall-detection output 140 including fall alarms/notifications 140-1 and sanitized video clips 140-2 when human falls are detected; [0052] Note that various embodiments of the disclosed embedded fall-detection system are based on implementing various deep-learning-based fast neural networks; [0062] In some embodiments, after receiving the fall-detection output from server 204, mobile app 212 can play the received sanitized video clip on one or more mobile devices 206-210; [0021] In some embodiments, if a fall is detected for the person, the process further includes generating a sanitized video clip by: identifying a common background image for the sequence of video images; and superimposing the set of skeleton diagrams of the detected person corresponding to the sequence of video images onto the common background image to obtain a sequence of sanitized video images.). As per claim 7, Ki, Ng, and Wang teach the method of claim 1. Ng teaches wherein the document comprises presentation media and the processing of the document comprises at least one of translation, captioning, or detecting portions of the document to apply a plurality of rendering resources ([0159] Bus 1002 is also coupled to camera system 1010. Camera system 1010 is configured to capture a sequence of video images at predetermined resolutions and couple the captured video images to various components within hardware environment 1000 via bus 1002, such as to memory 1006 for buffering and to processors 1004 and neural network accelerators 1012 for various deep-learning and neural network-based operations; [0054] Embedded fall-detection system 100 can generate fall-detection output 140 including fall alarms/notifications 140-1 and sanitized video clips 140-2 when human falls are detected; [0052] Note that various embodiments of the disclosed embedded fall-detection system are based on implementing various deep-learning-based fast neural networks; [0062] In some embodiments, after receiving the fall-detection output from server 204, mobile app 212 can play the received sanitized video clip on one or more mobile devices 206-210; [0021] In some embodiments, if a fall is detected for the person, the process further includes generating a sanitized video clip by: identifying a common background image for the sequence of video images; and superimposing the set of skeleton diagrams of the detected person corresponding to the sequence of video images onto the common background image to obtain a sequence of sanitized video images.). As per claim 8, it is a system claim of claim 1, so it is rejected for similar reasons. Additionally, Ki teaches system comprising: a memory component; and a processing device coupled to the memory component to perform operations comprising ([0069] The host 424 may also include local storage 468 which may be implemented, for example, with any type of storage device(s) based on any type of memory and/or storage media including solid state media, magnetic media, optical media, and/or the like; [0071] The CPU 458, local storage 468, and/or network interface 432 may communicate, for example, through a system bus 470.): distributing the plurality of partitions to a computing device ([0028] In some embodiments, a workflow and/or a model such as a graph, a machine learning model, a neural network, and/or the like, may be partitioned between multiple accelerator devices and/or virtual accelerators in accordance with example embodiments of the disclosure. For example, a host may partition a model between virtual accelerators based on the memory requirements and/or compute times of the portions of the model, as well as the memory resources and/or cores of the virtual accelerators. In some embodiments, based on the partitioning, the host may generate a clustered graph with data groups to be executed by the virtual accelerators and scheduled by a memory manager; [0030] analyze graph processing and/or deep learning (DL) applications (e.g., deep neural networks (DNNs)) in which computations and/or portions of a DL model may be distributed across multiple machines such as accelerator devices). As per claim 9, Ki, Ng, and Wang teach the system of claim 8. Ki teaches wherein the operations further comprise: executing each of the plurality of partitions of the machine learning model using the runtime environment corresponding to the respective runtime requirements to process the document ([0028] In some embodiments, a workflow and/or a model such as a graph, a machine learning model, a neural network, and/or the like, may be partitioned between multiple accelerator devices and/or virtual accelerators in accordance with example embodiments of the disclosure. For example, a host may partition a model between virtual accelerators based on the memory requirements and/or compute times of the portions of the model, as well as the memory resources and/or cores of the virtual accelerators. In some embodiments, based on the partitioning, the host may generate a clustered graph with data groups to be executed by the virtual accelerators and scheduled by a memory manager; [0003] Some data processing workloads, such as machine learning workloads, may involve the use of models; [0127] At operation 887 (Part 3.), the device (e.g., the device implementing the virtual NPUs) may extract one or more operational parameters from the graph partitions provided by the host such as memory usage, task dependencies, timing information for one or more graph partitions, and/or the like for use at runtime; [0040] For example, in some embodiments, the device 200 may be used for ML training and/or inference (e.g., DL and/or DNNs), speech recognition, language processing, image recognition). Additionally, Ng teaches rendering output based on the processing of the document ([0054] Embedded fall-detection system 100 can generate fall-detection output 140 including fall alarms/notifications 140-1 and sanitized video clips 140-2 when human falls are detected; [0052] Note that various embodiments of the disclosed embedded fall-detection system are based on implementing various deep-learning-based fast neural networks; [0054] Fall-detection engine 101 receives video images 104 as input; [0062] In some embodiments, after receiving the fall-detection output from server 204, mobile app 212 can play the received sanitized video clip on one or more mobile devices 206-210). As per claim 13, it is a system claim of claim 6, so it is rejected for similar reasons. As per claim 14, it is a system claim of claim 7, so it is rejected for similar reasons. As per claim 15, it is a non-transitory computer-readable medium claim of claim 1, so it is rejected for similar reasons. Additionally, Ki teaches a non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations ([0069] The host 424 may also include local storage 468 which may be implemented, for example, with any type of storage device(s) based on any type of memory and/or storage media including solid state media, magnetic media, optical media, and/or the like; [0071] The CPU 458, local storage 468, and/or network interface 432 may communicate, for example, through a system bus 470). As per claim 16, it is a non-transitory computer-readable medium claim of claim 2, so it is rejected for similar reasons. As per claim 19, it is a non-transitory computer-readable medium claim of claim 6, so it is rejected for similar reasons. As per claim 20, it is a non-transitory computer-readable medium claim of claim 7, so it is rejected for similar reasons. Claims 3, 10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Ki, Ng, and Wang, as applied to claims 1, 8 and 15 above, in view of Fu et al. (US 20190325276 A1 hereinafter Fu). Fu was cited in a prior office action. As per claim 3, Ki, Ng, and Wang teach the method of claim 1. Ki, Ng, and Wang fail to teach wherein at least one of the plurality of partitions is reused among a plurality of machine learning models. However, Fu teaches wherein at least one of the plurality of partitions is reused among a plurality of machine learning models ([0030] respective labeled neural network layers can be reused to train part of a neural network for new visual data; [0015] Aspects of the present disclosure are directed toward training neural networks). It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined Ki, Ng, and Wang with the teachings of Fu to implement machine learning quickly (see Fu [0023] applying deep learning to the IoT devices quickly and easily. For example, aspects of the present disclosure can benefit picture regeneration by back propagation and image compression.). As per claims 10 and 17, they are system and non-transitory computer-readable medium claims of claim 3, so they are rejected for similar reasons. Claims 4, 5, 11, 12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Ki, Ng, and Wang, as applied to claims 1 and 15 above, in view of Mody et al. (US 20190005375 A1 hereinafter Mody). Mody was cited in a prior office action. As per claim 4, Ki, Ng, and Wang teach the method of claim 1. Ki, Ng, and Wang fail to teach further comprising applying a partition-specific security profile to at least one of the plurality of partitions. However, Mody teaches further comprising applying a partition-specific security profile to at least one of the plurality of partitions ([0044] As described herein, the key management 408 may receive encrypted keys from the external memories such as the external memory 210. At the secure IP block 202, and during the signal processing, different keys may be supplied for each layer of the multi-layer CNN data.). It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined Ki, Ng, and Wang with the teachings of Mody to provide security (see Mody [0027] Referencing the image frame 102 of FIG. 1 above, the image classification may be performed through secure decryptions and/or encryptions by the secure IP block 202 with the use of corresponding keys as further discussed below.) As per claim 5, Ki, Ng, Wang, and Mody teach the method of claim 4. Mody teaches wherein the partition-specific security profile comprises applying encryption to at least some layers of the machine learning model ([0044] As described herein, the key management 408 may receive encrypted keys from the external memories such as the external memory 210. At the secure IP block 202, and during the signal processing, different keys may be supplied for each layer of the multi-layer CNN data). As per claims 11 and 18, they are system and non-transitory computer-readable medium claims of claim 4, so they are rejected for similar reasons. As per claim 12, it is a system claim of claim 5, so it is rejected for similar reasons. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HSING CHUN LIN whose telephone number is (571)272-8522. The examiner can normally be reached Mon - Fri 9AM-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /H.L./Examiner, Art Unit 2195 /Aimee Li/Supervisory Patent Examiner, Art Unit 2195
Read full office action

Prosecution Timeline

Show 6 earlier events
Jan 21, 2026
Final Rejection mailed — §103, §112
Feb 17, 2026
Interview Requested
Mar 12, 2026
Applicant Interview (Telephonic)
Mar 15, 2026
Examiner Interview Summary
Apr 14, 2026
Request for Continued Examination
Apr 15, 2026
Response after Non-Final Action
Jun 11, 2026
Non-Final Rejection mailed — §103, §112
Jul 10, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705102
PARALLELISM WITH TASK DEPENDENCIES IN A CURATED EXPERIENCE
3y 4m to grant Granted Aug 11, 2026
Patent 12693880
PLURALITY OF SMART NETWORK INTERFACE CARDS ON A SINGLE COMPUTE NODE
4y 7m to grant Granted Jul 28, 2026
Patent 12681757
ACCELERATED MEMORY ALLOCATION
3y 11m to grant Granted Jul 14, 2026
Patent 12675310
VIRTUAL MACHINE DEPLOYMENT BASED ON WORKLOAD AND HARDWARE IN A HYPER-CONVERGED INFRASTRUCTURE (HCI) ENVIRONMENT
3y 8m to grant Granted Jul 07, 2026
Patent 12670036
Identifying Cluster Idleness For Cluster Shutdown
3y 10m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
61%
Grant Probability
99%
With Interview (+81.4%)
3y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 117 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month