Prosecution Insights
Last updated: August 17, 2026
Application No. 18/399,061

ELECTRONIC DEVICE AND CONTROLLING METHOD OF ELECTRONIC DEVICE

Non-Final OA §102§112
Filed
Dec 28, 2023
Priority
Jul 29, 2022 — RE 10-2022-0094786 +3 more
Examiner
LU, HWEI-MIN
Art Unit
Tech Center
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
63%
Grant Probability
Moderate
1-2
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
146 granted / 233 resolved
+2.7% vs TC avg
Strong +40% interview lift
Without
With
+39.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
24 currently pending
Career history
263
Total Applications
across all art units

Statute-Specific Performance

§101
9.5%
-30.5% vs TC avg
§103
50.2%
+10.2% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
29.0%
-11.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 233 resolved cases

Office Action

§102 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is in responsive to communication(s): original application filed on 12/28/2023, said application claims a priority filing date of 07/29/2022. Claims 1-20 are pending. Claims 1, 9, and 15 are independent. Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: 136, 138, and 101 in ¶ [0144]. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification Applicant is reminded of the proper language and format for an abstract of the disclosure. The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details. The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided. The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: ELECTRONIC DEVICE AND CONTROLLING METHOD OF ELECTRONIC DEVICE FOR PROCESSING DATA. The use of the term "Wi-Fi" in ¶ [0058], "Bluetooth" in ¶ [0058], and "Zigbee" in ¶ [0059], which is a trade name or a mark used in commerce, has been noted in this application. The term should be accompanied by the generic terminology; furthermore the term should be capitalized wherever it appears or, where appropriate, include a proper symbol indicating use in commerce such as ™, SM , or ® following the term. Although the use of trade names and marks used in commerce (i.e., trademarks, service marks, certification marks, and collective marks) are permissible in patent applications, the proprietary nature of the marks should be respected and every effort made to prevent their use in any manner which might adversely affect their validity as commercial marks. The disclosure is objected to because of the following informalities: in ¶ [0090], "… the third preprocessing process and the fourth preprocessing process followed by the third preprocessing process …" appears to be "… the third preprocessing process and the fourth preprocessing process followed by the second preprocessing process …"; in ¶ [0123]. "… In FIG. 7, for convenience, the operation of calculating the first operation speed and the operation of calculating the second operation are separately specified …" appears to be "… In FIG. 7, for convenience, the operation of calculating the first operation speed and the operation of calculating the second operation speed are separately specified …"; in ¶ [0145], "When the instruction is executed by the processor (e.g., the processor 130), the processor may perform a function corresponding to …" appears to be "When the instruction is executed by the processor (e.g., the first processor 130), the processor may perform a function corresponding to …". Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3, 5, 11, 13, 17, and 19 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 3, 11, and 17 recite the limitation "… obtain/obtaining an output value corresponding to the at least one input value by performing …" in lines 3, 2-3, and 3-4, respectively, which rendering these claims indefinite because "… obtain an output value corresponding to the at least one input value by receiving …" is also recited in their respective based claim, and it is unclear whether "an output value corresponding to the at least one input value" recited here is the same as or different to "an output value corresponding to the at least one input value " recited in its based claim. Clarification is required. Claims 5, 13, and 19 recite the limitation "…a number of input values that can be processed at an operation speed equal to a difference between the first operation speed and the second operation speed for a predetermined time period …" in lines 3-5, 3-5, and 4-6 respectively, which rendering these claims indefinite because "… determine a number of the at least one input value for the at least one preprocessing process to be transmitted to the external device based on at least one of a difference between the first operation speed and the second operation speed …" is also recited in their respective based claim, and it is unclear whether "a difference between the first operation speed and the second operation speed" recited here is the same as or different to "a difference between the first operation speed and the second operation speed" recited in its based claim. Clarification is required . Claims 5, 13, and 19 recite the limitation "… a number of input values that can be transmitted through a bandwidth of the network…" in lines 6-7, 6-7, and 7-8 respectively, which rendering these claims indefinite because "… a transmission amount of data transmittable through a bandwidth of a network connecting …" is also recited in their respective based claim, and it is unclear whether "a bandwidth" recited here is the same as or different to "a bandwidth" recited in its based claim. Clarification is required . Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-4, 6-12, 14-18, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by ZHANG (CN112508188A, pub. date: 03/16/2021), hereinafter ZHANG Independent Claims 1, 9, and 15 ZHANGE discloses an electronic device (ZHANGE, ¶ [0132] with FIG. 11: an electronic device) comprising: a communicator (ZHANGE, ¶ [0132] with 1120 and 1140 in FIG. 11: an electronic device including a communication interface 1120 and a communication bus 1140; ¶¶ [0142]-[0143]: the communication bus can be divided into address bus, data bus, control bus, etc.; the communication interface is used for communication between the aforementioned terminal and other devices); at least one memory (ZHANGE, ¶ [0132] with 1130 in FIG. 11: an electronic device including a memory 1130; ¶ [0144]: the memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device) configured to store data for a neural network model (ZHANGE, ¶¶ [0133]-[0134] with FIG. 11: memory 1130 is used to store computer programs; obtaining original training samples from a preset storage device according to the sample request information; ¶ [0063]: the target model can be a deep learning neural network model; ¶¶ [0002] and [0005]-[0006]: store the preprocessed original training samples in the server's memory; copy the preprocessed original training samples from memory to the GPU memory; retrieves/obtaining original training samples from a preset storage device according to the sample requirement information); at least one first processor (ZHANGE, ¶ [0132] with 1110 in FIG.11: an electronic device including a processor 1110; ¶¶ [0130] and [0138]-[0139]: control/controlling the graphics processing unit (GPU) to train a second sub-model using the intermediate training samples through the training process; the GPU reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶ [0108] with FIG. 2: the GPU 140 reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶ [0089]: control the GPU to train a second sub-model using the intermediate training samples through the training process; ¶ [0047]: training process 130 can call GPU 140 to train the target model; ¶ [0039]: the GPU reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶¶ [0037], [0015], and [0010]-[0011]: control/controlling the graphics processing unit (GPU) to train a second sub-model using the intermediate training samples; the GPU reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶¶ [0002]-[0005]: models are typically trained using GPUs (Graphics Processing Units); call the GPU so that the GPU can use the preprocessed original training samples in the GPU memory to train the model; controls a graphics processing unit (GPU) to train a second sub-model using the intermediate training samples) configured to perform a training process of the neural network model (ZHANGE, ¶ [0134]: inputting the original training samples into a first sub-model and obtaining intermediate training samples output by the first sub-model; wherein the first sub-model is a part of the target model to be trained; sending the intermediate training samples to the training process at the second training end so that the training process trains a second sub-model according to the intermediate training samples; wherein the second sub-model is another part of the target model; ¶ [0063]: the target model can be a deep learning neural network model; ¶¶ [0002]-[0006]: sending the intermediate training samples to the training process at the second training end, so that the training process trains a second sub-model according to the intermediate training samples); and at least one second processor (ZHANGE, ¶ [0132] with 1110 in FIG.11: an electronic device including a processor 1110; ¶ [0145]: the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components; ¶ [0084]: a large amount of CPU for dataset preprocessing; ¶ [0049]: the execution entity is the working process created by the CPU of the first training end; ¶ [0046] with FIG. 2: each CPU of the first training end creates and runs worker process 120, and the CPU of the second training end creates and runs training process 130 and master process 110 ; ¶¶ [0002]-[0005]: before training the model, a training process needs to be created using the CPU (central processing unit) on the server; when there are many preprocessing items, which consumes a lot of CPU resources; the CPU acquires and processes raw/original training samples) configured to: perform a plurality of preprocessing processes for the training process (ZHANGE, ¶¶ [0002]-[0003]: the first training thread is used to obtain the original training samples from the external storage device, preprocess the original training samples, and store the preprocessed original training samples in the server's memory; when there are many preprocessing items, which consumes a lot of CPU resources; ¶ [0045]: the first sub-model is used to preprocess the original training samples; ¶ [0047] with FIG. 2: worker process 120 continuously reads and preprocesses the dataset according to the sample requirement information, and then sends the intermediate training samples to training process 130; ¶¶ [0065]-[0066]: intermediate training samples can be samples that have been preprocessed by the first sub-model; the first sub-model can be a sub-model to be trained in the target model, or it can be a sub-model in the target model used to preprocess the original training samples; ¶ [0068]: when the first sub-model is a sub-model used for preprocessing in the target model, it is sufficient to obtain the intermediate training samples output by the first sub-model; preprocessing includes, but is not limited to: sample feature normalization, sample feature verification, and illegal value detection; ¶¶ [0083]-[0084]: a preprocessed sample format is customized to be compatible with both dense and sparse features, that is, to be compatible with both common and unique features of the samples, thereby enabling to handle various sample scenarios; the efficiency of GPU deep learning training can be improved, avoiding the training bottleneck caused by the need for a large amount of CPU for dataset preprocessing, improving the training speed of the model, and improving the iteration speed of the model; ¶ [0111]; the functions of sample acquisition and preprocessing are distributed to different working processes on different physical machines (first training end) for processing, which solves the problem of limited CPU processing efficiency on a single physical machine and improves the efficiency of model training; ¶ [0118]: if a worker process is slow in acquiring and preprocessing, the amount of sample file names distributed can be reduced, while a worker process with faster processing speed can distribute more sample file names), determine a first operation speed of at least one preprocessing process of the plurality of preprocessing processes performed by the at least one second processor and a second operation speed of the training process performed by the at least one first processor (ZHANGE, ¶ [0012]: receiving a processing completion message from the working process; determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; ¶ [0040]: after receiving the processing completion message, the main control process determines the processing time of the worker process for the original training samples based on the reception time of the processing completion message and the sending time of the sample demand information; based on the processing time of the worker process for the original training samples, it determines the sample demand information to be sent to the worker process next; wherein, the sample demand information carries the number of original training samples to be acquired by the worker process corresponding to the processing time; ¶ [0091]: after the sub-models are deployed, the processing time for each sub-model is determined; according to the processing time from front to back, the training end where the first sub-model is located obtains the original training samples, sends the intermediate training samples output by the sub-model of the first training end to the second training end, the intermediate training samples output by the sub-model of the second training end to the third training end, and so on, until the sub-model of the last training end has finished processing; ¶¶ [0116]-[0117] with FIGS. 7-8: the second training end determines the processing time of the original training sample by the first training end through the main control process based on the reception time of the processing completion message and the sending time of the sample requirement information; the second training end determines the sample request information to be sent to the first training end next, based on the processing time of the original training samples by the first training end through the main control process; wherein, the sample request information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; ¶ [0140]: determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time), based on the first operation speed being slower than the second operation speed, control the communicator to transmit at least one input value for the at least one preprocessing process to an external device connected to the electronic device (ZHANGE, ¶¶ [0003]-[0004]: when there are many preprocessing items, which consumes a lot of CPU resources; this makes it impossible for the CPU to acquire and process the original training samples at a speed that can keep up with the training speed of the GPU, resulting in a waste of GPU computing power; address the problem that the speed at which the CPU acquires and processes raw training samples cannot meet the training speed of the GPU; ¶¶ [0019], [0043], and [0082]: avoid the problem that the CPU processing speed cannot meet the GPU processing speed caused by training the target model on a single training endpoint, thus solving the training bottleneck of the target model; ¶ [0084]: the efficiency of GPU deep learning training can be improved, avoiding the training bottleneck caused by the need for a large amount of CPU for dataset preprocessing, improving the training speed of the model, and improving the iteration speed of the model; ¶ [0099] with FIG. 2: in the sample request information sent to the first training terminal 210 with a slower processing speed, the number of original training samples to be obtained is small, while in the sample request information sent to the first training terminal 210 with a faster processing speed, the number of original training samples to be obtained is large; ¶ [0118]: the distribution speed of sample file names is controlled by the main control process; this allows for control over the training speed of different worker processes; if a worker process is slow in acquiring and preprocessing, the amount of sample file names distributed can be reduced, while a worker process with faster processing speed can distribute more sample file names; this balances the processing speed of different worker processes and improves the overall model training efficiency), and obtain an output value corresponding to the at least one input value by receiving, through the communicator, a processing result of the external device for the at least one input value from the external device (ZHANGE, ¶¶ [0005]-[0010]: retrieve intermediate training samples output by the first sub-model, and sends the intermediate training samples to the training process; the training process receives the intermediate training samples from the worker process; inputting the original training samples into a first sub-model and obtaining intermediate training samples output by the first sub-model; sending the intermediate training samples to the training process at the second training end, so that the training process trains a second sub-model according to the intermediate training samples; after sending the intermediate training samples to the training process of the second training end, sending a preset processing completion message to the master control process of the second training end in order to receive the next sample request information from the master control process of the second training end; generating transmission information with a preset sample format based on the intermediate training samples; and sending the transmission information to the training process; wherein the transmission information with the sample format includes dense feature vectors and/or sparse feature vectors corresponding to the intermediate training samples; receiving intermediate training samples from the worker process via a training process; wherein the intermediate training samples are the output results of the first sub-model after inputting the original training samples; ¶¶ [0014]-[0015]: a second acquiring module for inputting the original training samples into a first sub-model and acquiring intermediate training samples output by the first sub-model; a first sending module for sending the intermediate training samples to the training process at the second training end, so that the training process trains a second sub-model according to the intermediate training samples; receive intermediate training samples from the working process via the training process; wherein the intermediate training samples are the results output by the first sub-model after inputting original training samples; ¶¶ [0036]-[0037] with FIG. 1: inputs the original training samples into the first sub-model, retrieves the intermediate training samples output by the first sub-model, and sends the intermediate training samples to the training process 130; training process 130 receives intermediate training samples from worker process 120; controls the graphics processing unit (GPU) to train a second sub-model using the intermediate training samples; ¶ [0040]-[0041]: the worker process sends a preset processing completion message to the main control process simultaneously or after sending the intermediate training samples to the training process; the working process generates transmission information with a preset sample format based on the intermediate training samples; and sends the transmission information to the training process; wherein the transmission information in the sample format includes dense feature vectors and/or sparse feature vectors corresponding to the intermediate training samples; ¶[0047]: sends the intermediate training samples to training process 130 so that training process 130 can call GPU 140 to train the target model; ¶ [0049]: the first training end is the first physical machine or the first container in the training physical machine used for model training; ¶ [0052]:the first training end is the first physical machine or the first container in the training physical machine used for model training; ¶ [0062] with S330 in FIG. 3: : input the original training samples into the first sub-model and obtain the intermediate training samples output by the first sub-model; the first sub-model is a part of the target model to be trained; ¶ [0065]: intermediate training samples refer to samples processed by the first sub-model; e.g., intermediate training samples can be samples that have been preprocessed by the first sub-model; ¶ [0067]: after obtaining the intermediate training samples output by the first sub-model, the loss value of the first sub-model is determined based on the intermediate training samples; ¶ [0070] with S340 in FIG. 4: the intermediate training samples are sent to the training process of the second training end so that the training process can train the second sub-model based on the intermediate training samples; ¶¶ [0072]-[0073]: based on the ranking of multiple sub-models, the training end jointly deployed by the first N (N≥1) sub-models is used to obtain the original training samples and output intermediate training samples; the training ends deployed by the subsequent sub-models are used to receive the intermediate training samples output by the previous training end and output the processed intermediate sequence samples to the next training end; since distributed training is executed on different physical machines or containers, in order for the intermediate training samples output by the sub-model of the previous physical machine or container to be directly processed by the sub-model in the next physical machine or container, before sending the intermediate training samples to the next training end (second training end), transmission information with a preset sample format is generated based on the intermediate training samples; the transmission information is sent to the training process of the second training end; the transmission information with the sample format includes the dense feature vector and/or sparse feature vector corresponding to the intermediate training samples; ¶¶ [0088]-[0089] with S420-S430 in FIG. 4: receive intermediate training samples from the working process through the training process; the intermediate training samples are the results output by the first sub-model after the original training samples are input into the first sub-model); Control the GPU to train a second sub-model using the intermediate training samples through the training process; ¶ [0091]: send the intermediate training samples output by the sub-model of the first training end to the second training end, the intermediate training samples output by the sub-model of the second training end to the third training end, and so on, until the sub-model of the last training end has finished processing; ¶¶ [0101]-[0105] with S530-S550 in FIG. 5: inputs the original training samples into the first sub-model, and obtains the intermediate training samples output by the first sub-model; the first training end 210 sends the intermediate training samples to the training process 130 of the second training end 220 through the worker process 120; the working process 120 can send the number of intermediate training samples to the training process 130 of the second training end 220 according to the number of samples sent in the sample requirement information, or according to the current network conditions; the second training end 220 receives intermediate training samples from the first training end 210 through the training process 130; the training process 130 of the second training end 220 receives the number of intermediate training samples sent and stores the number of intermediate training samples sent into the video memory). ZHANGE further discloses a non-transitory computer-readable recording medium (ZHANGE, ¶ [0132] with 1130 in FIG. 11: an electronic device including a memory 113; ¶ [0146]: a computer-readable storage medium is provided; ¶ [0148]: the computer-readable storage medium can be any available medium that a computer can access, or a data storage device) storing instructions (ZHANGE, ¶ [0133] with FIG. 11: memory 1130 is used to store computer programs; ¶ [0146]: a computer-readable storage medium stores instructions) that by at least one processor (ZHANGE, ¶ [0132] with 1110 in FIG.11: an electronic device including a processor 1110; ¶ [0145]: the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components;) execute a method of controlling an electronic device (ZHANGE, ¶ [0132] with FIG. 11: an electronic device) described above (ZHANGE, ¶ [0146]: a computer-readable storage medium is provided, which stores instructions that, when run on a computer, cause the computer to execute the method steps of the working process at the first training end as described, or to execute the method steps of the second training end as described). Claims 2, 10, and 16 ZHANG discloses all the elements as stated in Claims 1, 9, and 15 respectively and further discloses wherein the at least one second processor is further configured to: store a test value for the at least one preprocessing process in the at least one memory, and determine the first operation speed by repeatedly performing the at least one preprocessing process based on the stored test value (ZHANG, ¶ [0091]: after the sub-models are deployed, the processing time for each sub-model is determined; according to the processing time from front to back, the training end where the first sub-model is located obtains the original training samples, sends the intermediate training samples output by the sub-model of the first training end to the second training end, the intermediate training samples output by the sub-model of the second training end to the third training end, and so on, until the sub-model of the last training end has finished processing; ¶¶ [0110]-[0112] with FIG. 2: after the GPU 140 determines that the second sub-model training is complete, that is, after the target model training is complete, it sends a training completion message to the training process 130; the training process 130 sends a training stop message to the main control process 110; the main control process 110 sends a training stop notification message to the worker process 120 of the first training terminal 210; after receiving the training stop notification message, the worker process 120 of the first training terminal 210 stops retrieving original training samples from the storage device; the distribution of sample file names is controlled by a master process, which allows for flexible control over sample processing, including whether to repeat samples multiple times or to shuffle the order of sample files; ¶¶ [0116]-[0118] with FIGS. 7-8: the first training end sends a preset processing completion message to the main control process of the second training end through the worker process, so as to receive the next sample request information from the main control process of the second training end; the second training end receives a processing completion message from the first training end through the main control process; the second training end determines the processing time of the original training sample by the first training end through the main control process based on the reception time of the processing completion message and the sending time of the sample requirement information; the second training end determines the sample request information to be sent to the first training end next, based on the processing time of the original training samples by the first training end through the main control process; wherein, the sample request information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; ¶ [0140]: determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; ¶ [0040]: after receiving the processing completion message, the main control process determines the processing time of the worker process for the original training samples based on the reception time of the processing completion message and the sending time of the sample demand information; based on the processing time of the worker process for the original training samples, it determines the sample demand information to be sent to the worker process next; wherein, the sample demand information carries the number of original training samples to be acquired by the worker process corresponding to the processing time; ¶ [0012]: receiving a processing completion message from the working process; determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time) (ZHANGE, ¶¶ [0003]-[0004]: when there are many preprocessing items, which consumes a lot of CPU resources; this makes it impossible for the CPU to acquire and process the original training samples at a speed that can keep up with the training speed of the GPU, resulting in a waste of GPU computing power; address the problem that the speed at which the CPU acquires and processes raw training samples cannot meet the training speed of the GPU; ¶¶ [0019], [0043], and [0082]: avoid the problem that the CPU processing speed cannot meet the GPU processing speed caused by training the target model on a single training endpoint, thus solving the training bottleneck of the target model; ¶ [0084]: the efficiency of GPU deep learning training can be improved, avoiding the training bottleneck caused by the need for a large amount of CPU for dataset preprocessing, improving the training speed of the model, and improving the iteration speed of the model; ¶ [0099] with FIG. 2: in the sample request information sent to the first training terminal 210 with a slower processing speed, the number of original training samples to be obtained is small, while in the sample request information sent to the first training terminal 210 with a faster processing speed, the number of original training samples to be obtained is large; ¶ [0118]: the distribution speed of sample file names is controlled by the main control process; this allows for control over the training speed of different worker processes; if a worker process is slow in acquiring and preprocessing, the amount of sample file names distributed can be reduced, while a worker process with faster processing speed can distribute more sample file names; this balances the processing speed of different worker processes and improves the overall model training efficiency). Claims 3, 11, and 17 ZHANG discloses all the elements as stated in Claims 2, 10, and 16 respectively and further discloses wherein the at least one second processor is further configured to, based on the first operation speed being faster than the second operation speed, obtain an output value corresponding to the at least one input value by performing the at least one preprocessing process on the at least one input value (ZHANGE, ¶¶ [0003]-[0004]: GPUs are faster than CPUs, and in many cases, multiple GPUs are used to train a model; CPU computing power is relatively limited, especially when there are many preprocessing items, which consumes a lot of CPU resources; this makes it impossible for the CPU to acquire and process the original training samples at a speed that can keep up with the training speed of the GPU, resulting in a waste of GPU computing power; address the problem that the speed at which the CPU acquires and processes raw training samples cannot meet the training speed of the GPU; ¶¶ [0019], [0043], and [0082]: avoid the problem that the CPU processing speed cannot meet the GPU processing speed caused by training the target model on a single training endpoint, thus solving the training bottleneck of the target model; ¶ [0084]: the efficiency of GPU deep learning training can be improved, avoiding the training bottleneck caused by the need for a large amount of CPU for dataset preprocessing, improving the training speed of the model, and improving the iteration speed of the model; ¶ [0099] with FIG. 2: in the sample request information sent to the first training terminal 210 with a slower processing speed, the number of original training samples to be obtained is small, while in the sample request information sent to the first training terminal 210 with a faster processing speed, the number of original training samples to be obtained is large; ¶ [0118]: the distribution speed of sample file names is controlled by the main control process; this allows for control over the training speed of different worker processes; if a worker process is slow in acquiring and preprocessing, the amount of sample file names distributed can be reduced, while a worker process with faster processing speed can distribute more sample file names; this balances the processing speed of different worker processes and improves the overall model training efficiency). Claims 4, 12, and 18 ZHANG discloses all the elements as stated in Claims 2, 10, and 16 respectively and further discloses wherein the at least one second processor is further configured to: control the communicator to transmit the test value to the external device, based on the at least one preprocessing process for the test value being performed by the external device, receive, from the external device through the communicator, information indicating a third operation speed, and determine a number of the at least one input value for the at least one preprocessing process to be transmitted to the external device based on at least one of a difference between the first operation speed and the second operation speed, the third operation speed, and a transmission amount of data transmittable through a bandwidth of a network connecting the electronic device and the external device (ZHANGE, ¶ [0012]: receiving a processing completion message from the working process; determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; ¶ [0040]: after receiving the processing completion message, the main control process determines the processing time of the worker process for the original training samples based on the reception time of the processing completion message and the sending time of the sample demand information; based on the processing time of the worker process for the original training samples, it determines the sample demand information to be sent to the worker process next; wherein, the sample demand information carries the number of original training samples to be acquired by the worker process corresponding to the processing time; ¶ [0091]: after the sub-models are deployed, the processing time for each sub-model is determined; according to the processing time from front to back, the training end where the first sub-model is located obtains the original training samples, sends the intermediate training samples output by the sub-model of the first training end to the second training end, the intermediate training samples output by the sub-model of the second training end to the third training end, and so on, until the sub-model of the last training end has finished processing; ¶¶ [0116]-[0118] with FIGS. 7-8: the second training end determines the processing time of the original training sample by the first training end through the main control process based on the reception time of the processing completion message and the sending time of the sample requirement information; the second training end determines the sample request information to be sent to the first training end next, based on the processing time of the original training samples by the first training end through the main control process; wherein, the sample request information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; the distribution speed of sample file names is controlled by the main control process; this allows for control over the training speed of different worker processes; if a worker process is slow in acquiring and preprocessing, the amount of sample file names distributed can be reduced, while a worker process with faster processing speed can distribute more sample file names; this balances the processing speed of different worker processes and improves the overall model training efficiency; ¶ [0140]: determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time) (ZHANGE, ¶¶ [0003]-[0004]: when there are many preprocessing items, which consumes a lot of CPU resources; this makes it impossible for the CPU to acquire and process the original training samples at a speed that can keep up with the training speed of the GPU, resulting in a waste of GPU computing power; address the problem that the speed at which the CPU acquires and processes raw training samples cannot meet the training speed of the GPU; ¶¶ [0019], [0043], and [0082]: avoid the problem that the CPU processing speed cannot meet the GPU processing speed caused by training the target model on a single training endpoint, thus solving the training bottleneck of the target model; ¶ [0084]: the efficiency of GPU deep learning training can be improved, avoiding the training bottleneck caused by the need for a large amount of CPU for dataset preprocessing, improving the training speed of the model, and improving the iteration speed of the model; ¶ [0099] with FIG. 2: in the sample request information sent to the first training terminal 210 with a slower processing speed, the number of original training samples to be obtained is small, while in the sample request information sent to the first training terminal 210 with a faster processing speed, the number of original training samples to be obtained is large) (ZHANGE, ¶¶ [0005]-[0010]: retrieve intermediate training samples output by the first sub-model, and sends the intermediate training samples to the training process; the training process receives the intermediate training samples from the worker process; inputting the original training samples into a first sub-model and obtaining intermediate training samples output by the first sub-model; sending the intermediate training samples to the training process at the second training end, so that the training process trains a second sub-model according to the intermediate training samples; after sending the intermediate training samples to the training process of the second training end, sending a preset processing completion message to the master control process of the second training end in order to receive the next sample request information from the master control process of the second training end; generating transmission information with a preset sample format based on the intermediate training samples; and sending the transmission information to the training process; wherein the transmission information with the sample format includes dense feature vectors and/or sparse feature vectors corresponding to the intermediate training samples; receiving intermediate training samples from the worker process via a training process; wherein the intermediate training samples are the output results of the first sub-model after inputting the original training samples; ¶¶ [0014]-[0015]: a second acquiring module for inputting the original training samples into a first sub-model and acquiring intermediate training samples output by the first sub-model; a first sending module for sending the intermediate training samples to the training process at the second training end, so that the training process trains a second sub-model according to the intermediate training samples; receive intermediate training samples from the working process via the training process; wherein the intermediate training samples are the results output by the first sub-model after inputting original training samples; ¶¶ [0036]-[0037] with FIG. 1: inputs the original training samples into the first sub-model, retrieves the intermediate training samples output by the first sub-model, and sends the intermediate training samples to the training process 130; training process 130 receives intermediate training samples from worker process 120; controls the graphics processing unit (GPU) to train a second sub-model using the intermediate training samples; ¶ [0040]-[0041]: the worker process sends a preset processing completion message to the main control process simultaneously or after sending the intermediate training samples to the training process; the working process generates transmission information with a preset sample format based on the intermediate training samples; and sends the transmission information to the training process; wherein the transmission information in the sample format includes dense feature vectors and/or sparse feature vectors corresponding to the intermediate training samples; ¶[0047]: sends the intermediate training samples to training process 130 so that training process 130 can call GPU 140 to train the target model; ¶ [0049]: the first training end is the first physical machine or the first container in the training physical machine used for model training; ¶ [0052]:the first training end is the first physical machine or the first container in the training physical machine used for model training; ¶ [0062] with S330 in FIG. 3: : input the original training samples into the first sub-model and obtain the intermediate training samples output by the first sub-model; the first sub-model is a part of the target model to be trained; ¶ [0065]: intermediate training samples refer to samples processed by the first sub-model; e.g., intermediate training samples can be samples that have been preprocessed by the first sub-model; ¶ [0067]: after obtaining the intermediate training samples output by the first sub-model, the loss value of the first sub-model is determined based on the intermediate training samples; ¶ [0070] with S340 in FIG. 4: the intermediate training samples are sent to the training process of the second training end so that the training process can train the second sub-model based on the intermediate training samples; ¶¶ [0072]-[0073]: based on the ranking of multiple sub-models, the training end jointly deployed by the first N (N≥1) sub-models is used to obtain the original training samples and output intermediate training samples; the training ends deployed by the subsequent sub-models are used to receive the intermediate training samples output by the previous training end and output the processed intermediate sequence samples to the next training end; since distributed training is executed on different physical machines or containers, in order for the intermediate training samples output by the sub-model of the previous physical machine or container to be directly processed by the sub-model in the next physical machine or container, before sending the intermediate training samples to the next training end (second training end), transmission information with a preset sample format is generated based on the intermediate training samples; the transmission information is sent to the training process of the second training end; the transmission information with the sample format includes the dense feature vector and/or sparse feature vector corresponding to the intermediate training samples; ¶¶ [0088]-[0089] with S420-S430 in FIG. 4: receive intermediate training samples from the working process through the training process; the intermediate training samples are the results output by the first sub-model after the original training samples are input into the first sub-model); Control the GPU to train a second sub-model using the intermediate training samples through the training process; ¶ [0091]: send the intermediate training samples output by the sub-model of the first training end to the second training end, the intermediate training samples output by the sub-model of the second training end to the third training end, and so on, until the sub-model of the last training end has finished processing; ¶¶ [0101]-[0105] with S530-S550 in FIG. 5: inputs the original training samples into the first sub-model, and obtains the intermediate training samples output by the first sub-model; the first training end 210 sends the intermediate training samples to the training process 130 of the second training end 220 through the worker process 120; the working process 120 can send the number of intermediate training samples to the training process 130 of the second training end 220 according to the number of samples sent in the sample requirement information, or according to the current network conditions; the second training end 220 receives intermediate training samples from the first training end 210 through the training process 130; the training process 130 of the second training end 220 receives the number of intermediate training samples sent and stores the number of intermediate training samples sent into the video memory). Claims 6, 14, and 20 ZHANG discloses all the elements as stated in Claims 4, 12, and 18 respectively and further discloses wherein the at least one second processor is further configured to, based on receiving, from the external device through the communicator, a processing result of the external device for the test value, determine the third operation speed based on a time point at which the test value is transmitted to the external device and a time point at which the processing result is received from the external device (ZHANGE, ¶ [0012]: receiving a processing completion message from the working process; determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; ¶ [0040]: after receiving the processing completion message, the main control process determines the processing time of the worker process for the original training samples based on the reception time of the processing completion message and the sending time of the sample demand information; based on the processing time of the worker process for the original training samples, it determines the sample demand information to be sent to the worker process next; wherein, the sample demand information carries the number of original training samples to be acquired by the worker process corresponding to the processing time; ¶ [0091]: after the sub-models are deployed, the processing time for each sub-model is determined; according to the processing time from front to back, the training end where the first sub-model is located obtains the original training samples, sends the intermediate training samples output by the sub-model of the first training end to the second training end, the intermediate training samples output by the sub-model of the second training end to the third training end, and so on, until the sub-model of the last training end has finished processing; ¶¶ [0116]-[0118] with FIGS. 7-8: the second training end determines the processing time of the original training sample by the first training end through the main control process based on the reception time of the processing completion message and the sending time of the sample requirement information; the second training end determines the sample request information to be sent to the first training end next, based on the processing time of the original training samples by the first training end through the main control process; wherein, the sample request information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; the distribution speed of sample file names is controlled by the main control process; this allows for control over the training speed of different worker processes; if a worker process is slow in acquiring and preprocessing, the amount of sample file names distributed can be reduced, while a worker process with faster processing speed can distribute more sample file names; this balances the processing speed of different worker processes and improves the overall model training efficiency; ¶ [0140]: determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time) (ZHANGE, ¶¶ [0003]-[0004]: when there are many preprocessing items, which consumes a lot of CPU resources; this makes it impossible for the CPU to acquire and process the original training samples at a speed that can keep up with the training speed of the GPU, resulting in a waste of GPU computing power; address the problem that the speed at which the CPU acquires and processes raw training samples cannot meet the training speed of the GPU; ¶¶ [0019], [0043], and [0082]: avoid the problem that the CPU processing speed cannot meet the GPU processing speed caused by training the target model on a single training endpoint, thus solving the training bottleneck of the target model; ¶ [0084]: the efficiency of GPU deep learning training can be improved, avoiding the training bottleneck caused by the need for a large amount of CPU for dataset preprocessing, improving the training speed of the model, and improving the iteration speed of the model; ¶ [0099] with FIG. 2: in the sample request information sent to the first training terminal 210 with a slower processing speed, the number of original training samples to be obtained is small, while in the sample request information sent to the first training terminal 210 with a faster processing speed, the number of original training samples to be obtained is large). Claims 7 ZHANG discloses all the elements as stated in Claim 1 and further discloses wherein the at least one second processor is further configured to, based on information on an operation speed of a processing process having the same operation structure as that of the at least one preprocessing process being stored in the at least one memory, determine the first operation speed based on the information on the operation speed stored in the at least one memory (ZHANGE, ¶ [0012]: receiving a processing completion message from the working process; determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; ¶ [0040]: after receiving the processing completion message, the main control process determines the processing time of the worker process for the original training samples based on the reception time of the processing completion message and the sending time of the sample demand information; based on the processing time of the worker process for the original training samples, it determines the sample demand information to be sent to the worker process next; wherein, the sample demand information carries the number of original training samples to be acquired by the worker process corresponding to the processing time; ¶ [0091]: after the sub-models are deployed, the processing time for each sub-model is determined; according to the processing time from front to back, the training end where the first sub-model is located obtains the original training samples, sends the intermediate training samples output by the sub-model of the first training end to the second training end, the intermediate training samples output by the sub-model of the second training end to the third training end, and so on, until the sub-model of the last training end has finished processing; ¶¶ [0116]-[0118] with FIGS. 7-8: the second training end determines the processing time of the original training sample by the first training end through the main control process based on the reception time of the processing completion message and the sending time of the sample requirement information; the second training end determines the sample request information to be sent to the first training end next, based on the processing time of the original training samples by the first training end through the main control process; wherein, the sample request information carries the number of original training samples to be acquired by the first training end corresponding to the processing time; the distribution speed of sample file names is controlled by the main control process; this allows for control over the training speed of different worker processes; if a worker process is slow in acquiring and preprocessing, the amount of sample file names distributed can be reduced, while a worker process with faster processing speed can distribute more sample file names; this balances the processing speed of different worker processes and improves the overall model training efficiency; ¶ [0140]: determining the processing time of the original training samples by the first training end based on the receiving time of the processing completion message and the sending time of the sample demand information; determining the sample demand information to be sent to the working process next based on the processing time of the original training samples by the first training end; wherein the sample demand information carries the number of original training samples to be acquired by the first training end corresponding to the processing time) (ZHANGE, ¶¶ [0003]-[0004]: when there are many preprocessing items, which consumes a lot of CPU resources; this makes it impossible for the CPU to acquire and process the original training samples at a speed that can keep up with the training speed of the GPU, resulting in a waste of GPU computing power; address the problem that the speed at which the CPU acquires and processes raw training samples cannot meet the training speed of the GPU; ¶¶ [0019], [0043], and [0082]: avoid the problem that the CPU processing speed cannot meet the GPU processing speed caused by training the target model on a single training endpoint, thus solving the training bottleneck of the target model; ¶ [0084]: the efficiency of GPU deep learning training can be improved, avoiding the training bottleneck caused by the need for a large amount of CPU for dataset preprocessing, improving the training speed of the model, and improving the iteration speed of the model; ¶ [0099] with FIG. 2: in the sample request information sent to the first training terminal 210 with a slower processing speed, the number of original training samples to be obtained is small, while in the sample request information sent to the first training terminal 210 with a faster processing speed, the number of original training samples to be obtained is large). Claims 8 ZHANG discloses all the elements as stated in Claim 1 and further discloses wherein the at least one first processor comprises a graphics processing unit (GPU) or a neural processing unit (NPU) (ZHANGE, ¶ [0132] with 1110 in FIG.11: an electronic device including a processor 1110; ¶¶ [0130] and [0138]-[0139]: control/controlling the graphics processing unit (GPU) to train a second sub-model using the intermediate training samples through the training process; the GPU reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶ [0108] with FIG. 2: the GPU 140 reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶ [0089]: control the GPU to train a second sub-model using the intermediate training samples through the training process; ¶ [0047]: training process 130 can call GPU 140 to train the target model; ¶ [0039]: the GPU reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶¶ [0037], [0015], and [0010]-[0011]: control/controlling the graphics processing unit (GPU) to train a second sub-model using the intermediate training samples; the GPU reads the intermediate training samples from the video memory and uses the intermediate training samples to train the second sub-model; ¶¶ [0002]-[0005]: models are typically trained using GPUs (Graphics Processing Units); call the GPU so that the GPU can use the preprocessed original training samples in the GPU memory to train the model; controls a graphics processing unit (GPU) to train a second sub-model using the intermediate training samples), and the at least one second processor comprises a central processing unit (CPU) or a microprocessor unit (MPU) (ZHANGE, ¶ [0132] with 1110 in FIG.11: an electronic device including a processor 1110; ¶ [0145]: the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components; ¶ [0084]: a large amount of CPU for dataset preprocessing; ¶ [0049]: the execution entity is the working process created by the CPU of the first training end; ¶ [0046] with FIG. 2: each CPU of the first training end creates and runs worker process 120, and the CPU of the second training end creates and runs training process 130 and master process 110 ; ¶¶ [0002]-[0005]: before training the model, a training process needs to be created using the CPU (central processing unit) on the server; when there are many preprocessing items, which consumes a lot of CPU resources; the CPU acquires and processes raw/original training samples). Allowable Subject Matter Claims 5, 13, and 19 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Claims 5, 13, and 19 ZHANG discloses all the elements as stated in Claims 4, 12, and 18 respectively above. MAHESH (US 2022/0172325 A1, filed on 02/16/2022) discloses in ¶¶ [0002]-[0004] and [0015]-[0020] that (1) machine learning is typically performed in a server environment with training server nodes having one or more central processing units (CPUs) that perform iterative processing, which may include offloading processing tasks to a graphics processing unit (GPU) or other accelerator; (2) effective machine learning involves data preprocessing and data augmentation to diversify and randomize the input data; (3) the data may also need preprocessing operations to clean up the data for use in the training model; (4) traditionally, preprocessing can be offloaded to a dedicated data preprocessing accelerator on the training server node, which increases the cost to put a dedicated preprocessing processor on the server node; (5) preprocessing can also be offloaded the GPU, which requires the GPU to support the preprocessing capabilities, and which also "steals" processing cycles from the GPU by using processing bandwidth of the GPU that will then be unavailable for processing; (6) executes a distributed application can have a processor node to generate a request for data and a storage node that stores the requested data; (7) the distributed application can be a machine learning model or other distributed application that executes on a processor node, where the execution involves repeated access to data from the storage node; (8) the processor node or the training node can offload preprocessing for the machine learning to the storage node; (9) the processor node will use the data for iterative processing to train a machine learning model; (10) the storage node receives the request for the data, reads the data, preprocess the data to perform data transformation on the data, and provides the preprocessed data to the processor node for the iterative processing; (11) in contrast to traditional systems that can offload preprocessing to a dedicated data preprocessing accelerator on the training server node or to a GPU (graphics processing unit), the system can offload preprocessing to the external storage; (12) offloading to the dedicated accelerator or to the GPU both involve increased system cost, while offloading to the storage node can free up processing bandwidth with minimal or no additional system cost; (13) the preprocessing can be preprocessing of random data samples requested by a machine learning application for generation of a machine learning model; (14) the preprocessing is contrasted to, and can be in addition to, dedicated functions such as data compression or data encryption for communication or for storage; (15) rather, the preprocessing involves data transformation to prepare the data for iterative processing for the machine learning model; (16) the preprocessing by the external storage can enable data augmentation on training data before it is provided to a dedicated machine learning or artificial intelligence (AI) processor, such as a central processing unit (CPU) executing a machine learning application, a GPU accelerator, or other dedicated accelerator processor; (17) the external storage includes a CPU or dedicated special purpose processor to manage the interface to the storage nodes; (18) the CPU or special purpose processor can perform the preprocessing; (19) the external storage cluster can include a controller, which can be implemented as a processor or server to manage the distribution of data and data access requests to the storage nodes, and each storage node can include a processor or CPU to manage operations distributed to it; (20) enable on-demand control over specific preprocessing operations; e.g., one or more triggers in a command identify a specific operation to be performed on the data prior to providing it to the processor node for use in training; and (21) thus, the processor node can generate a request for data with a specific request for a type of preprocessing, and the storage node can access the data and perform the specific operation on the data before sending it to the processor node. MAHESH further discloses in ¶¶ [0025]-[0041] with FIG. 1 that (1) external storage 110 represents a cluster of P storage nodes, specifically, storage node 120[1:P], collectively storage nodes 120, where P is an integer; (2) the cluster of storage nodes 120 provides storage of data for use by the cluster of processing nodes or processor nodes, identified as N training nodes 140[1:N], collectively training nodes 140, where N is an integer; (3) training nodes 140[1:N] include corresponding CPUs (central processing units) 142[1:N], collectively CPUs 142; (4) training nodes 140 each include M accelerators, which can be accelerator hardware such as a FPGA (field programmable gate array) or other programmable logic, graphics processors or GPUs, or other hardware to offload computations; (5) system 100 represents the accelerators in each training node 140[1:N] as accelerators (ACC) 144[1:M], collectively accelerators 144, where M is an integer; (5) training nodes 140 couple to external storage 110 over storage fabric 130; (6) system 100 can include other communication fabrics that are not specifically illustrated, such as a fabric interconnecting training nodes to allow the transfer of tasks between different processing nodes; (6) storage fabric 130 represents a communication network or communication infrastructure to connect storage nodes 120 to training nodes 140; (7) system 100 performs machine learning by iterative processing operations in training nodes 140 based on data obtained from storage nodes 120; (8) training data can be split or sharded into multiple serialized files to feed to the different training nodes parallelly, wherein splitting and sending of data in parallel can allow training nodes 140 to continue processing data instead of being starved, waiting for data; (9) storage nodes 120 perform preprocessing operations on training data prior to sending the data to training nodes 140; (10) processor 124 performs one or more preprocessing operations to perform data transformation on the training data requested by the training nodes prior to providing the requested data; (11) block 126, illustrated in gray, represents the preprocessed object or object file. Preprocessed (PRE-PROC) data 132 and preprocessed (PRE-PROC) data 134 represent objects or files or streams of data provided to training nodes 140 from one or more storage nodes 120 over storage fabric 130: (12) storage nodes 120 perform the preprocessing on the training data, and as such, storage nodes 120 can optimize data for the data pipeline within the training framework being used; (13) the operation by storage nodes 120 prevent CPUs 142 from needing to preprocess the data while the current batch is being trained on accelerators 144; (14) the on-demand preprocessing can be performed with specifically requested data transformations for the training flow; (15) system 100 allows training based on large volumes of data (e.g., in the multiple terabytes), where the data does not fit in the volatile memory (e.g., DRAM (dynamic random access memory)) of the training node, requiring repeated fetching from storage; (16) the repeated fetching from storage nodes 120 involves preprocessing at the storage nodes to prepare the data for training on CPUs 142; (17) thus, system 100 offloads preprocessing to storage nodes 120, which improves latency of the data pipeline, by not starving out CPUs 142 or accelerators 144 due to CPU processing bottlenecks that traditionally occur due to the CPU performing the data preprocessing; (18) system 100 can enable direct loading of data into accelerators 144 from storage nodes 120; (19) storage nodes 120 independently apply requested preprocessing functions on demand to data objects fetched from their independent storage resources; e.g., different training nodes 140 can make requests for data with different preprocessing requests to different storage nodes, or to the same storage node at different times; (20) as such, each data request can include a request for preprocessing that a receiving storage node 120 can apply to the data; (21) processor nodes (training nodes 140) can execute machine learning application software that generates a request for data from external storage 110: and (22) external storage 110 can direct the request to an individual storage node 120 or a selected storage node that stores the data, to access and preprocess the data. MAHESH also discloses in ¶¶ [0042]-[0057] with FIG. 2 that (1) External storage 210 includes P storage nodes, illustrated as storage nodes 220[1:P], collectively storage nodes 220; (2) storage nodes 220 include objects 222 or other data stored at the nodes, processor (PROC) 224 to perform preprocessing operations on data read or fetched from storage, and block 226 to represent preprocessed data; (3) processor node 240 represents a training node; (4) processor node 240 includes CPU 242, which represents one or more processor devices for the processor node; (5) processor node 240 includes M accelerators, identified as accelerators (ACC) 244[1 :M], collectively accelerators 244; (6) CPU 242 can offload iterative machine learning processing tasks to accelerators 244; (7) processor node 240 includes software stack 250, which represents a software stack for the machine learning training; (8) each software component can be an agent or application executing on CPU 242 or accelerator 244 or both CPU 242 and accelerator 244; (9) the software components represent the control logic provided for machine learning, which is executed by hardware resources of processor node 240; (10) software stack 250 includes training application 252, which represents a training application that manages the storing of data in external storage 210, the retrieving of data from the external storage, and the operations within processor node 240 by CPU 242 and accelerators 244 for training; (11) software stack 250 can include training framework 254 to provide software features callable or executable by training application 252; (12) training framework 254 can include multiple operators 256, which represent functions controllable by the data pipeline for data; (13) plugin 258 represents one or more software components that interface with external storage 210, enabling the generation of specific commands to store or retrieve data from the COS; (14) training application 252 stores a serialized data record file or multiple data records file as one or more objects into external storage 210 (COS); (15) the storing can be performed by an API (application programming interface) of training application 252, training framework 254, and plugin 258; e.g., enable training application 252 to store data through a REST API via framework 254 and plugin 258; (16) through plugin 258, the REST API PUT call to external storage 210 can include additional custom metadata to indicate the format of the object and the ability requested from storage for preprocessing on the stored object on demand; (17) storing an entire object in a selected storage node can enable the storage node to hold the entire data record and enable subsequent preprocessing without having to access the data from other storage nodes, which enable preprocessing of an object in its entirety on a selected storage node in response to a subsequent data request command (e.g., a GET request); (18) having the whole object in one storage node avoids the need for reassembling data from across storage nodes 220 before processing, which can avoid access latency. In one example, external storage 210 will spread different objects across different storage nodes 220 for parallel throughput; (19) training application 252 can invoke operators 256 on demand on training data and the software layer will pass on handling of the operations to external storage 210 via API calls or other commands; (20) operators 256 can enable offloading preprocessing from processor node 240 to a selected storage node 220; and (21) plugin 258 and operators 256 provide API extension for training framework 254 for training application 252 to invoke to support preprocessing offload to external storage 210. Anil et al. (US 2023/0105476 A1, filed on 03/06/2020) discloses in ¶¶ [0003]-[0015] that (1) improving the computational throughput of specialized hardware accelerators by using a plurality of available computing devices to provide preprocessed data for hardware accelerator consumption; (2) receiving a processing pipeline of preprocessing operations and operation data representing operations to be executed on a plurality of hardware accelerators, wherein the operations represented by the operation data can be, e.g., operations for training or executing a machine learning model; (3) the operations specified by the operation data can include operations that the hardware accelerators are optimized to perform, e.g., matrix multiplication; (4) the processing pipeline can specify the operation data as a computational graph having nodes representing operations and edges representing data dependencies between the operations; (5) computing devices assigned the preprocessing operations can be configured to continuously preprocess input for eventual processing by hardware accelerators assigned the operation data for the processing pipeline; (6) the hardware accelerators can be configured to fetch enqueued input from multiple computing devices assigned the preprocessing operations; (7) because the specialized hardware accelerators are generally capable of processing data at speeds that are orders of magnitude faster than speeds at which general-purpose devices can transform raw input data into preprocessed inputs for the hardware accelerators, a computational bottleneck can occur in which a hardware accelerator is intermittently idle as it waits for more input to become available, i.e., the hardware accelerator experiences "starvation" of available input for processing; (8) this computational bottleneck is only worsened as incremental improvements to hardware accelerators continue to outpace computational improvements to general-purpose devices; (9) similar bottlenecks occur based on memory-read speed and network bandwidth; (10) by employing the techniques described in this specification, each hardware accelerator assigned the operation data can have the number of computing devices the hardware accelerator fetches the inputs from scaled to limit or prevent starvation, i.e., improving computational throughput of preprocessed inputs to the hardware accelerator so the time the hardware accelerator is idle and waiting for input is limited or eliminated altogether; (11) in addition to reducing a computational bottleneck, the distribution of data generation across general-purpose computing devices can alleviate the memory bottleneck that often limits the processing speed of a computing device processing input data to the amount of memory that can be cached at any one time on the computing device, because the collective memory of multiple computing devices assigned to a hardware accelerator can be leveraged; (12) because hardware accelerators are generally more expensive to build and maintain than inexpensive and widely available general-purpose devices, techniques described in this specification can cheaply and efficiently scale a number of computing devices providing data for hardware accelerator consumption; (13) as hardware accelerators are replaced with improved hardware, additional computing devices can be programmatically assigned to keep up with the increased computational demand of the accelerator; and (14) a distributed computing system implementing techniques described in this specification can be configured to identify a respective ratio of computing device to hardware accelerator assignments for each hardware accelerator, which can be flexibly implemented across a group of hardware accelerators of different architectures and computational characteristics. Anil further discloses in ¶¶ [0023]-[] with FIG. 1 that (1) the processing pipeline 105 includes preprocessing operations 120 and operation data 125; (2) the processing pipeline 105 can be represented as a software object defined by functions of a processing Application Program Interface ("API") 130 implemented by the distributed computing system 100; (3) the processing pipeline 105 can be represented as a software object defined by functions of a processing Application Program Interface ("API") 130 implemented by the distributed computing system 100; (4) the preprocessing operations are any operations that need to be performed on input data to prepare the data for processing by a machine learning model; (5) the preprocessing operations 120 can receive, as input, raw data, and generate, as output, preprocessed data that filters out extraneous or irrelevant information; (6) in addition to filtering out useless information, a computing device executing the preprocessing operations 120 can process raw data and perform data augmentation; (7) if the processing pipeline 105 represents operations for training or executing a machine learning model, e.g., a neural network, then the preprocessing operations 120 can be operations for preparing raw input data to be received as properly-formatted input for the machine learning model being deployed on the hardware accelerators of the distributed computing system 100; (8) examples of preprocessing operations can include binarizing, standardizing, or normalizing input data; restructuring input data into an acceptable data format, e.g., a tensor; and encoding features of the input data, e.g., one-hot encoding of categorical features; and (9) the metadata 127 can specify information that can be used by a scheduling engine 145 in scheduling the processing pipeline 105 for execution in the distributed computing system 100, e.g., the identity of a user associated with the processing pipeline 105, a priority level for execution of the processing pipeline 105, and preferred types of computing devices to execute the processing pipeline 105 on. Anil also discloses in ¶¶ [0051]-[0093] with FIGS. 2A-2C that (1) consider a processing pipeline representing operations for training a neural network: (i) preprocessing operations for preparing training examples; (ii) operation data representing operations for training the neural network; and (iii) additional metadata; (2) a distributed computing system that includes the computing devices 200 and the hardware accelerators 210 can assign, through a scheduler engine, (i) each computing device the preprocessing operations and (ii) each hardware accelerator with the computational graph; (3) each of the computing devices 200 can generate preprocessed input as training examples; (4) specifically, each of the computing devices 200 can process input data to generate a plurality of training examples such that a hardware accelerator assigned the operation data can receive the training examples, as input, and process the training examples according to the operation data; (5) each of the computing devices 200 can execute the preprocessing operations to generate training examples as separate stateless services; (6) specifically, each of the computing devices 200 can generate training examples independent of one another, e.g., as randomized training examples, even if the computing devices 200 are preprocessing the same dataset; (7) as a result, the hardware accelerators 210 are free to fetch training examples from any computing device with negligible risk of two hardware accelerators training on substantially similar batches of training examples; (8) each of the computing devices 200, after generating a training example, can enqueue the training example in a respective queue maintained by the computing device; (9) elements of a queue maintained by a computing device can be individual training examples, e.g., in the example in which the computational graph is for training a neural network, or elements of the queue can be batches of training examples; (10) at any point, a queue includes zero or more training examples awaiting processing by a hardware accelerator; (11) the scheduler engine can assign each of the hardware accelerators the operation data of the processing pipeline; (12) the hardware accelerators 210 can fetch training examples from one or more computing devices; (13) how many and which computing devices a particular hardware accelerator fetches data from can be set automatically by the scheduler engine, based on load-balancing and the individual capacity of the hardware accelerator; (14) in general, the scheduler engine can calculate a ratio for each hardware accelerator, representing the number of computing devices that, when assigned to the hardware accelerator, can provide preprocessing inputs to the hardware accelerator so as to mitigate or eliminate starvation of the hardware accelerator; (15) consider hardware accelerator A 210A to be a first-generation tensor processing unit, and hardware accelerators B 210B and C 210C to be second generation tensor processing units, wherein "generation" refers to incremental changes in performance for a same type of hardware accelerator, e.g., a TPU; (16) a later generation hardware accelerator generally performs better than an earlier generation hardware accelerator, e.g., because of improvements in computational capacity; (15) based on a respective type of each hardware accelerator, the scheduler engine can assign a corresponding number of computing devices to match the computational capacity of each hardware accelerator; e.g., the hardware accelerator A is assigned one computing device, computing device B 200B, and however, hardware accelerator B and C, both having a higher computational capacity than the hardware accelerator A are assigned two computing devices each: computing device A 200A and computing device D 200D to hardware accelerator B, and computing device C 200C and computing device E 200E to hardware accelerator C; (16) the ratio for a hardware accelerator can be determined by first determining how many training examples the hardware accelerator can process in a given period of time, e.g., 1 second; e.g., if a hardware accelerator can process a batch of 1024 examples in 110 milliseconds, then the approximate number of training examples that the hardware accelerator can process in a second (1000 milliseconds) is 1024 * (1000/110), or approximately 9,309 training examples a second; (17) from the computed rate, the scheduler engine can assign available computing devices to the hardware accelerator until at least approximately 9,309 training examples can be provided by the assigned devices and to the hardware accelerator each second while the accelerator is processing; (18) the training example processing rate can be pre-computed for each type of hardware accelerator implemented, and the scheduler engine can automatically assign an appropriate ratio of computing devices to a particular hardware accelerator based on the pre-computed rate; (19) by scaling the number of assigned computing devices to each hardware accelerator based on its computational capacity, the distributed computing system can improve computational throughput of the hardware accelerator and mitigate or prevent under-utilization (or "starvation") because the hardware accelerator can process training examples faster than what can be provided; (20) in some implementations two hardware accelerators can share the same computing device, and similar to how the distributed computing system can assign computing devices to each hardware accelerator based on the computational capacity of the hardware accelerator, the distributed computing system can also assign computing devices to each hardware accelerator based on the computational capacity of each computing device; e.g., if a computing device has a computational capacity meeting a predetermined computational threshold, then the distributing computing system can assign multiple hardware accelerators to the computing device; (21) the distributed computing system can determine the computational threshold based on the computational capacity of each hardware accelerator; (22) in addition or alternatively, each hardware accelerator of the distributed computing system can be configured to fetch enqueued data elements from each available computing device; e.g., a hardware accelerator can be configured to fetch enqueued data elements randomly from all available computing devices; (23) alternatively, the hardware accelerator can be configured to fetch enqueued data elements from each available computing device sequentially, e.g., based on identifiers for the computing devices, e.g., a network address; (24) by fetching data uniformly from each computing device, the hardware accelerators can reduce network congestion caused multiple hardware accelerators attempt to fetch data from the same computing device; (25) the distributed computing system can shift the computational demand away from the one or more host CPUs of a hardware accelerator, to free the host CPUs to perform other tasks; (26) the distributed computing system can also overcome memory bottlenecks by reducing the number of external memory reads needed by the host CPU for a hardware accelerator to perform the preprocessing operations on a set of data; (27) the host CPU for a hardware accelerator can batch preprocessed input from multiple computing devices assigned to its corresponding hardware accelerator so that the hardware accelerator is free to fetch preprocessed input data from a plurality of different computing devices and alleviate network traffic to any one particular computing device; (28) the distributed computing system that includes the computing devices 200 and the hardware accelerators 210 can automatically assign new devices connected to the network interconnecting the hardware accelerators 210 and the computing devices 200; e.g., as shown in FIG. 2B, the computing device 200F is shown to be providing input training examples to the hardware accelerator A 210A, denoted by the dashed line; (29) the distributed computing system can automatically handle reassignment, if necessary, of computing devices to the hardware accelerators in situations in which one or more computing devices become unavailable for preprocessing; (30) regardless of the circumstances leading to a computing device's unavailability, the distributed computing system can detect when a computing device becomes unavailable, and re-adjust computing-device-to-hardware accelerator assignment automatically, in response; (31) the distributed computing system can be configured to perform this re-assignment according to a variety of different approaches that can, e.g., favor overall utilization of the hardware accelerators overall, or prioritize utilization of some hardware accelerators that are better-optimized for performing operations in the computational graph of a processing pipeline; (32) in the example illustrated in FIG. 2C, the computing device A 200A and the computing device D 200D are unavailable, indicated by crossed-out arrows; (33) the distributed computing system can determine that computing devices previously assigned to the hardware accelerator B 210B are unavailable, and in response re-assign the computing device F 200F from the hardware accelerator A 210A to the hardware accelerator B 210B; (34) the distributed computing system can make the decision of which of the computing devices 200 to assign to the hardware accelerator B 210B based on the relative overall utilization of the hardware accelerators 210; and (35) if the hardware accelerator A 210A has a relatively smaller computational capacity than the hardware accelerator C 210C, the distributed computing system can decide to reassign a computing device from the hardware accelerator A 210A, e.g., the computing device F 200F, to the hardware accelerator B 210B, instead of reassigning from a computing device assigned to the hardware accelerator C 210C. Dar et al. (US 2020/0134467 A1, pub. date: 04/30/2020) discloses in ¶¶ [0006]-[0016] that (1) receiving a plurality of NN training tasks, each training task including (i) a respective preprocessing phase that preprocesses data to be provided as input data to the NN, and (ii) a respective computation phase that trains the NN using the preprocessed data; (2) the plurality of NN training tasks is executed, including: (a) a commonality is identified between the input data required by computation phases of two or more of the training tasks, and (b) in response to identifying the commonality, one or more preprocessing phases are executed that produce the input data jointly for the two or more training tasks; (3) in response to identifying the commonality, assigning two or more computation tasks to a same group of one or more processors; (4) assigning the computation phases to the processors in accordance with a predefined assignment criterion, wherein (i) the assignment criterion aims to minimize a total execution time of the training tasks; (ii) the assignment criterion aims to minimize idle times during computation phases; (iii) estimating durations of execution of the computation phases, and assigning the computation phases to the processors based on the durations of execution; and (iv) re-estimating the durations during execution of the computation phases, and reassigning one or more of the computation phases to the processors based on the re-estimated durations of execution; (5) assigning the preprocessing phases to the processors in accordance with a predefined assignment criterion, wherein (i) the assignment criterion aims to minimize a total execution time of the training tasks; and (ii) estimating durations of execution of the preprocessing phases, and assigning the preprocessing phases to the processors based on the durations of execution; and (6) deciding on a maximal number of training tasks for which to produce the input data jointly, based on a total execution time of the training tasks. Dar further discloses in ¶¶ [0022]-[0042] with FIG. 1 that (1) each step in the training process of NN includes two phases: preprocessing and NN computation; (2) the data preprocessing stage is typically performed by a Central Processing Unit (CPU), but can also be done on any other processing device; (3) non-trainable NN layers that are shared between multiple tasks can also be regarded as part of preprocessing; (4) the NN training may be performed using dedicated hardware ( e.g., GPU, TPU) that runs a software, such as "tensorflow," "pytorch," and others; (5) preprocessing phases typically run in parallel to computation phases and may create a bottleneck in the pipeline; (6) there is a pipeline of preprocessing and computation tasks, and preprocessing of a next step runs in parallel to the computation of the current step; (6) some example scenarios for bottlenecks caused by preprocessing are:(i) the data needs to be read from disk or from remote servers; (ii) the preprocessing is performed on low-end CPUs; (iii) the preprocessing is performed on a distributed computing system, which creates latency and increases the preprocessing time; (iv) when the duration of a preprocessing time is longer than the duration of the respective NN computation time, the high-end GPUs (or other platforms that run the NN computation) may become idle; and (v).as a result, the total time and cost of training the NN increases; (7) utilize similarity between different training tasks to optimize a total running time and cost of the NN training process; (8) allocates groups of training tasks to run together on a same hardware while sharing parts of the preprocessing and/or NN computation; and (9) the system can be optimized to achieve a shortest training time or to achieve a training task at a lower training cost. Dar further discloses in ¶¶ [0071]-[0078] with FIG. 4 that (1) a work-flow for training neural networks (NN) including allocation of resources of a distributed computing system per a given number of training tasks; (2) a processor of the system ( e.g., a CPU in preprocessing unit 30) splits the tasks into task-groups the present example two groups denoted 520, 522), wherein each group of tasks may have some shared preprocessing portion to be determined; (3) in analysis stages (530, 532), the processor splits tasks within each group (520, 522) among preprocessing groups 540, 542, 544, and 546 so as to optimize preprocessing running time; (4) the estimation of durations of execution of the computation phases is done dynamically as training progresses; (5) the estimation process then comprises re-estimating the durations during computation, and reassigning the computation phases to the processors based on the re-estimated durations of execution; e.g., when preprocessed data is shared between tasks, the preprocessing runtime may change during computation phase of training since runtime may depend on network load and configuration; (6) dynamically receive information reported by the NN, such as typically reported parameters at every beginning or end of a step or an epoch, on the runtimes of the shared and non-shared parts and modify "Execution groups"; (7) the algorithm can then stop training runs, optimize their resource allocation and grouping, and rerun them accordingly; (8) the new runs can continue from previous state by using a suitable checkpoint and recovery mechanism; (9) next, preprocessing groups 540, 542, 544, and 546 serve as input for computation groups 550, 551, 552, and 554; and (10) the partitioning into groups 550, 551, 552, and 554 aims to optimize the overall cost of training; e.g., (i) computation group 550 may receive two training tasks, each having a shared preprocessed portion (in the present example preprocessing group 540) and run on eight Nvidia k80 GPUs, allocating six GPUs to the first training task, two GPUs to the second task, and so on; and (ii) computation group 551 may receive a single training task (in the present example preprocessing group 542) and run on a single Nvidia v100 GPU only; (iii) computation group 552 may receive a single training task (in the present example preprocessing group 544) and run on a single TPU only; and (iv) computation group 554 may receive two training tasks, each having a shared preprocessed portion (in the present example preprocessing group 546) and run on two Nvidia k80 GPUs. Ibrahim et al. ("Preprocessing Pipeline Optimization for Scientific Deep Learning Workloads", 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS), May 30 – Jun 3, 2022, pp. 1118-1128) discloses in ABSTRACT of Page 1118 that (1) scientific machine learning performance is often constrained by data movement overheads, particularly on existing and emerging hardware-accelerated systems; (2) focus on optimizing the data movement across storage and memory systems, by developing domain-specific data encoder/decoders; (3) these plugins have the dual benefit of significantly reducing communication while enabling efficient decoding on the accelerated hardware; (4) explore detailed performance analysis for two important scientific learning workloads from cosmology and climate analytics, CosmoFlow and DeepCAM, on the GPU-enabled Summit and Cori supercomputers; (5) results demonstrate that our optimizations can significantly improve overall performance by up to 10× compared with the default baseline, while preserving convergence behavior; and (6) overall, this methodology can be applied to various machine learning domains and emerging AI technologies. Ibrahim further discloses in Section I of Page 1118 that (1) the development of domain-specific encoding and decoding plugins, which have the advantages of reducing the required data movement, thus allowing a larger set to fit in the memory system, and accelerating data prepossessing execution on the learning engine; (2) our encoders also apply fusion and reordering of the decompression steps with application-specific preprocessing computations to reduce computation and data movement; (3) additionally, our approach enables a data decoding scheme implemented on the hardware accelerators, further reducing the overall runtime; (4) analyzing the contents of samples data for two scientific deep learning workloads, CosmoFlow and DeepCAM, to devise a compressed format suitable for execution on GPU-accelerated architectures; (5) implementing specialized preprocessing plugin to improve the speed of feeding data to the deep learning pipeline for the two applications; (6) demonstrating the preservation (or improvement) of convergence properties; and (7) evaluating our implemented schemes on three HPC platforms, using a variety of configurations, and demonstrating speedups of up to 3× and 10× for DeepCAM and CosmoFlow, respectively, compared with the baseline. Ibrahim also discloses in Section II with FIG. 1 of Pages 1118-1119 that (1) in an HPC environment shown in Figure 1.a, a training sample could originate from a shared file system traversing multiple hops until reaching the accelerator optimized for conducting a training or inference task; (2) for training, a sample is repeatedly accessed, proportional to the number of training epochs; (3) the ability to cache the training set depends on the number of samples assigned to a node and the capacity of the storage or memory hierarchy; e.g., (i) if the samples assigned to a node fit in the host CPU memory, a sample traverses step 1 & 2 once, while step 3 & 4 are repeated throughout the training session; and (ii) if the dataset per node fits in the node NVMe, but not in memory, the step 2 & 3 & 4 are repeated; (4) a challenge associated with this migration path is the potentially large arithmetic intensity required to hide the memory latency; (5) the potential of caching data in the memory systems depends on the sample size and the number of samples assigned to a node; (6) in general, reducing the input sample size, for instance through compression, enables caching more samples in the host CPU memory; (7) the preprocessing logic, step 3 , may involve decompression of a sample or applying some augmentation operators; e.g., (i) in CosmoFlow, the preprocessing involves applying a log operator to all points within a sample; and (ii) in image classification applications, preprocessing of an image may involve rotating, resizing, flipping, etc.; (8) some of these preprocessing steps carry serialization logic and as such can be executed more efficiently on the CPU; (9) however, when possible, efficiency can be increased by offloading preprocessing computations to the accelerators; (10) Figure 1.b shows the optimized migration path developed in this work; (11) the samples are initially encoded in step b.1 to enable efficient migration and processing on the target accelerator architecture; (12) this enables reducing the data movement, allowing the dataset caching on the nearest level of the memory system, and enabling efficient decoding (or decompressing samples) on the accelerator; (13) additionally, we perform complex preprocessing operators on the GPU, moving step a.3 to step b.5 for accelerated preprocessing; (14) comparing Figure 1.a & 1.b, a reduced sample size potentially lower the impact of the bottleneck transfer at step a. 2 and increase the frequency of performing step b. 4, leveraging bandwidth that is one to two orders of magnitude larger; (15) moreover, offloading the sample decoding and preprocessing to GPUs adds the additional benefit of fewer data moved across the bandwidth-constrained system resources (PCIe, NVLink, etc.); and (16) it also leverages the accelerator computational power in preprocessing samples. However, closest arts of records, as discussed above, singly or in combination do not teach or suggest at least following features "determine/determining, as the number of the at least one input value to be transmitted to the external device, a minimum number among a number of input values that can be processed at an operation speed equal to a difference between the first operation speed and the second operation speed for a predetermined time period, a number of input values that can be processed at the third operation speed during the predetermined time period, and a number of input values that can be transmitted through a bandwidth of the network during the predetermined time period". Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HWEI-MIN LU/Primary Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

Dec 28, 2023
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §102, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705533
SYSTEMS AND METHODS FOR IMPROVING PREDICTION PROCESS USING AUTOMATED RULE LEARNING FRAMEWORK
3y 8m to grant Granted Aug 11, 2026
Patent 12700003
SYSTEMS AND METHODS FOR FREQUENT MACHINE LEARNING MODEL RETRAINING AND RULE OPTIMIZATION
4y 2m to grant Granted Aug 04, 2026
Patent 12694335
SYSTEMS AND METHODS FOR REPURPOSING A MACHINE LEARNING MODEL
3y 6m to grant Granted Jul 28, 2026
Patent 12682260
ADJUDICATION ALGORITHM BYPASS CONDITIONS
4y 1m to grant Granted Jul 14, 2026
Patent 12675690
ANOMALY DETECTION WITH MODEL HYPERPARAMETER SELECTION
4y 3m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
63%
Grant Probability
99%
With Interview (+39.6%)
2y 11m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 233 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month