Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. KR 10-2023-0181900, filed on December 14th, 2023.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 10, 11, and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over “Jaehoon Jung, Jinpyo Kim, and Jaejin Lee. 2023. DeepUM: Tensor Migration and Prefetching in Unified Memory. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS ’23), March 25–29, 2023, Vancouver, BC, Canada. ACM, New York, NY, USA, 15 pages” (hereafter Jung),
in view of “Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, Xuehai Qian. 2020. Capuchin: Tensor-based GPU Memory Management for Deep Learning. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’20), March 16–20, 2020, Lausanne, Switzerland. ACM, New York, NY, USA, 15 pages” (hereafter Peng),
and Minkin et al. Pub. No. US 2023/0289292 A1 (hereafter Minkin).
Regarding claim 1, Jung teaches “A processor-implemented method comprising: … perform an initial iteration of training of a deep learning application, wherein the initial iteration is performed through execution of a plurality of kernels and, …; storing pattern information for allocating the … kernels to the memory in the initial iteration; and prefetching … based on the pattern information to perform a next iteration of training of the deep learning application ([Pg. 208, Introduction] teaches that kernel execution patterns and their memory access patterns are mostly fixed and can be repeated in the DNN training workload, thus memorizing the repeated patterns and exploiting the information through correlation prefetching is desirable such that an initial iteration of training must occur in order to record the pattern information. It states that DeepUM’s correlation tables record the history of the kernel executions and their page accesses during the training phase of a DNN and prefetches information from the correlation tables by predicting which kernels will execute next such that it is to perform a next iteration)”.
Jung does teach of kernel execution patterns and the memory access patterns within the kernels, however, Jung may not explicitly teach of tensor access patterns and tensor prefetching.
Peng teaches tensor access patterns and prefetching tensors such that it teaches the limitation “A processor-implemented method comprising: allocating tensors to a memory to perform an initial iteration of training of a deep learning application ([Pg. 891, Abstract] teaches a tensor-based GPU memory management module that reduces memory footprint via tensor eviction/prefetching. It makes decisions based on dynamic tensor access patterns tracked at runtime such that tensors were allocated during a training iteration. Also see [Pg. 895, section 4.2] regarding dividing the training into two phases, wherein the first iteration is for observing the tensor access sequence) … storing pattern information for allocating the tensors … and prefetching the tensors based on the pattern information to perform a next iteration of training of the deep learning application ([Pg. 895, section 4.2] teaches that there are two phases of training; measured execution and guided execution, wherein measured execution observes for tensor access and records the tensor access sequence with additional information for each tensor such as access count, timestamp, operation that produced a tensor and its input tensor for recomputation decisions, and such)”.
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Peng’s Capuchin - tensor-based access pattern tracking and prefetching with Jung’s DeepUM correlation-based prefetching framework as both references address the same problem of reducing memory-management overhead during iterative DNN training by exploiting recurring memory-access behavior. In particular, Peng’s Capuchin teaches that tensor accesses exhibit regular and repeated patterns across training iterations and dynamically tracks the tensor accesses to make management decision, such as tensor prefetching. DeepUM similarly recognizes that DNN training exhibits repeated kernel execution and memory-access patterns, records the history of kernel executions with their corresponding memory accesses and predicts the next kernel to execute, prefetching data expected to be accessed by that kernel. Thus, it would have been obvious to have stored both tensor and kernel pattern information for prefetching in future iterations of training. A person having ordinary skill in the art would have been motivated to make this combination in order to reduce memory footprint, allowing for fine-grain and flexible control of when and how to perform memory optimization techniques within deep learning (Peng, Pg. 891, Abstract). Since the teachings were analogous art known at the filing time of the invention, one of ordinary skill could have applied said teachings to achieve expected results.
The combination may not explicitly teach tensors corresponding to each of the kernels used to execute the kernels.
Minkin teaches of a tensor memory access unit (TMAU) that provides hardware circuitry for moving data blocks between global and shared memory, wherein a kernel may access the data that may represent tensors such that it teaches the limitation “in response to each kernel being executed, tensors corresponding to the each kernel are used to execute the kernels; ([0040] teaches that a kernel may access multidimensional data structures which may include matrices that may represent tensors. [0096-0097] also teaches that the kernel may obtain pointers to tensor descriptors, which is then provided to the TMAU for subsequent memory access requests)”.
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have applied the teachings of Minkin to the combination of Jung and Peng, to implement a tensor memory access unit (TMAU) that can provide a relationship between a kernel’s execution and the particular tensor blocks that the kernel accesses. A person having ordinary skill in the art would have been motivated to make this combination in order to associate the memory access behavior of a kernel taught by Jung’s DeepUM with the tensor access information tracked by Peng’s Capuchin, thereby using kernel-specific tensor access behavior to inform prefetching decisions, as well as providing efficient data transfer mechanisms, offloading a significant portion of data access operations from kernels running on streaming multiprocessors (Minkin [0042]). Since the teachings were analogous art known at the filing time of the invention, one of ordinary skill could have applied said teachings to achieve expected results. Together, Jung in combination with Peng and Minkin teach every limitation of the claimed invention. Since the teachings were analogous art known at the filing time of the invention, one of ordinary skill could have applied said teachings to achieve expected results.
Regarding claim 10, the combination teaches “A processor-implemented method comprising implementing the trained deep learning application, wherein the deep learning application is trained by the method of claim 1 ([Peng Pg. 892, right column] teaches Capuchin, a tensor-based GPU memory management module for deep learning frameworks. Or [Jung pg., 208, right column] teaches DeepUM that exploits CUDA UM to allow GPU memory oversubscription for DNNs. Or [Minkin 0164-0165] teaches parallel processing (PPU) configured to implement large neural networks in deep learning applications)”.
Regarding claim 11, the combination teaches “A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1 (Minkin [0219-0220] may not explicitly mention non-transitory, but list examples that are non-transitory)”.
Claim 13 is similar to claim 1 and is rejected for the same reasons. Claim 13 is directed towards “An electronic device comprising: one or more processors (Minkin [Abs])”.
Regarding claim 2, the combination teaches “The method of claim 1, wherein the storing of the pattern information for allocating the tensors and the kernels to the memory comprises: generating unique information of the tensors; and generating and storing the pattern information based on the unique information (Peng [Pg. 895, section 4.2] teaches that during measured execution, Capuchin keeps the tensor access sequence along with additional information for each tensor)”.
Claim 14 is similar to claim 2 and is rejected for the same reasons.
Regarding claim 3, the combination teaches “The method of claim 2, wherein the unique information comprises feature information of a tensor and stack information about a process of allocating the tensor to the memory (Peng [Pg. 895, section 4.2] teaches that during measured execution, Capuchin keeps the tensor access sequence along with additional information for each tensor, including access count, timestamp, and operation that produced a tensor)”.
Claim 15 is similar to claim 3 and is rejected for the same reason.
Regarding claim 12, it is similar to claims 1, 2, and 3 combined, and is rejected for similar reasons.
Regarding claim 4, the combination teaches “The method of claim 1, wherein the pattern information comprises a kernel table for storing an execution order of the kernels in the initial iteration and a tensor table for storing tensors corresponding to the each kernel (Jung [Pg. 210, section 3.1] teaches correlation tables managed by DeepUM record the history of kernel execution and their page accesses during training phase. Peng [Pg. 895, section 4.2] teaches that Capuchin keeps tensor access sequence information with additional information including operation that produced a tensor and its input tensor such that the operation that produced the tensor may be a kernel that was executed. Minkin [0097] teaches that the kernel may obtain pointers to tensor descriptors such that the kernel is associated with tensors. Thus when combined, Jung’s correlation table of kernel executions to memory accesses may correspond to the tensor access information as it contains information pertaining to how the tensor was produced, as well as the hardware circuitry of Minkin, the TMAU that is used to map the kernel to the tensor via pointers during memory access requests)”.
Claim 16 is similar to claim 4 and is rejected for the same reason.
Regarding claim 5, the combination teaches “The method of claim 4, wherein the prefetching of the tensors comprises predicting kernels to be executed based on the kernel table and the tensor table and prefetching tensors corresponding to the predicted kernels ([Jung Pg. 212, section 4.2] teaches of correlation prefetching, wherein it has an execution ID correlation table that records the history of execution IDs of the kernels. There also exists a block table for each execution ID and records a history of UM block accesses within the corresponding kernel. When a predicted kernel is to be executed based on the table, prefetching occurs by prefetching UM blocks corresponding to their respective kernels within the correlation table. Peng has previously taught of tensor access patterns recorded and prefetched for future iterations, and Minkin has taught that memory access requests of kernels may be for tensors such that when combined, it would be plausible that the block table of Jung contains tensor information for prefetching the specific blocks corresponding to their kernels. Also see [Peng section 5.2] regarding a (tensor_id, access_count) pairing wherein it may represent a block access within memory as it would require a tensor to access that specific memory block)”.
Claim 17 is similar to claim 5 and is rejected for the same reason.
Regarding claim 7, the combination teaches “The method of claim 2, further comprising managing the tensors with structures including the unique information ([Peng Pg. 898, section 5.2 and figure 5] teach of a tensor structure. Or [Jung, Pg. 212-213, section 4.2] teaches of a block table)”.
Claim 19 is similar to claim 7 and is rejected for the same reason.
Regarding claim 9, the combination teaches “The method of claim 1, further comprising training the deep learning application using the prefetched tensors ([Jung Pg. 213, section 4.2] teaches of prefetching UM block accesses corresponding to kernels and [Peng Pg. 897, section 4.3-4.5] teaches that tensor access information is used to determine which tensors should be available in advance during the guided execution of training such that prefetch tensors are used for training the deep learning application)”.
Claims 6 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Jung, Peng, and Minkin as used above in claims 1 and 5, and in further view of Eidus et al. Pub. No. US 2010/0223432 A1 (hereafter Eidus).
Regarding claim 6, the combination teaches of a tensor table (Jung and Peng combined), however it may not explicitly teach of using a self-balancing binary tree to generate the table.
Eidus teaches a self-balancing binary search tree for managing pages such that it teaches the limitation “The method of claim 5, wherein the tensor table is generated through a search using a self-balancing binary search tree ([0024-0030] teaches of a memory manager that maintains data structures for sorting pages, wherein a data structure may be a red-black tree that can be used to improve search performance. When combined with the block table of Jung’s DeepUM [Section 4.2], it may be used to search up addresses within the unified memory block to find start and end addresses that point to page faults and prefetched blocks accordingly, such that the block table may be made up of the page faults that have been recorded and searched by the binary search tree)”.
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have applied the teachings of Eidus to the combination of Jung, Peng and Minkin, to implement self-balancing binary search tree as the data structure for storing pages. A person having ordinary skill in the art would have been motivated to make this combination in order to improve search performance regarding pages (Eidus [0025]), especially when a page fault occurs and it needs to be recorded and prefetched.
Claim 18 is similar to claim 6 and is rejected for the same reason.
Claims 8 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jung, Peng, Minkin as used above in claims 1 and 7, and in further view of Eidus et al. Pub. No. US 2010/0223432 A1 (hereafter Eidus) and PATEL et al. Pub. No. US 20240220571 A1 (hereafter Patel).
Regarding claim 8, the combination may not explicitly teach of using a self-balancing search tree for searching for a page fault.
Eidus teaches using a red-black tree for managing pages such that it teaches the limitation “The method of claim 7, wherein the managing of the tensors comprises, in response to a page fault occurring, managing the structures with a self-balancing binary search tree for searching for a structure in which the page fault occurred ([0024-0025] teaches a red-black tree used for searching pages)“.
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have applied the teachings of Eidus to the combination of Jung, Peng and Minkin, to implement self-balancing binary search tree as the data structure for storing pages. A person having ordinary skill in the art would have been motivated to make this combination in order to improve search performance regarding pages (Eidus [0025]), especially in a deep learning application when page faults occur and pages correlating to the faulted block needs to be prefetched (Jung [Section 4.2]).
The combination may not explicitly teach of managing structures with a hash table.
Patel teaches of an input tensor with elements of a kernel to determine a generated output corresponding to each kernel point using a hash table such that it teaches the limitation “and managing the structures with a hash table to search for a tensor using the feature information ([0041-0044] teaches of determining affected output elements for each valid input element, based on an input tensor and kernel. Respective tables are then generated for each valid input element, where the tables may be implemented as hash tables the affected output elements are then subsequently used to generate intermediate values and a convolution output associated with each kernel point, such that these intermediate values may form the convolution output tensor)”.
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have applied the teachings of Patel to the combination of Jung, Peng, Minkin, and Eidus to implement a hash table based organization of tensor related information to the tensor structures. A person having ordinary skill in the art would have been motivated to make this combination in order to provide efficient lookup of tensors and tensor related information during kernel/convolution processing, reducing redundancy/latency associated with locating tensors (Patel [0016 & 0057]).
Claim 20 is similar to claim 8 and is rejected for the same reasons.
References not cited but are pertinent to the art are as follows:
WO 2023121831 A1
Teaches
CONFIGURING A PREFETCHER ASSOCIATED WITH A PROCESSOR CORE
US 6549996 B1
teaches
Scalable Multiple Address Space Server
US 20220350744 A1
teaches
TECHNIQUES FOR PRE-FETCHING INFORMATION USING PATTERN DETECTION
US 20210157734 A1
teaches
METHOD AND APPARATUS FOR CONTROLLING MEMORY USING PREFETCH INFORMATION
US 20200302304 A1
teaches
METHODS OF OPERATING A GRAPHICS PROCESSING UNIT (GPU) TO TRAIN A DEEP NEURAL NETWORK USING A GPU LOCAL MEMORY AND RELATED ARTICLES OF MANUFACTURE
US 20220284263 A1
teaches
NEURAL NETWORK OPERATION APPARATUS AND METHOD
US 20200202198 A1
teaches
NEURAL NETWORK PROCESSOR
US 20230274129 A1
teaches
METHOD FOR EXECUTION OF COMPUTATIONAL GRAPH IN NEURAL NETWORK MODEL AND APPARATUS THEREOF
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRANDON A NGUYEN whose telephone number is (571)272-6074. The examiner can normally be reached Mon-Fri (10am-6pm).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRANDON NGUYEN/Examiner, Art Unit 2195
/Aimee Li/Supervisory Patent Examiner, Art Unit 2195