DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/11/2026 has been entered.
Claims 1-20 are pending and they are presented for examination.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 (similarly claim(s) 8 and 14) recite: “return an indication of which of one or more hardware non-uniform memory access (NUMA) storages and one or more hardware graphics processing unit (GPU) storages”. The limitation is ambiguous, the examiner is unclear if the word “which” is meant to pick particular one NUMA storage from one or more NUMA storages, or pick particular one NUMA storage and particular one GPU storages, etc.
Claim 9 (similarly claims 10 and 11) recite: “API call, are to indicate”. The examiner is unclear if API call is a singular term or a plural term. Since, API call is a single call, the examiner is unclear why “are” is use.
Claims 2-7, 9-13 and 15-20 are rejected based on rejection of its corresponding dependent claim.
Response to Amendment
Applicant's arguments with respect to claims have been considered but are moot in view of the new ground(s) of rejection.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayanan et al. (Pub 20230004417) (hereafter Narayanan) in view of Wagle et al. (Pub 20160371194) (hereafter Wagle) and further in view of Dragon: Breaking GPU Memory Capacity Limits with Direct NVM Access (IEEE 2018 Pak Markthub) (hereafter Pak).
As per claim 1, Narayanan teaches:
A processor, comprising:
circuitry to, in response to receiving an application programming interface (API) call, return an indication of which of one or more hardware non-uniform memory access (NUMA) storages and one or more hardware graphics processing unit (GPU) storages correspond to one or more virtual memory addresses indicated by the API call. ([Paragraph 1], Virtualization allows system software called a virtual machine monitor (VMM), also known as a hypervisor, to create multiple isolated execution environments called virtual machines (VMs) in which operating systems (OSs) and applications can run. [Paragraph 65], user priority inputs (high/medium/low) are read from the VM configuration file, the selection policy is intelligent and the NUMA preferred node of the running VM and input device class for which an ADI is to be selected are determined. [Paragraph 81], The Graphics Processor Unit (GPU) module 810 may include one or more GPU cores and a GPU cache which may store graphics related data for the GPU core. The GPU core may internally include one or more execution units and one or more instruction and data caches. Additionally, the Graphics Processor Unit (GPU) module 810 may contain other graphics logic units that are not shown in FIG. 8, such as one or more vertex processing units, rasterization units, media processing units, and codecs. [Paragraph 52], In addition, in order to preserve priorities that may be determined at the application level by users, an optional user configuration file can be used to specify a priority High, Medium, Low) for a VM and an option to incorporate priorities in the selection policy, or override the selection policy with priorities. If there is a conflict in resource allocation, the priority assigned to each VM is used to determine the best ADI to assign. For example, if two VMs request a VDEV composition, based upon the VM's priority (High/Med/Low), the ADI is assigned to the VM with the higher priority. [Paragraph 24], System software can use an API to request intelligent selection of an ADI for a device class by traversing the ADI selection tree to select the ADI for the device class. The traverse of the ADI selection tree is performed in response to the request from an application through an application programming interface (API) call.)
Narayanan teaches issuing an API call and having a user preferred nodes which includes NUMA storages and GPUs. [Narayanan paragraph 43, 65, 81].
However, Narayanan does not explicitly disclose return an indication of which of one or more non-uniform memory access (NUMA) storages and one or more graphics processing unit (GPU) storages correspond to one or more storages indicated by the API call.
Wagle teaches return an indication of which of one or more non-uniform memory access (NUMA) storages and one or more graphics processing unit (GPU) storages correspond to one or more storages indicated by the API call. ([Paragraph 27], The NUMA API currently supports four different policy flavors: default: Allocate on the local node (the node the thread is running on)… bind: Allocate on a specific set of nodes… preferred: Try to allocate on a node first… [Paragraph 32], FIG. 3, NUMA policy is set for virtual memory addresses using APIs from the ‘libnuma’ library. FIG. 4 illustrates use of a new API “ALLOC_ON_NUMA_NODE (<N>)” which wraps these libnuma APIs in order to set NUMA policy for virtual memory addresses. [Paragraph 39], First, at S535, the memory manager sets the NUMA policy to ‘preferred’ by calling a libnuma API (e.g., numa_set_bind_policy(0)). Next, at S540, the allocation bit mask is set to the first node (e.g., numa_bitmask_setbit(mask, 7)). The reserved memory is bound to the first node at S545 using, e.g., the numa_tonodemask_memory(ptr, <size>, mask) API call. S550 then includes setting a second NUMA policy to act as a fallback policy in a low memory or out-of-memory condition. In the present example, the second NUMA policy is the interleaved policy but embodiments are not limited thereto. [Paragraph 41], The portion may have a size corresponding to the Small, Medium or Big sub-allocators. The address pointer of the identified portion is then returned to the worker thread at S565 (e.g., RESPONSE of FIG. 6).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the invention, to combine the teachings of Narayanan wherein an API(s) is/are utilized to indicate an API corresponds to a NUMA node (with processor and storage), priority level for NUMA preference is determined at an application level by users and prefetch buffers exist to store prefetched data, into teachings of Wagle wherein an indication of preferred storages are returned via address pointers, because this would enhance the teachings of Narayanan wherein by returning an indication of storages based on preference and fallback policy allows flexibility of requesting and subsequently allocating/assigning a storage in an order of preference/priority based on availability. [Wagle paragraph 39-40]
Although Narayanan and Wagle discloses of hardware NUMA storages and GPU storages/memory(ies) such as NVM (non-volatile memory devices). ([Narayanan paragraph 43, 81 ] [Wagle paragraph 4,]
However, Narayanan and Wagle do not explicitly disclose one or more virtual memory addresses indicated by the API call.
Pak teaches one or more virtual memory addresses indicated by the API call. ([Page 3], Changes required for CUDA applications to work with the DRAGON framework are minimal. Applications can take advantage of DRAGON by using the dragon_map() API function with a file path and other optimization parameters (explained later) to map previously dumped binary data into a unified memory space. This function internally finds an unmapped virtual address space, registers this space along with the file path to the internal tracking mechanism of our driver extension, and returns the virtual address to the application. Once mapped, the memory range is accessible by both GPUs and CPUs, and DRAGON transparently allows direct access down to NVM devices during GPU kernel execution… In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space.)
Since Narayanan discloses NVM devices being utilized as NUMA storages.
Pak also teaches one or more hardware non-uniform memory access (NUMA) storages and one or more hardware graphics processing unit (GPU) storages correspond to one or more virtual memory addresses indicated by the API call. ([Page 2], Going beyond system memory capacity limits: Active-Pointers [28] is the only study that considered expanding GPU addressable memory range by mapping it to a file-system. Active Pointers is based on GPUs [29] and relies on in-kernel software address translation (i.e., SW-based page faulting). It requires existing CUDA kernels to be modified to use custom pointers and APIs. Memory references are captured on the fly via operator overloading, and a page-fault handling kernel-code is executed for every access. This approach incurs significant overhead in addition to the extensive cost of modifying existing GPU kernels; therefore, it is relatively inefficient as it does not exploit the new hardware paging support in GPUs, as demonstrated in § IV.
[Page 3], In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space. NVIDIA driver for Pascal and Volta architectures captures page faults from both GPUs and CPUs, and DRAGON relies on this hardware-based page faulting mechanism to handle the accesses from GPU kernels to the mapped files. The memory consistency between GM and HM is handled by the original UM-P module, whereas data swapping and consistency between HM and NVM are handled by our driver extension (as explained in detail in § III-B).)
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the invention, to combine the teachings of Narayanan and Wagle wherein an API call(s) is/are utilized to indicate an API corresponds to a NUMA node (with processor and storage), priority level for NUMA preference is determined at an application level by users, prefetch buffers exist to store prefetched data and return and indication of assigned/allocated preferred storage (via address), into teachings of Pak wherein one or more virtual memory addresses indicated by the API call correspond to NUMA and/or GPU storages, because this would enhance the teachings of Narayanan wherein by implementing a unified virtual memory to create a memory range that is accessible by both GPU(s) and CPU(s), it allows transparent direct access to NVM devices during GPU kernel execution and allows utilization of heterogeneous memory hierarchy to take advantage of large capacity and high bandwidth of NVM devices. [Pak page 2-4]
As per claim 2, rejection of claim 1 is incorporated:
Narayanan teaches prefetch buffers ([Paragraph 77], Although not shown, the processor core 802 may internally include one or more instruction/data caches (L1 cache), execution units, prefetch buffers, instruction queues, branch address calculation units, instruction decoders, floating point units, retirement units, etc.). Pak also discloses prefetching [Page 5]
Pak teaches wherein the circuitry, in response to receiving the API call, is to indicate a NUMA node to which data was last prefetched in response to another API call. ([Page 5], DRAGON also allows programmers to implement more advanced prefetching behavior by using cudaMemAdvise() and cudaMemPrefetchAsync(), which are provided by the original UM-P [10]. The advice parameter will hint the UM-P driver (and hence DRAGON) about the data access behavior. When cudaMemPrefetchAsync() is called, DRAGON will replicate the intended effects of advice for the addresses residing on NVM. This behavior can be utilized to minimize NVM read overhead for scattered data, where Linux’s read-ahead will not help due to nonconsecutive access patterns. [Page 3], Changes required for CUDA applications to work with the DRAGON framework are minimal. Applications can take advantage of DRAGON by using the dragon_map() API function with a file path and other optimization parameters (explained later) to map previously dumped binary data into a unified memory space. This function internally finds an unmapped virtual address space, registers this space along with the file path to the internal tracking mechanism of our driver extension, and returns the virtual address to the application. Once mapped, the memory range is accessible by both GPUs and CPUs, and DRAGON transparently allows direct access)
As per claim 3, rejection of claim 1 is incorporated:
Narayanan teaches preferred location of one or more memory addresses by API(s) ([Paragraph 1], Virtualization allows system software called a virtual machine monitor (VMM), also known as a hypervisor, to create multiple isolated execution environments called virtual machines (VMs) in which operating systems (OSs) and applications can run. [Paragraph 65], user priority inputs (high/medium/low) are read from the VM configuration file, the selection policy is intelligent and the NUMA preferred node of the running VM and input device class for which an ADI is to be selected are determined. [Paragraph 81], The Graphics Processor Unit (GPU) module 810 may include one or more GPU cores and a GPU cache which may store graphics related data for the GPU core. The GPU core may internally include one or more execution units and one or more instruction and data caches. Additionally, the Graphics Processor Unit (GPU) module 810 may contain other graphics logic units that are not shown in FIG. 8, such as one or more vertex processing units, rasterization units, media processing units, and codecs. [Paragraph 52], In addition, in order to preserve priorities that may be determined at the application level by users, an optional user configuration file can be used to specify a priority High, Medium, Low) for a VM and an option to incorporate priorities in the selection policy, or override the selection policy with priorities. If there is a conflict in resource allocation, the priority assigned to each VM is used to determine the best ADI to assign. For example, if two VMs request a VDEV composition, based upon the VM's priority (High/Med/Low), the ADI is assigned to the VM with the higher priority.)
Pak teaches wherein the circuitry, in response to receiving the API call, is to indicate a NUMA node set as a preferred location of one or more memory addresses in response to another API call. ([Page 3], NVM-specific optimizations into modern GPUs to exploit the unique benefits of the underlying heterogeneous memory hierarchy. For the proof-of-concept implementation, DRAGON is developed as an extension to Unified Memory for Pascal [8], and it is made accessible to users via a separate user-level host API. Our extension and modifications to the NVIDIA driver are limited to the MIT-licensed open-source nvidia-uvm submodule. DRAGON preserves all existing CUDA features without any interface change or performance penalty. Changes required for CUDA applications to work with the DRAGON framework are minimal. Applications can take advantage of DRAGON by using the dragon_map() API function with a file path and other optimization parameters (explained later) to map previously dumped binary data into a unified memory space. This function internally finds an unmapped virtual address space, registers this space along with the file path to the internal tracking mechanism of our driver extension, and returns the virtual address to the application. Once mapped, the memory range is accessible by both GPUs and CPUs, and DRAGON transparently allows direct access down to NVM devices during GPU kernel execution. [Page 2], In such cases (e.g., in situ visualization), using node-local NVMs as the primary memory location will minimize the data movement and also allow full access to the processed data on the later steps of the workflow. [Page 5], dragon_map() accepts an additional parameter, named flags, to indicate the data access type. D_READ and D_WRITE correspond to the read-only and write-only optimizations, respectively (§ III-C). Combining them (the default value) tells DRAGON that the mapped data are for both reading and writing and thus do not apply those optimizations. The presence of D_VOLATILE tells DRAGON to apply the intermediate data optimization (§ III-D). The other two API functions, dragon_sync() and dragon_unmap(), allow applications to manually flush GM and HM contents and release all memory space occupied by the driver, respectively. DRAGON also allows programmers to implement more advanced prefetching behavior by using cudaMemAdvise() and cudaMemPrefetchAsync(), which are provided by the original UM-P [10]. The advice parameter will hint the UM-P driver (and hence DRAGON) about the data access behavior. When cudaMemPrefetchAsync() is called, DRAGON will replicate the intended effects of advice for the addresses residing on NVM. This behavior can be utilized to minimize NVM read overhead for scattered data, where Linux’s read-ahead will not help due to nonconsecutive access patterns. )
As per claim 4, rejection of claim 1 is incorporated:
Pak teaches wherein the circuitry, in response to receiving the API call is to indicate a location type and a location identity of virtual memory accessible by one or more central processing units (CPUs) and the one or more hardware GPUs. ([Page 2], Going beyond system memory capacity limits: Active-Pointers [28] is the only study that considered expanding GPU addressable memory range by mapping it to a file-system. Active Pointers is based on GPUs [29] and relies on in-kernel software address translation (i.e., SW-based page faulting). It requires existing CUDA kernels to be modified to use custom pointers and APIs. Memory references are captured on the fly via operator overloading, and a page-fault handling kernel-code is executed for every access. This approach incurs significant overhead in addition to the extensive cost of modifying existing GPU kernels; therefore, it is relatively inefficient as it does not exploit the new hardware paging support in GPUs, as demonstrated in § IV.
[Page 3], In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space. NVIDIA driver for Pascal and Volta architectures captures page faults from both GPUs and CPUs, and DRAGON relies on this hardware-based page faulting mechanism to handle the accesses from GPU kernels to the mapped files. The memory consistency between GM and HM is handled by the original UM-P module, whereas data swapping and consistency between HM and NVM are handled by our driver extension (as explained in detail in § III-B).)
As per claim 5, rejection of claim 1 is incorporated:
Pak teaches wherein the API call comprises one or more parameters that indicate a range of memory. ( [Page 5], dragon_map() accepts an additional parameter, named flags, to indicate the data access type. D_READ and D_WRITE correspond to the read-only and write-only optimizations, respectively (§ III-C). Combining them (the default value) tells DRAGON that the mapped data are for both reading and writing and thus do not apply those optimizations. The presence of D_VOLATILE tells DRAGON to apply the intermediate data optimization (§ III-D). The other two API functions, dragon_sync() and dragon_unmap(), allow applications to manually flush GM and HM contents and release all memory space occupied by the driver, respectively. DRAGON also allows programmers to implement more advanced prefetching behavior by using cudaMemAdvise() and cudaMemPrefetchAsync(), which are provided by the original UM-P [10]. The advice parameter will hint the UM-P driver (and hence DRAGON) about the data access behavior. When cudaMemPrefetchAsync() is called, DRAGON will replicate the intended effects of advice for the addresses residing on NVM. This behavior can be utilized to minimize NVM read overhead for scattered data, where Linux’s read-ahead will not help due to nonconsecutive access patterns. [Page 3], Changes required for CUDA applications to work with the DRAGON framework are minimal. Applications can take advantage of DRAGON by using the dragon_map() API function with a file path and other optimization parameters (explained later) to map previously dumped binary data into a unified memory space. This function internally finds an unmapped virtual address space, registers this space along with the file path to the internal tracking mechanism of our driver extension, and returns the virtual address to the application. Once mapped, the memory range is accessible by both GPUs and CPUs, and DRAGON transparently allows direct access down to NVM devices during GPU kernel execution… In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space. In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space. NVIDIA driver for Pascal and Volta architectures captures page faults from both GPUs and CPUs, and DRAGON relies on this hardware-based pagefaulting mechanism to handle the accesses from GPU kernels to the mapped files. The memory consistency between GM and HM is handled by the original UM-P module, whereas data swapping and consistency between HM and NVM are handled by our driver extension (as explained in detail in § III-B).)
As per claim 6, rejection of claim 1 is incorporated:
Pak teaches wherein the one or more virtual memory addresses is a range of virtual memory accessible by one or more central processing units (CPUs) and the one or more hardware GPUs. ([Page 3], Changes required for CUDA applications to work with the DRAGON framework are minimal. Applications can take advantage of DRAGON by using the dragon_map() API function with a file path and other optimization parameters (explained later) to map previously dumped binary data into a unified memory space. This function internally finds an unmapped virtual address space, registers this space along with the file path to the internal tracking mechanism of our driver extension, and returns the virtual address to the application. Once mapped, the memory range is accessible by both GPUs and CPUs, and DRAGON transparently allows direct access down to NVM devices during GPU kernel execution… In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space.)
As per claim 7, rejection of claim 1 is incorporated:
Wagle discloses returning an address.
Pak teaces wherein the circuitry, in response to receiving the API call, is to return an identifier of a location type and an identifier of a location identity based, at least in part, on one or more parameters that indicate a range of memory. ([Page 3], Changes required for CUDA applications to work with the DRAGON framework are minimal. Applications can take advantage of DRAGON by using the dragon_map() API function with a file path and other optimization parameters (explained later) to map previously dumped binary data into a unified memory space. This function internally finds an unmapped virtual address space, registers this space along with the file path to the internal tracking mechanism of our driver extension, and returns the virtual address to the application. Once mapped, the memory range is accessible by both GPUs and CPUs, and DRAGON transparently allows direct access down to NVM devices during GPU kernel execution… In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space.)
As per claim(s) 8-10, 12 and 13 is/are system claim(s) corresponding to processor claim(s) 1-6. Therefore, rejected based on similar rationale.
As per claim 11, rejection of claim 8 is incorporated:
Pak teaches wherein the one or more processors, in response to receiving the API call, are to indicate a NUMA node to which a range of virtual memory indicated by the one or more users was last prefetched. ([Page 3], Changes required for CUDA applications to work with the DRAGON framework are minimal. Applications can take advantage of DRAGON by using the dragon_map() API function with a file path and other optimization parameters (explained later) to map previously dumped binary data into a unified memory space. This function internally finds an unmapped virtual address space, registers this space along with the file path to the internal tracking mechanism of our driver extension, and returns the virtual address to the application. Once mapped, the memory range is accessible by both GPUs and CPUs, and DRAGON transparently allows direct access down to NVM devices during GPU kernel execution… In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space. [Page 5], advanced prefetching behavior by using cudaMemAdvise() and cudaMemPrefetchAsync(), which are provided by the original UM-P [10]. The advice parameter will hint the UM-P driver (and hence DRAGON) about the data access behavior. When cudaMemPrefetchAsync() is called, DRAGON will replicate the intended effects of advice for the addresses residing on NVM. This behavior can be utilized to minimize NVM read overhead for scattered data, where Linux’s read-ahead will not help due to nonconsecutive access patterns.)
As per claim(s) 14-17, 19 is/are method claim(s) corresponding to processor claim(s) 1, 2, 5, 6. Therefore, rejected based on similar rationale.
[Claim 15 Narayanan discloses managed memory location paragraph 11, fig. 3 and Roberts paragraph 11]
As per claim 18, rejection of claim 14 is incorporated:
Narayanan teaches wherein the one or more hardware NUMA storages are physical memories included in one or more NUMA nodes that each include one or more central processing units (CPUs). ([Paragraph 31], With software-based Input/Output (I/O) virtualization, the VMM exposes a virtual device (such as network interface controller (NIC) functionality, for example) to a VM. Some examples of a NIC are part of an Infrastructure Processing Unit (IPU) or data processing unit (DPU) or utilized by an IPU or DPU. An IPU or DPU can include a network interface, memory devices, and one or more programmable or fixed function processors (e.g., CPU or XPU) to perform offload of operations that could have been performed by a host CPU or XPU or remote CPU or XPU. In some examples, the IPU or DPU can perform virtual switch operations, manage storage transactions (e.g., compression, cryptography, virtualization), and manage operations performed on other IPUs, DPUs, servers, or devices. [Paragraph 76], The SoC 804 includes at least one Central Processing Unit (CPU) module 808, a memory controller 814, and a Graphics Processor Unit (GPU) module 810. In other embodiments, the memory controller 814 may be external to the SoC 804 and the GPU module may be external to the SoC 804. The CPU module 808 includes at least one processor core 802 and a level 2 (L2) cache 806.)
Pak also teaches ([Page 2], Going beyond system memory capacity limits: Active-Pointers [28] is the only study that considered expanding GPU addressable memory range by mapping it to a file-system. Active Pointers is based on GPUs [29] and relies on in-kernel software address translation (i.e., SW-based page faulting). It requires existing CUDA kernels to be modified to use custom pointers and APIs. Memory references are captured on the fly via operator overloading, and a page-fault handling kernel-code is executed for every access. This approach incurs significant overhead in addition to the extensive cost of modifying existing GPU kernels; therefore, it is relatively inefficient as it does not exploit the new hardware paging support in GPUs, as demonstrated in § IV. [Page 3], In DRAGON, there are three locations where data can reside in: GPU memory (GM), host memory (HM), and nonvolatile memory (NVM). The content of the file that is mapped via dragon_map() is visible to both CPUs and GPUs under the same unified virtual address space. NVIDIA driver for Pascal and Volta architectures captures page faults from both GPUs and CPUs, and DRAGON relies on this hardware-based page faulting mechanism to handle the accesses from GPU kernels to the mapped files. The memory consistency between GM and HM is handled by the original UM-P module, whereas data swapping and consistency between HM and NVM are handled by our driver extension (as explained in detail in § III-B).)
As per claim 20, rejection of claim 14 is incorporated:
Narayanan teaches A non-transitory computer-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least perform the method of claim 14. ([Paragraph 86], A non-transitory machine-readable storage medium can cause a machine to perform the functions or operations described, and includes any mechanism that stores information in a form accessible by a machine (for example, computing device, electronic system, etc.), such as recordable/non-recordable media (for example, read only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.).)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONG U KIM whose telephone number is (571)270-1313. The examiner can normally be reached 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at 5712723338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DONG U KIM/Primary Examiner, Art Unit 2197