DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1-20 are pending.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on June 29, 2023 is/are in compliance with the provisional of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 8-11, 13-15, and 17-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vakharwala et al. (US 2022/0414020) (hereinafter Vak) (published December 29, 2022) in view of Tsirkin et al. (US 2017/0357579) (hereinafter Tsirkin) (published December 14, 2017).
Regarding Claim 1, Vak discloses an apparatus comprising: at least one accelerator to perform operations on data; and
“SVM allows CPUs and XPUs to manipulate complex data structures (e.g., trees) in memory without needless data copying. SVM also enables XPUs to handle page-faults, removing the requirement for XPUs to pin their memory, and thereby, for example, have a much larger working set size” (Vak [0011])
“As further illustrated, host processor 210 couples to multiple accelerators 250A,B. Although two accelerators are shown in this implementation, additional accelerators may be present in a particular implementation” (Vak [0026])
an address translation cache (ATC) coupled to the at least one accelerator, the ATC to store address translations, wherein the ATC is to:
“Similarly, root complex 140 includes another MMU, namely an IOMMU 142, that may store address translations on behalf of XPUs 130. Thus as shown, requests for translation may be received in root complex 140 from given XPUs 130 and in turn IOMMU 142 provides a physical address. Such translations may be stored in a TLB within XPU 130, referred to as a device TLB or more particularly herein, an address translation cache (ATC) 132” (Vak [0014])
“In any event, in the high level view shown in FIG. 2, direct communication paths are provided between host processor 210 and accelerators 250. More particularly as described herein, host processor 210 may be in direct communication with ATCs 260A,B included within accelerators 250A,B, via a direct software-to-ATC communication interface. In this way, the overhead and complexity of communicating indirectly between host processors and ATCs through an IOMMU or other intermediary can be avoided” (Vak [0027] the ATC is coupled in the accelerator)
send a command to a pending request queue (PRQ) stored in a memory coupled to the apparatus, the PRQ associated with a process of a software; and
“As further shown in FIG. 5, at block 530 the ATC may write a command into the page request queue. This command may be for the host processor to perform a page request operation such as obtaining or creating a new page translation for storage in the ATC. Thereafter at block 540, the ATC may update its tail pointer of the configuration register in the ATC” (Vak [0061])
“Understand that at this point, the software in execution on the host processor may access the command from the page request queue and perform the requested operation, e.g., providing a page translation for a given page for storage in the page request queue, handling a page fault or so forth. Then the software sends a configuration register write request to the ATC to update a head pointer at block 570. Understand while shown at this high level in the embodiment of FIG. 5, many variations and alternatives are possible” (Vak [0062])
send an interrupt to inform the process regarding the command.
“Still referring to FIG. 5, control next passes to block 550 where the ATC may send an interrupt directly to the host processor to inform software regarding presence of the command” (Vak [0062])
But does not disclose that the software is guest software. Tsirkin discloses a guest software associated with a request buffer.
“Typically, requests are in guest memory and are passed by the guest virtual machine using a guest address (e.g., guest physical address, guest bus address), which is typically stored in a device request buffer of the virtual device in guest memory” (Tsirkin [0011])
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify Vak in view of Tsirkin such that the process associated with the page request queue (PRQ) is guest software executing in a guest virtual machine. Vak already provides a mechanism in which an accelerator’s address translation cache (ATC) places a page request command into a page request queue and notifies software executing on the host processor so that the software can service the request. Tsirkin teaches implementing device request handling in a virtualized environment, in which requests are generated by a guest virtual machine and stored in a device request buffer in guest memory. Thus, applying Tsirkin’s guest-virtual-machine request handling to Vak’s page-request mechanism would have amounted to applying a known virtualization technique to Vak’s existing page-request architecture to permit guest software to generate and/or be associated with requests serviced through the PRQ.
The motivation for doing so would have been to enable Vak’s accelerator address-translation and page-fault handling techniques to operate in a virtualized computing environment, thereby allowing guest software executing in a guest virtual machine to access accelerator resources while retaining the benefits of Vak’s PRQ-based request handling. Such a modification would have been predictable because Vak already contemplates software servicing page requests generated by an accelerator, while Tsirkin expressly teaches passing device requests from a guest virtual machine using addresses associated with guest memory.
Regarding Claim 2, Vak further discloses wherein the apparatus is to send the command to a location in the PRQ based on information in a PRQ descriptor stored in the memory, the PRQ descriptor associated with the process.
“As illustrated, method 500 begins by receiving a configuration register write request in the ATC (block 510). This register write request may be used to identify metadata of a page request queue stored in a memory. For example, this metadata may include a base location, and initial head and tail pointers, among potentially other information. Next at block 520 this metadata may be stored into fields of one or more configuration registers. Thus at this point the ATC is ready to issue commands such as page requests using communications along the interface” (Vak [0060])
“As further shown in FIG. 5, at block 530 the ATC may write a command into the page request queue. This command may be for the host processor to perform a page request operation such as obtaining or creating a new page translation for storage in the ATC. Thereafter at block 540, the ATC may update its tail pointer of the configuration register in the ATC” (Vak [0061])
Regarding Claim 3, Vak further discloses wherein the apparatus is to send the command to the location in the PRQ based at least in part on a tail pointer of the PRQ descriptor.
“As further shown in FIG. 5, at block 530 the ATC may write a command into the page request queue. This command may be for the host processor to perform a page request operation such as obtaining or creating a new page translation for storage in the ATC. Thereafter at block 540, the ATC may update its tail pointer of the configuration register in the ATC” (Vak [0061] the command would be written at the location of the tail pointer and have the pointer updated)
Regarding Claim 8, Vak further discloses wherein the ATC is to: request an address of the PRQ from a translation agent; and in response, receive an address of a PRQ descriptor, the PRQ descriptor associated with the PRQ.
“One of the many challenges that legacy XPUs face when trying to take advantage of SVM is to build a PCIe ATS. ATS allows XPUs to request address translations from the IOMMU (aka Translation Agent—TA) and cache the results in a translation cache, ATC, which decouples the translation caching requirement of XPUs from the translation caches available in Root Complex IOMMU” (Vak [0016])
“As discussed, PCIe ATS may allow XPUs to build an ATC to improve performance. However, legacy definitions of the ATS may not allow system software to communicate with an ATC. Instead, legacy definitions may require system software to use a IOMMU as a “middle-man,” and all the communication between system software and ATC may occur via IOMMU” (Vak [0017])
“To perform initialization of device queue 245, software 220 may write, e.g., via a configuration register write, directly to ATC 260.sub.A to write this initialization information regarding the device queue 245A into configuration register(s) 264A. Thereafter, when commands are written into device queue 245A, software 220 may send a configuration register write, e.g., to update the tail pointer, to indicate presence of this additional command” (Vak [0032] for legacy definitions this initialization would happen via the IOMMU)
“As illustrated, method 300 begins by initializing a device queue (which in this embodiment is an invalidation queue) in memory (block 310). Such initialization may be used to identify a base location for this queue, along with its parameters, including its size, capabilities and so forth” (Vak [0045])
Vak at paragraph [0017] explains that, under the legacy ATS definitions, the IOMMU serves as the intermediary for communication between system software and the ATC. Thus, when the software initializes the device queue as described in [0045] by identifying the queue's base location, the corresponding address information would be communicated to the ATC through the IOMMU. In this manner, the IOMMU, identified by Vak as the Translation Agent in [0016], provides the address translation to the ATC for the queue's base location.
Regarding Claim 9, Vak further discloses wherein the apparatus further comprises a device process information cache, wherein the apparatus is to store the address of the PRQ descriptor in the device process information cache.
“As illustrated, each ATC 260 may include a cache memory 262, one or more configuration registers 264, and a cache controller 266. Although embodiments are not limited in this regard, configuration registers 264 may be implemented as part of an accelerator's PCIe configuration space and may provide storage for various information, including a queue base, head and tail pointers and certain process address space identifier (PASID) and privilege information associated with this device queue” (Vak [0031] the configuration registers would be the process information cache)
Regarding Claim 10, Vak discloses a method comprising: generating, in an address translation cache (ATC) of an accelerator, a command to be communicated to a software in execution on a host processor coupled to the accelerator, the command for a memory location in a memory;
“As further illustrated, host processor 210 couples to multiple accelerators 250A,B. Although two accelerators are shown in this implementation, additional accelerators may be present in a particular implementation” (Vak [0026])
“In an embodiment, ATC 260 may constantly monitor Head and Tail registers. For example, if Head=Tail−1, ATC 260 knows that DevPRQ is full, and it needs to wait and not generate any new Page Requests. If there is space in DevPRQ, ATC 260 can write a new Page Request into DevPRQ by issuing a Memory Write (which may go through IOMMU DMA remapping process just like any other DMA write) to an address calculated by adding Tail to the Base register. ATC 260 then sends an interrupt to software asking for processing of commands in DevPRQ” (Vak [0057] ATC generates new page requests to memory locations when the DevPRQ is not full)
accessing a pending request queue (PRQ) descriptor to read pointer information and a pending request status, the PRQ descriptor associated with a PRQ of the memory, the PRQ associated with the accelerator and a process of the software; and
“Thus as shown in FIG. 2, for each ATC 260 the software may use a standard circular buffer in memory with head/tail pointers to put commands for the ATC. Software may then ask the ATC to take the commands and process them by updating the tail pointer, which is a register in ATC” (Vak [0030])
“Although embodiments are not limited in this regard, configuration registers 264 may be implemented as part of an accelerator's PCIe configuration space and may provide storage for various information, including a queue base, head and tail pointers and certain process address space identifier (PASID) and privilege information associated with this device queue” (Vak [0031])
writing the command to the PRQ using the pointer information.
“As shown in FIG. 4, ATC 260 may send command(s) to software 220 in accordance with various embodiments herein. For each ATC 260, software 220 may use a standard circular buffer in memory 240 with head/tail pointers to receive commands from ATC 260. ATC 260 writes the commands into this buffer using Memory Write opcodes, updates the tail pointer, and then sends an interrupt to inform software 220 that there are commands to process in the buffer” (Vak [0049])
But does not disclose that the software is guest software. Tsirkin discloses a guest software associated with a request buffer.
“Typically, requests are in guest memory and are passed by the guest virtual machine using a guest address (e.g., guest physical address, guest bus address), which is typically stored in a device request buffer of the virtual device in guest memory” (Tsirkin [0011])
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify Vak in view of Tsirkin such that the process associated with the page request queue (PRQ) is guest software executing in a guest virtual machine. Vak already provides a mechanism in which an accelerator’s address translation cache (ATC) places a page request command into a page request queue and notifies software executing on the host processor so that the software can service the request. Tsirkin teaches implementing device request handling in a virtualized environment, in which requests are generated by a guest virtual machine and stored in a device request buffer in guest memory. Thus, applying Tsirkin’s guest-virtual-machine request handling to Vak’s page-request mechanism would have amounted to applying a known virtualization technique to Vak’s existing page-request architecture to permit guest software to generate and/or be associated with requests serviced through the PRQ.
The motivation for doing so would have been to enable Vak’s accelerator address-translation and page-fault handling techniques to operate in a virtualized computing environment, thereby allowing guest software executing in a guest virtual machine to access accelerator resources while retaining the benefits of Vak’s PRQ-based request handling. Such a modification would have been predictable because Vak already contemplates software servicing page requests generated by an accelerator, while Tsirkin expressly teaches passing device requests from a guest virtual machine using addresses associated with guest memory.
Regarding Claim 11, Vak further discloses further comprising sending an interrupt to inform the guest software regarding the command based at least in part on the pending request status.
“Still referring to FIG. 5, control next passes to block 550 where the ATC may send an interrupt directly to the host processor to inform software regarding presence of the command” (Vak [0062])
Regarding Claim 13, Vak further discloses further comprising: requesting, via the ATC, an address of the PRQ in the memory from a translation agent; and in response to requesting the address of the PRQ, receiving an address of the PRQ descriptor and storing the address of the PRQ descriptor in a device process information cache of the ATC, the PRQ descriptor associated with the accelerator and the process of the guest software and comprising the address of the PRQ.
“One of the many challenges that legacy XPUs face when trying to take advantage of SVM is to build a PCIe ATS. ATS allows XPUs to request address translations from the IOMMU (aka Translation Agent—TA) and cache the results in a translation cache, ATC, which decouples the translation caching requirement of XPUs from the translation caches available in Root Complex IOMMU” (Vak [0016])
“As discussed, PCIe ATS may allow XPUs to build an ATC to improve performance. However, legacy definitions of the ATS may not allow system software to communicate with an ATC. Instead, legacy definitions may require system software to use a IOMMU as a “middle-man,” and all the communication between system software and ATC may occur via IOMMU” (Vak [0017])
“To perform initialization of device queue 245, software 220 may write, e.g., via a configuration register write, directly to ATC 260.sub.A to write this initialization information regarding the device queue 245A into configuration register(s) 264A. Thereafter, when commands are written into device queue 245A, software 220 may send a configuration register write, e.g., to update the tail pointer, to indicate presence of this additional command” (Vak [0032] for legacy definitions this initialization would happen via the IOMMU, also see Fig. 2 the queue is associated with the accelerator and software)
“As illustrated, method 300 begins by initializing a device queue (which in this embodiment is an invalidation queue) in memory (block 310). Such initialization may be used to identify a base location for this queue, along with its parameters, including its size, capabilities and so forth” (Vak [0045])
Vak at paragraph [0017] explains that, under the legacy ATS definitions, the IOMMU serves as the intermediary for communication between system software and the ATC. Thus, when the software initializes the device queue as described in [0045] by identifying the queue's base location, the corresponding address information would be communicated to the ATC through the IOMMU. In this manner, the IOMMU, identified by Vak as the Translation Agent in [0016], provides the address translation to the ATC for the queue's base location.
Regarding Claim 14, Vak further discloses further comprising storing, in the device process information cache, a plurality of PRQ descriptor addresses, each of the plurality of PRQ descriptor addresses associated with a process of the guest software.
“As illustrated, each ATC 260 may include a cache memory 262, one or more configuration registers 264, and a cache controller 266. Although embodiments are not limited in this regard, configuration registers 264 may be implemented as part of an accelerator's PCIe configuration space and may provide storage for various information, including a queue base, head and tail pointers and certain process address space identifier (PASID) and privilege information associated with this device queue” (Vak [0031] the configuration registers would be the process information cache)
Regarding Claim 15, Vak further discloses further comprising receiving, from the translation agent, an address of the PRQ descriptor, the PRQ descriptor comprising a head pointer, a tail pointer, and a pending request status indicator.
“As discussed, PCIe ATS may allow XPUs to build an ATC to improve performance. However, legacy definitions of the ATS may not allow system software to communicate with an ATC. Instead, legacy definitions may require system software to use a IOMMU as a “middle-man,” and all the communication between system software and ATC may occur via IOMMU” (Vak [0017])
“As illustrated, each ATC 260 may include a cache memory 262, one or more configuration registers 264, and a cache controller 266. Although embodiments are not limited in this regard, configuration registers 264 may be implemented as part of an accelerator's PCIe configuration space and may provide storage for various information, including a queue base, head and tail pointers and certain process address space identifier (PASID) and privilege information associated with this device queue” (Vak [0031] pending status can be indicated by the positions of the head and tail pointers as described in [0040] and [0057])
Regarding Claim 17, Vak discloses a system comprising: a memory; a central processing unit (CPU) coupled to the memory, the CPU having at least one core to execute instructions, the CPU to execute a software; and
“As seen in the high level of FIG. 2, computing system 200 (which may be any type of computing device ranging from small portable devices such as smartphones, tablets or so forth to larger systems including client or other desktop systems, server systems or so forth) includes a host processor 210 that couples to a memory 240. As an example, host processor 210 may be some type of multicore processor or other system on chip (SoC) that in turn is coupled to memory 240, which may be implemented as a dynamic random access memory (DRAM)” (Vak [0025])
“Core 215 may be any type of processing core and in different implementations may be an in-order or out-of-order processing core. A software 220 is illustrated that may execute on core 215. In various implementations, software 220 may be a system software such as an operating system, hypervisor, firmware or so forth” (Vak [0028])
a processing circuit coupled to the CPU, the processing circuit comprising: an accelerator to perform operations on data; and
“SVM allows CPUs and XPUs to manipulate complex data structures (e.g., trees) in memory without needless data copying. SVM also enables XPUs to handle page-faults, removing the requirement for XPUs to pin their memory, and thereby, for example, have a much larger working set size” (Vak [0011])
“As further illustrated, host processor 210 couples to multiple accelerators 250A,B. Although two accelerators are shown in this implementation, additional accelerators may be present in a particular implementation. As one example, accelerators 250 may be graphics processors, where each accelerator 250 includes a plurality of independent graphics processing units (GPUs)” (Vak [0026])
an address translation cache (ATC) coupled to the accelerator, the ATC to store address translations,
“Similarly, root complex 140 includes another MMU, namely an IOMMU 142, that may store address translations on behalf of XPUs 130. Thus as shown, requests for translation may be received in root complex 140 from given XPUs 130 and in turn IOMMU 142 provides a physical address. Such translations may be stored in a TLB within XPU 130, referred to as a device TLB or more particularly herein, an address translation cache (ATC) 132” (Vak [0014])
“In any event, in the high level view shown in FIG. 2, direct communication paths are provided between host processor 210 and accelerators 250. More particularly as described herein, host processor 210 may be in direct communication with ATCs 260A,B included within accelerators 250A,B, via a direct software-to-ATC communication interface. In this way, the overhead and complexity of communicating indirectly between host processors and ATCs through an IOMMU or other intermediary can be avoided” (Vak [0027])
wherein the ATC is to directly send a command to a process of the software via a pending request queue (PRQ) stored in the memory, the PRQ associated with the process of the software.
“As further shown in FIG. 5, at block 530 the ATC may write a command into the page request queue. This command may be for the host processor to perform a page request operation such as obtaining or creating a new page translation for storage in the ATC. Thereafter at block 540, the ATC may update its tail pointer of the configuration register in the ATC” (Vak [0061])
“Understand that at this point, the software in execution on the host processor may access the command from the page request queue and perform the requested operation, e.g., providing a page translation for a given page for storage in the page request queue, handling a page fault or so forth. Then the software sends a configuration register write request to the ATC to update a head pointer at block 570. Understand while shown at this high level in the embodiment of FIG. 5, many variations and alternatives are possible” (Vak [0062])
But does not disclose that the software is guest software. Tsirkin discloses a guest software associated with a request buffer.
“Typically, requests are in guest memory and are passed by the guest virtual machine using a guest address (e.g., guest physical address, guest bus address), which is typically stored in a device request buffer of the virtual device in guest memory” (Tsirkin [0011])
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify Vak in view of Tsirkin such that the process associated with the page request queue (PRQ) is guest software executing in a guest virtual machine. Vak already provides a mechanism in which an accelerator’s address translation cache (ATC) places a page request command into a page request queue and notifies software executing on the host processor so that the software can service the request. Tsirkin teaches implementing device request handling in a virtualized environment, in which requests are generated by a guest virtual machine and stored in a device request buffer in guest memory. Thus, applying Tsirkin’s guest-virtual-machine request handling to Vak’s page-request mechanism would have amounted to applying a known virtualization technique to Vak’s existing page-request architecture to permit guest software to generate and/or be associated with requests serviced through the PRQ.
The motivation for doing so would have been to enable Vak’s accelerator address-translation and page-fault handling techniques to operate in a virtualized computing environment, thereby allowing guest software executing in a guest virtual machine to access accelerator resources while retaining the benefits of Vak’s PRQ-based request handling. Such a modification would have been predictable because Vak already contemplates software servicing page requests generated by an accelerator, while Tsirkin expressly teaches passing device requests from a guest virtual machine using addresses associated with guest memory.
Regarding Claim 18, Vak further discloses further comprising a translation agent coupled to the processing circuit, the translation agent to send the ATC an address of a PRQ descriptor stored in the memory, the PRQ descriptor associated with the process.
“One of the many challenges that legacy XPUs face when trying to take advantage of SVM is to build a PCIe ATS. ATS allows XPUs to request address translations from the IOMMU (aka Translation Agent—TA) and cache the results in a translation cache, ATC, which decouples the translation caching requirement of XPUs from the translation caches available in Root Complex IOMMU” (Vak [0016])
“As discussed, PCIe ATS may allow XPUs to build an ATC to improve performance. However, legacy definitions of the ATS may not allow system software to communicate with an ATC. Instead, legacy definitions may require system software to use a IOMMU as a “middle-man,” and all the communication between system software and ATC may occur via IOMMU” (Vak [0017])
“To perform initialization of device queue 245, software 220 may write, e.g., via a configuration register write, directly to ATC 260.sub.A to write this initialization information regarding the device queue 245A into configuration register(s) 264A. Thereafter, when commands are written into device queue 245A, software 220 may send a configuration register write, e.g., to update the tail pointer, to indicate presence of this additional command” (Vak [0032] for legacy definitions this initialization would happen via the IOMMU)
“As illustrated, method 300 begins by initializing a device queue (which in this embodiment is an invalidation queue) in memory (block 310). Such initialization may be used to identify a base location for this queue, along with its parameters, including its size, capabilities and so forth” (Vak [0045])
Vak at paragraph [0017] explains that, under the legacy ATS definitions, the IOMMU serves as the intermediary for communication between system software and the ATC. Thus, when the software initializes the device queue as described in [0045] by identifying the queue's base location, the corresponding address information would be communicated to the ATC through the IOMMU. In this manner, the IOMMU, identified by Vak as the Translation Agent in [0016], provides the address translation to the ATC for the queue's base location.
Regarding Claim 19, Vak further discloses wherein the processing circuit further comprises a process information cache to store the address of the PRQ descriptor, the PRQ descriptor comprising an address of the PRQ and an indicator to indicate whether the ATC is to send an interrupt to inform the process of the command.
“As illustrated, each ATC 260 may include a cache memory 262, one or more configuration registers 264, and a cache controller 266. Although embodiments are not limited in this regard, configuration registers 264 may be implemented as part of an accelerator's PCIe configuration space and may provide storage for various information, including a queue base, head and tail pointers and certain process address space identifier (PASID) and privilege information associated with this device queue” (Vak [0031] the configuration registers would be the process information cache)
Claims 4 and 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vak (published December 29, 2022) and Tsirkin (published December 14, 2017) as applied to claim 2 above, and further in view of Whitney et al. (US 4,637,797) (hereinafter Whitney) (published January 20, 1987).
Regarding Claim 4, the combination of Vak and Tsirkin disclosed the apparatus of claim 2, but does not explicitly state wherein the ATC is to send the interrupt when a pending request interrupt indicator of the PRQ descriptor indicates that there are no unserviced pending request interrupts.
Whitney discloses wherein the ATC is to send the interrupt when a pending request interrupt indicator of the PRQ descriptor indicates that there are no unserviced pending request interrupts.
“Referring to FIG. 11, the Send Byte Interrupt routine begins by saving the computer's registers, setting an In Progress flag, and disabling the generation of additional Send Byte interrupts (box 270). The disabling of the interrupt is necessary to make sure that a glitch on the interrupt line does not cause the Send Byte routine to reenter itself in the middle of its operation--a problem which can could otherwise cause the recording process to fail occasionally” (Whitney col 15 lines 23-31)
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify the system in the combination of Vak and Tsirkin in view of Whitney such that the ATC sends the interrupt based on the status of a pending request interrupt indicator associated with the PRQ descriptor. Applying Whitney's interrupt-status technique to Vak's existing ATC interrupt mechanism would therefore have provided a predictable way to determine whether a pending request interrupt had already been serviced before generating another interrupt. The motivation for doing so would have been to improve the reliability and efficiency of Vak's interrupt mechanism by avoiding redundant or reentrant interrupts when a pending request interrupt is already being serviced. Whitney expressly recognizes that preventing additional interrupts during processing avoids a glitch or repeated interrupt from causing the interrupt-handling routine to reenter itself and potentially disrupt operation.
Regarding Claim 5, the combination of Vak and Tsirkin disclosed the apparatus of claim 2, and Vak further discloses wherein the ATC is to: send a second command to the PRQ; and
“the core is to receive an interrupt from a first ATC, the interrupt to indicate presence of a second command from the first ATC for a software in execution on the core to perform another operation” (Vak [0085])
But does not explicitly state not send another interrupt to inform the process regarding the second command when a pending request interrupt indicator of the PRQ descriptor indicates that there is at least one pending request interrupt outstanding to guest software.
Whitney discloses not send another interrupt to inform the process regarding the second command when a pending request interrupt indicator of the PRQ descriptor indicates that there is at least one pending request interrupt outstanding to guest software.
“Referring to FIG. 11, the Send Byte Interrupt routine begins by saving the computer's registers, setting an In Progress flag, and disabling the generation of additional Send Byte interrupts (box 270). The disabling of the interrupt is necessary to make sure that a glitch on the interrupt line does not cause the Send Byte routine to reenter itself in the middle of its operation--a problem which can could otherwise cause the recording process to fail occasionally” (Whitney col 15 lines 23-31)
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify the system in the combination of Vak and Tsirkin in view of Whitney such that the ATC sends the interrupt based on the status of a pending request interrupt indicator associated with the PRQ descriptor. Applying Whitney's interrupt-status technique to Vak's existing ATC interrupt mechanism would therefore have provided a predictable way to determine whether a pending request interrupt had already been serviced before generating another interrupt. The motivation for doing so would have been to improve the reliability and efficiency of Vak's interrupt mechanism by avoiding redundant or reentrant interrupts when a pending request interrupt is already being serviced. Whitney expressly recognizes that preventing additional interrupts during processing avoids a glitch or repeated interrupt from causing the interrupt-handling routine to reenter itself and potentially disrupt operation.
Claims 6 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vak (published December 29, 2022) and Tsirkin (published December 14, 2017) as applied to claim 2 and 10 above, and further in view of Tati et al. (US 9,086,981) (hereinafter Tati) (published July 21, 2015).
Regarding Claim 6, the combination of Vak and Tsirkin disclosed the apparatus of claim 2, but does not explicitly state wherein the apparatus is to send the command in response to an access permission violation to a location in the memory, the memory comprising a shared memory to be shared by the apparatus and a host processor on which the guest software is to execute.
Vak and Tati discloses wherein the apparatus is to send the command in response to an access permission violation to a location in the memory, the memory comprising a shared memory to be shared by the apparatus and a host processor on which the guest software is to execute.
“In an example, the ATC is to send a command for storage in the queue and update the pointer stored in the configuration register to indicate presence of the command in the queue” (Vak [0098])
“In an example, the command comprises a page request and the ATC is to identify completion of a page request operation by the processor on receipt of a page response command from the software, the page response command stored in another queue in the shared memory” (Vak [0099])
“Additional examples of page fault causes include the absence of a virtual to physical page mapping, an access permission violation, or writing to read-only memory. When a page fault occurs, the VMkernel locates the data that corresponds to the faulted page within the non-volatile storage, reads the data from the non-volatile storage, and transfers the data into the volatile memory” (Tati col5 lines 59-65, access permission violation condition that triggers page-fault is handled via the page responses in Vak)
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify the system in the combination of Vak and Tsirkin in view of Tati such that the ATC sends the page request command in response to an access permission violation. A person of ordinary skill would therefore have recognized that Vak's page request mechanism could be used to request servicing of a page fault caused by the access permission violation disclosed by Tati, thereby providing the claimed response to the access permission violation. The motivation for combining the teachings would have been to provide an established mechanism for handling page faults, including page faults resulting from access permission violations, in the accelerator system of Vak.
Regarding Claim 16, the combination of Vak and Tsirkin disclosed the method of claim 10, but does not explicitly state further comprising generating the command in response to an access permission violation for the memory location in the memory.
Vak and Tati further comprising generating the command in response to an access permission violation for the memory location in the memory.
“In an example, the ATC is to send a command for storage in the queue and update the pointer stored in the configuration register to indicate presence of the command in the queue” (Vak [0098])
“In an example, the command comprises a page request and the ATC is to identify completion of a page request operation by the processor on receipt of a page response command from the software, the page response command stored in another queue in the shared memory” (Vak [0099])
“Additional examples of page fault causes include the absence of a virtual to physical page mapping, an access permission violation, or writing to read-only memory. When a page fault occurs, the VMkernel locates the data that corresponds to the faulted page within the non-volatile storage, reads the data from the non-volatile storage, and transfers the data into the volatile memory” (Tati col5 lines 59-65, access permission violation condition that triggers page-fault is handled via the page responses in Vak)
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify the system in the combination of Vak and Tsirkin in view of Tati such that the ATC sends the page request command in response to an access permission violation. A person of ordinary skill would therefore have recognized that Vak's page request mechanism could be used to request servicing of a page fault caused by the access permission violation disclosed by Tati, thereby providing the claimed response to the access permission violation. The motivation for combining the teachings would have been to provide an established mechanism for handling page faults, including page faults resulting from access permission violations, in the accelerator system of Vak.
Claims 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vak (published December 29, 2022), Tsirkin (published December 14, 2017), Tati (published July 21, 2015)and as applied to claim 6 above, and further in view of Kuwahara et al. (US 2019/0347215) (hereinafter Kuwahara) (published November 14, 2019).
Regarding Claim 7, the combination of Vak, Tsirkin, and Tati disclosed the apparatus of claim 6, but does not explicitly state wherein the ATC is to receive a translation completion comprising a virtual address-to-physical address translation for the location in the memory, the translation completion to indicate the access permission violation.
Kuwahara discloses wherein the ATC is to receive a translation completion comprising a virtual address-to-physical address translation for the location in the memory, the translation completion to indicate the access permission violation.
“Fundamentally, only a virtual address that is input and a physical address that is to be accessed are required as address information when translation is successfully completed. However, a real address is also stored in the TLB in case of translation failure such as an access permission violation, and this results in a problem in which the circuit area and the power consumption increase” (Kuwahara [0012] the result of address translation includes the virtual and physical addresses and the real address is use with respect to the translation to determine access permission violation)
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify the system in the combination of Vak, Tsirkin, and Tati in view of Kuwahara such that the ATC receives a translation completion comprising a virtual address-to-physical address translation and indicating an access permission violation. By applying Kuwahara's translation-result and access-permission checking technique to Vak's address-translation mechanism would have provided the ATC with translation information that identifies both the translated address and an access permission violation. The motivation for combining the teachings would have been to enable Vak's ATC and page-request mechanism to detect and appropriately handle access permission violations encountered during address translation.
Claims 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vak (published December 29, 2022) and Tsirkin (published December 14, 2017) as applied to claim 11 above, and further in view of DALY et al. (US 2020/0403940) (hereinafter Daly) (published December 24, 2020).
Regarding Claim 12, the combination of Vak and Tsirkin disclosed the method of claim 11, but does not explicitly state further comprising sending the interrupt to a host software, the host software to provide the interrupt to the guest software.
Daly discloses further comprising sending the interrupt to a host software, the host software to provide the interrupt to the guest software.
“Virtio-net Interrupt Request (IRQ) Relay in Contemporary Linux-Based Systems. For both versions of virtio-net, the IRQ is filtered through QEMU just as the notify kick. This IRQ relay translates the hardware interrupt in the host into a software fast interrupt request path (IRQFD) in the guest” (Daly [0033])
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify the system in the combination of Vak and Tsirkin in view of Daly such that the interrupt generated by the ATC is provided to host software, which in turn provides the interrupt to the guest software. By applying Daly's host-to-guest interrupt relay mechanism to Vak's ATC-generated interrupt would have allowed the interrupt generated by Vak's ATC to be received by host software and subsequently delivered to the guest software of Tsirkin. The motivation for combining the teachings would have been to enable Vak's interrupt-based page-request mechanism to operate in the virtualized environment taught by Tsirkin, such that guest software could receive notifications of requests generated by the accelerator.
Claims 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vak (published December 29, 2022) and Tsirkin (published December 14, 2017) as applied to claim 17 above, and further in view of Trikalinou et al. (US 2021/0026543) (hereinafter Trikalinou) (published January 28, 2021).
Regarding Claim 20, the combination of Vak and Tsirkin disclosed the system of claim 17, but does not explicitly state wherein the memory is to store a process address space identifier (PASID) table, the PASID table to store a plurality of entries, at least one of which is associated with the process and to store the address of the PRQ descriptor.
Vak and Trikalinou discloses wherein the memory is to store a process address space identifier (PASID) table, the PASID table to store a plurality of entries, at least one of which is associated with the process and to store the address of the PRQ descriptor.
“Although embodiments are not limited in this regard, configuration registers 264 may be implemented as part of an accelerator's PCIe configuration space and may provide storage for various information, including a queue base, head and tail pointers and certain process address space identifier (PASID) and privilege information associated with this device queue” (Vak [0031])
“In one embodiment, IOMMU 310 walks the context tables using the Bus, Device, Function and Process Address Space ID (PASID) information included in a Requestor Identifier (ReqID) received in the transaction. Thus, IOMMU 310 finds a PASID table entry that includes metadata for a device process that generated the device memory request” (Trikalinou [0030])
It would have been obvious before the effective filing date of the invention to one of ordinary skill in the art to modify the system in the combination of Vak and Tsirkin in view of Trikalinou such that the memory stores a PASID table having a plurality of entries, including an entry associated with the process and storing address information corresponding to the PRQ descriptor. Vak teaches storing a queue base, head and tail pointers, and PASID information associated with a device queue, while Trikalinou teaches a PASID table having entries that include metadata for respective device processes. Accordingly, by applying Trikalinou's PASID-table organization to Vak's existing PASID and queue information would have provided a predictable manner of associating the queue/address information with the particular process identified by the corresponding PASID entry. The motivation for combining the teachings would have been to provide a structured and scalable mechanism for associating Vak's device-queue information with individual device processes in a system supporting multiple processes and address spaces.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
KAMINSKI et al. (US 2011/0161619) discloses the accelerator device generating a page fault and sending it to the driver when the TLB does not include address translation entry or has insufficient access rights
KAKAIYA et al. (US 2018/0321985) discloses handling of interrupts and communications between host and guest drivers
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY LI whose telephone number is (571)270-5967. The examiner can normally be reached Monday to Friday 10:00 AM to 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Arpan P Savla can be reached at (571) 272-1077. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.L./Examiner, Art Unit 2137
/PRASITH THAMMAVONG/Primary Examiner, Art Unit 2137