DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 21-40 are pending in this application. Claims 1-20 were cancelled.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
The following title is suggested: COMPUTE KERNEL PARSING WITH LIMITS IN ONE OR MORE DIMENSIONS WITH EXECUTION OF BATCH OF WORKGROUPS USING ADJUSTED COORDINATES.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 21-40 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
Claim 21 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1, Statutory Category: Yes, the claim 21 is an apparatus that recites a series of steps and therefore falls in the statutory category of a machine.
Step 2A- Prong 1: Judicial Exception Recited: Yes, the claim recites: “; determine an offset value for a current sub-kernel of the compute kernel based on the sub-kernel limit value; adjust the received coordinates based on the offset value” As drafted, the claim as a whole recites an apparatus including steps that could be performed in the human mind, but for the recitation of generic computing components. The human mind can easily judging/evaluating/determining an offset value for a current sub-kernel of the compute kernel based on the sub-kernel limit value, changing/adjusting/modifying the received coordinates based on the offset value. Therefore, but for the recitation of generic computing components, these steps may be a Mental Processes that can be performed in the human mind (including an observation, evaluation, judgment, opinion).
Therefore, yes, the claims do recite judicial exceptions.
Step 2A- Prong 2: Integrated into a practical Application: No, this judicial exception is not integrated into a practical application. In particular, the claim recites an additional limitations that “receive coordinates of a batch of workgroups, wherein the coordinates include values for multiple dimensions; receive a sub-kernel limit value for a first dimension of the multiple dimensions” which is insignificant pre-solution data gathering (see MPEP § 2106.05(g)). In addition, the limitation of “graphics processor circuitry configured to execute instructions specified by workgroups; sub-kernel control circuitry configured to:, wherein the sub-kernel limit value is for compute kernel that is organized with multiple workgroups in a first dimension and multiple workgroups in a second dimension of the multiple dimensions” and “wherein the graphics processor circuitry is configured to retrieve the batch of workgroups for executing based on the adjusted coordinates” which is directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a generic computer as a tool to perform an abstract idea (see MPEP 2106.05(f)). Further, the limitation of
“transmit the adjusted coordinates to the graphics processor circuitry” which is insignificant extra solution activity (i.e., transmitting data) See MPEP 2106.05(g). Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they not impose any meaningful limits on practicing the abstract idea. Therefore, the claim is directed to the abstract idea.
Step 2B: Claim provides an Inventive Concept: No. The additional element “of “graphics processor circuitry configured to execute instructions specified by workgroups; sub-kernel control circuitry configured to:, wherein the sub-kernel limit value is for compute kernel that is organized with multiple workgroups in a first dimension and multiple workgroups in a second dimension of the multiple dimensions” and “wherein the graphics processor circuitry is configured to retrieve the batch of workgroups for executing based on the adjusted coordinates” which is directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a generic computer as a tool to perform an abstract idea (see MPEP 2106.05(f)). In addition, the limitation of “receive coordinates of a batch of workgroups, wherein the coordinates include values for multiple dimensions; receive a sub-kernel limit value for a first dimension of the multiple dimensions” which is insignificant pre-solution data gathering (see MPEP § 2106.05(g)) and the limitation of “transmit the adjusted coordinates to the graphics processor circuitry” (insignificant extra solution activity (i.e., transmitting data) See MPEP 2106.05(g)) which are well understood, routine, conventional activity (see MPEP § 2106.05(d)). Courts have identified “receiving and transmitting data, storing and retrieving information”, et cetera as well understood, routine, conventional and mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f))). These additional elements and combination of the elements does not amount to significant more than the exception itself or provide an inventive concept in Step 2B.
Under the 2019 PEG, a conclusion that an additional element is insignificant extra-solution activity in Step 2A should be re-evaluated in Step 2B. Here, the “receive” and “transmit” steps were considered to be extra-solution activity in Step 2A as insignificant data gathering and communication and are well understood, routine, conventional activity in the field. The “receive” and “transmit” steps are for the purpose of “communication” and “transmitting the data” and these can be reached on one of court case (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) see MPEP § 2106.05(d) II). Accordingly, a conclusion that “receive” and “transmit” are well understood, routine, conventional activity is supported under Berkheimer options 2.
For these reasons, there is no inventive concept in the claim, and thus the claim is ineligible.
Independent claims 30 and 38 are rejected for the same reason as claim 21 above. Claim 38 further recites “A non-transitory computer readable storage medium having stored thereon design information that specifies a design of at least a portion of a hardware integrated circuit in a format recognized by a semiconductor fabrication system that is configured to use the design information to produce the circuit according to the design”. These additional elements are directed to generic computing components/functions merely applying the abstract idea (MPEP § 2106.05(f)).
With respect to the dependent claim 22, the claim elaborates that first circuitry configured to determine, based on an increment amount and the sub-kernel limit value for the first dimension, a next position in the first dimension and an increment amount for the second dimension; second circuitry configured to determine, at least partially in parallel with the determination of the next position in the first dimension, next positions in the second dimension for multiple possible increment amounts in the second dimension; and select circuitry configured to: select one of the next positions generated by the second circuitry based on the determined increment amount for the second dimension from the first circuitry; and provide the coordinates of the batch of workgroups to the sub-kernel control circuitry based on the next position in the first dimension and the select next position in the second dimension (“determine, based on an increment amount and the sub-kernel limit value for the first dimension, a next position…”, “determine, at least partially in parallel…”, “select one of the next positions generated…” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind. Further, “provide the coordinates” which is insignificant extra solution activity (i.e., transmitting data) See MPEP 2106.05(g)) which are well understood, routine, conventional activity (see MPEP § 2106.05(d)). Courts have identified “receiving and transmitting data, storing and retrieving information”, et cetera as well understood, routine, conventional and mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f))).
With respect to the dependent claim 23, the claim elaborates that wherein the graphics processor circuitry includes multiple distributed shader processors and the transmission of the adjusted coordinates distributes workgroups from the batch of workgroups to multiple different distributed shader processors (these limitations are directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)). In addition, “distributes workgroups from the batch of workgroups to multiple different distributed shader processors” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind).
With respect to the dependent claim 24, the claim elaborates that compression circuitry configured to compress a block of output data generated by execution of workgroups of a given first sub-kernel. (these limitations are directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)).
With respect to the dependent claim 25, the claim elaborates that wherein the sub-kernel limit value corresponds to a sub-kernel shape that provides one or more performance or buffer size parameters for the compression circuitry (these limitations are directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)).
With respect to the dependent claim 26, the claim elaborates that workload parser circuitry configured to determine the sub-kernel limit value based on information in a compute command stream that includes the compute kernel (“determine the sub-kernel limit value based on information” are being treated as part of abstract idea and is analogous to Mental processes, such that concept can be performed in the human mind).
With respect to the dependent claim 27, the claim elaborates that wherein the current sub-kernel includes a limited number of workgroups in the second dimension that is smaller than the number of workgroups in the compute kernel in the second dimension (these limitations are directed to Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f)).
With respect to the dependent claim 28, the claim elaborates that wherein the sub-kernel control circuitry is configured to receive the sub-kernel limit value and transmit the adjusted coordinates in a single clock cycle (“receive” which is insignificant pre-solution data gathering (see MPEP § 2106.05(g)) and the limitation of “transmit” (insignificant extra solution activity (i.e., transmitting data) See MPEP 2106.05(g)) which are well understood, routine, conventional activity (see MPEP § 2106.05(d)). Courts have identified “receiving and transmitting data, storing and retrieving information”, et cetera as well understood, routine, conventional and mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (see MPEP 2106.05(f))).
With respect to the dependent claim 29, the claim elaborates that wherein the apparatus is a computing device that includes: a graphics processor that includes the graphics processor circuitry and the sub-kernel control circuitry; a display; a central processing unit; and network interface circuitry (These additional elements are directed to generic computing components/functions merely applying the abstract idea (MPEP § 2106.05(f)).
Dependent claims 31-37 recite the same features as applied to claims 22-28 respectively above, therefore they are also rejected under the same rationale.
Dependent claims 39-40 recite the same features as applied to claims 22 and 24 respectively above, therefore they are also rejected under the same rationale.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claim 21-40 are rejected under 35 U.S.C. 112(b), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
As per claims 21, 30 and 38 (line# refers to claim 21).
In line 8, it recites the phrase “a first dimension”. However, prior to this phrase at line 6, it recites “a first dimension”. Thus, it is unclear whether the second recitation of “a first dimension” is the same or different from the first recitation of “a first dimension”. If they are the same, the or said should be used.
As per claims 27 and 36(line# refers to claim 27):
In line 2, the phrase “the number of workgroups” lacks antecedence basis. It is uncertain if this term intent to refer to “a limited number of workgroups” as cited in claim 27, lines 1-2.
As per claims 22-29, 31-37 and 39-40:
They are method and non-transitory computer readable storage medium claims that depend from rejected claims and do not resolve the deficiencies thereof and are therefore rejected for the same reasons as above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 21, 30 and 38 are rejected under 35 U.S.C. 103 as being unpatentable over Ringseth (US Pub. 2012/0089961 A1) in view of YE (US Pub. 2022/0076487 A1).
As per claim 21, Ringseth teaches the invention substantially as claimed including An apparatus, comprising (Ringseth, Fig. 5, 100):
graphics processor circuitry configured to execute instructions specified by workgroups (Ringseth, Fig. 5, 120 compute engine, 121 compute node; Abstract, provides a tile communication operator that decomposes a computational space into sub-spaces (i.e., tiles) that may be mapped to execution structures (e.g., thread groups) of data parallel compute nodes; [0053] Computer system 100 includes a host 101 with one or more processing elements (PEs) 102 housed in one or more processor packages (not shown) and a memory system 104; [0056] The processing elements 102 in each processor package may have the same or different architectures and/or instruction sets. For example, the processing elements 102 may include any combination of in-order execution elements, superscalar execution elements, and data parallel execution elements (e.g., GPU execution elements; [0068] where the set of PEs 122 includes one or more GPUs (as graphics processor circuitry configured to execute instructions specified by workgroups (i.e., sub-spaces/tiles));
sub-kernel control circuitry configured to (Ringseth, Fig. 2, 12 tile communication operator; Fig. 5, 100, 12 within the compute node (as sub-kernel control circuitry); [0052] a computer system 100 configured to compile and execute data parallel code 10 that includes a tile communication operator):
receive coordinates of a batch of workgroups, wherein the coordinates include values for multiple dimensions (Ringseth, Fig. 1, 12 tile (indexable_type <N, T>,_Extent; Fig. 2, 14 to 12 (as receive); Abstract, a tile communication operator that decomposes a computational space (as whole as batch of workgroups) into sub-spaces (i.e., tiles) that may be mapped to execution structures; [0021] Input indexable type 14 has a rank (e.g., rank N in the embodiment of FIG. 1) and element type (e.g., element type T in the embodiment of FIG. 1) and defines the computational space that is decomposed by tile communication operator 12. For each input indexable type 14, the tile communication operator produces an output indexable type 18 with the same rank as input indexable type 14 and an element type that is a tile of input indexable type 14; see Fig. 3, 14 as multiple dimensions; [0020] an indexable type may be algebraically represented as the intersection of a finite number of half-planes formed by linear functions of the coordinate axes; [0032] tileIndex represents the coordinates of the tile containing _index;); ;
receive a sub-kernel limit value for a first dimension of the multiple dimensions (Ringseth, [0013] decomposes a computational space (represented by indexable_type <N, T> in the embodiment of FIG. 1) into sub-spaces (i.e., tiles) defined by an extent (represented by _Extent in the embodiment of FIG. 1). The tiles may be mapped to execution structures (e.g., thread groups (DirectX), thread blocks (CUDA); [0031] a tile may be set to be a multiple of the local structure constants so that the number of local structures that fit into a tile determines the unrolling factor of the implicit loop dimensions of the execution structure. This relationship may be seen in the local view decomposition: [0036] Assume there is a local view structure determined by 16.times.16 thread group dimensions (as sub-kernel limit value). Analogously, tile the matrices A, B and C into 16.times.16 tiles), wherein the sub-kernel limit value is for compute kernel that is organized with multiple workgroups in a first dimension and multiple workgroups in a second dimension of the multiple dimensions (Ringseth, Fig. 3, 18; [0036] Assume there is a local view structure determined by 16.times.16 thread group dimensions. Analogously, tile the matrices A, B and C into 16.times.16 tiles. (Assume for now that N is evenly divisible by 16--the general case checks for boundary conditions and the kernel exits early when a tile is not completely contained in the original data; [0041] a kernel that is dispatched with respect to thread group dimensions of 16.times.16, which means that 256 threads are logically executing the kernel simultaneously);
determine an offset value for a current sub-kernel of the compute kernel based on the sub-kernel limit value (Ringseth, [0029] the data structure "tile_range" forms the output indexable type 18 (also referred to as a "pseudo-field") for the tile communication operator 12 "tile". The index operator of tile_range takes: const index<_Rank>& _Index and forms a field or pseudo-field whose extent is _Tile and whose offset is: [0030] _Index*m_muitiplier within the _Parent it is constructed from; also see [0024] restricted to grid_tile, indexed by Indexable<N>/grid_tile. More particularly, if grid describes the shape of Indexable<N>, then range<N, Indexable<N>> is the collection of Indexable<N> restricted to grid_tile translated by offset in grid_range=(grid+grid_tile-1)/grid_tile. Accordingly, grid_range is the shape of range<N, Indexable<N>> when created by tile<grid_tile> (Indexable<N>);
adjust the received coordinates based on the offset value (Ringseth, [0032] _index=_tileIndex*thread_group_dimensions+localIndex, where _tileIndex represents the coordinates of the tile containing _index (i.e., input indexable type 14) and _localIndex is the offset within that tile (as to adjust the received coordinates based on the offset value to generate adjusted coordinates); and
wherein the graphics processor circuitry is configured to retrieve the batch of workgroups for executing based on the adjusted coordinates (Ringseth, Fig. 4 (as include adjusted coordinates, see [0029] and [0032]); Fig. 5, 121 compute node which includes 122 PE; [0005] decomposes a computational space into sub-spaces (i.e., tiles) that may be mapped to execution structures (e.g., thread groups) of data parallel compute nodes; [0013] The tiles may be mapped to execution structures (e.g., thread groups (DirectX), thread blocks (CUDA), work groups (OpenCL), or waves (AMD/ATI)) of data parallel (DP) optimal compute nodes such as DP optimal compute nodes 121; [0067] Each compute node 121 includes a set of one or more PEs 122 and a memory 124 that stores DP executable 138. PEs 122 execute DP executable 138 and store the results generated by DP executable 138 in memory 124. In particular, PEs 122 execute DP executable 138 to apply a tile communication operator 12 to an input indexable type 14 to generate an output indexable type 18 as shown in FIG. 5; [0068] where the set of PEs 122 includes one or more GPUs; [0070] the host compute node causes DP executable 138 and one or more indexable types 14 to be copied from memory system 104 to memory 124; see Fig. 5, PE communicate with memory for execution; [0072] instruction set of the GPUs for execution by the PEs 122 of the GPUs).
Although Ringseth teach adjusted coordinates, Ringseth fails to specifically teach that adjusted coordinates is transmit to the graphics processor circuitry.
However, YE teaches adjusted coordinates transmit to the graphics processor circuitry (YE, Claim 17, acquiring coordinate data of a heat point position by the CPU, and transmitting the coordinate data from the CPU to the GPU).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth with YE because YE’s teaching of transmit the coordinate to the graphics processor circuitry would have provided Ringseth’s system with the advantage and capability to utilizing the graphic processor to render the computing tasks/data based on the received coordinates in order to improving the resource utilization and system performance (see YE, [0004] and [0005]).
As per claim 30, it is a method claim of claim 21 above. Therefore, it is rejected for the same reason as claim 21 above.
As per claim 38, it is a non-transitory computer readable storage medium claim of claim 21 above. Therefore, it is rejected for the same reason as claim 21 above.
Claims 23 and 32 are rejected under 35 U.S.C. 103 as being unpatentable over Ringseth and YE, as applied to claims 21 and 30 respectively above, and further in view of Arvo (US Patent. 9,092,267 B2).
As per claim 23, Ringseth and YE teach the invention according to claim 21 above. Ringseth teaches adjusted coordinates (Ringseth, [0032] _index=_tileIndex*thread_group_dimensions+localIndex, where _tileIndex represents the coordinates of the tile containing _index (i.e., input indexable type 14) and _localIndex is the offset within that tile (as to adjust the received coordinates based on the offset value to generate adjusted coordinates). YE further teaches transmission of the adjusted coordinates (YE, Claim 17, acquiring coordinate data of a heat point position by the CPU, and transmitting the coordinate data from the CPU to the GPU).
Ringseth and YE fail to specifically teach wherein the graphics processor circuitry includes multiple distributed shader processors and distributes workgroups from the batch of workgroups to multiple different distributed shader processors.
However, Arvo teaches wherein the graphics processor circuitry includes multiple distributed shader processors and distributes workgroups from the batch of workgroups to multiple different distributed shader processors (Arvo, Abstract, processing data with a graphics processing unit (GPU); Col 1, line 23, using one or more shader processors residing in the GPU; Fig. 3 workgroups; Fig. 7, workgroups are distributed to multiple different distributed shader processors)).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth and YE with Arvo because Arvo’s teaching of executing the different workgroups among different shader processors in a GPU would have provided Ringseth and YE’s system with the advantage and capability to easily managing and scheduling different workgroups among the different shader processors in order to improving the system performance and efficiency.
As per claim 32, it is a method claim of claim 23 above. Therefore, it is rejected for the same reason as claim 23 above.
Claims 24, 33 and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Ringseth and YE, as applied to claims 21, 30 and 38 respectively above, and further in view of LU (US Pub. 2021/0018973 A1).
As per claim 24, Ringseth and YE teach the invention according to claim 21 above. Ringseth teaches execution of workgroups of a given first sub-kernel (Ringseth, Fig. 5, 120 compute engine, 121 compute node; Abstract, provides a tile communication operator that decomposes a computational space into sub-spaces (i.e., tiles) that may be mapped to execution structures (e.g., thread groups) of data parallel compute nodes; [0041] a kernel that is dispatched with respect to thread group dimensions of 16.times.16, which means that 256 threads are logically executing the kernel simultaneously; see Fig, 3, different shaded areas (as different sub-kernel)).
Ringseth and YE fail to specifically teach compression circuitry configured to compress a block of output data generated by execution.
However, LU teaches compression circuitry configured to compress a block of output data generated by execution (LU, Fig. 2, 105 a graphics compression circuit; [0021] the graphics compression circuit 105 is connected in series between the codec overlay hardware of the graphics processor 1021 and the graphics random access memory 1022, and is used to compress the image frames outputted by the graphics processor 1021 and output the compressed image frames to the graphics random access memory 1022).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth and YE with LU because LU’s teaching of providing graphics compression circuit to compress the image frames outputted by the graphics processor would have provided Ringseth and YE’s system with the advantage and capability to reducing the requirements for processing image frames and reduce the power consumption of the graphics compression circuit in order to improving the system performance and efficiency (see LU, [0022]).
As per claim 33, it is a method claim of claim 24 above. Therefore, it is rejected for the same reason as claim 24 above.
As per claim 40, it is a non-transitory computer readable storage medium claim of claim 24 above. Therefore, it is rejected for the same reason as claim 24 above.
Claims 25 and 34 are rejected under 35 U.S.C. 103 as being unpatentable over Ringseth, YE and LU, as applied to claims 24 and 33 respectively above, and further in view of Zhong et al. (US Pub. 2014/0301641 A1).
As per claim 25, Ringseth, YE and LU teach the invention according to claim 24 above. Ringseth, YE and LU fail to specifically teach wherein the sub-kernel limit value corresponds to a sub-kernel shape that provides one or more performance or buffer size parameters for the compression circuitry.
However, Zhong teaches wherein the sub-kernel limit value corresponds to a sub-kernel shape that provides one or more performance or buffer size parameters for the compression circuitry (Zhong, [0026] Experiments conducted by the inventor have also found certain factors useful in selecting a tile size. For example, experiments conducted by the inventor have shown that a tile that covers a span of 16 to 64 pixels is suited for many applications, especially in graphics composition. In many systems, a display sub-system will need to process pixels line-by-line in real time with a given refresh rate. The pixel data is typically stored in line buffers, which are on-chip memories that reside locally near the display sub-system. The line buffers generally will have limited size such that they can hold very few lines of pixels at a time, e.g. 1 or 2 lines. Accordingly, a height of a tile may advantageously be limited to 1 or 2 pixels tall per tile (as sub-kernel limit value corresponds to a sub-kernel shape), e.g. the number of lines in the line buffer. Therefore, a tile size of 8-64 pixels horizontally (e.g. a width of the line buffer) by 1-2 pixels vertically may advantageously make effective use of a line buffer of limited capacity dimension. (as buffer size parameters); [0028] Each tile of an image or a frame may be compressed and decompressed individually. Each compressed tile data may be transferred and written into memory locations that are aligned with subsequence tiles (as for compression; please note: compression circuitry was taught by LU).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth, YE and LU with Zhong because Zhong’s teaching of limit value corresponding to tile shape based on the buffer size for compression would have provided Ringseth, YE and LU’s system with the advantage and capability to improve compression/memory-processing efficiency and accommodate limited buffer capacity (see Zhong, [0025]).
As per claim 34, it is a method claim of claim 25 above. Therefore, it is rejected for the same reason as claim 25 above.
Claims 26 and 35 are rejected under 35 U.S.C. 103 as being unpatentable over Ringseth and YE, as applied to claims 21 and 30 respectively above, and further in view of YUDA (US Pub. 2015/0015571 A1).
YUDA was cited in the IDS filed on 08/21/2025.
As per claim 26, Ringseth and YE teach the invention according to claim 21 above. Ringseth and YE fail to specifically teach workload parser circuitry configured to determine the sub-kernel limit value based on information in a compute command stream that includes the compute kernel.
However, YUDA teaches workload parser circuitry configured to determine the sub-kernel limit value based on information in a compute command stream that includes the compute kernel (YUDA, Fig. 4a to 4b, 401 as first dimension; [0006] lines 1-5, a command receiving unit configured to receive a rendering command indicating a plurality of unit polygons (as whole as compute kernel) which are to be rendered in a rendering region and each of which includes one or more polygons (as information in a compute command stream that includes the compute kernel); [0120] lines 1-7, The divided area setting unit 115 sets a size and a division number for rectangular divided areas. The procedure of divided-area-based rendering is described with reference to FIG. 4A and FIG. 4B. In the divided-area-based rendering, a final rendering region 401 is divided into divided areas 402 (0 to 31) smaller than the final rendering region 401. Then, rendering is performed for each of the divided areas 402; Fig. 11A, Polygon 431, 432, 439, 435, 436, 437, portion of 438 and 433 etc; [0149] lines 1-2, calculating area information for various polygon patterns with reference to FIG. 11A and FIG. 11B; [0150] lines 1-2, All vertices of a polygon 431 are within a final rendering region 401; [0151] line 1, All vertices of a polygon 432 belong to the same divided area [Examiner noted: determine the sub-kernel limit value (i.e., polygons) in the first dimension based on information in a compute command stream that includes the compute kernel (i.e., rendering command indicating a plurality of unit polygons (as whole as compute kernel) which are to be rendered in a rendering region and each of which includes one or more polygons)]).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth and YE with YUDA because YUDA’s teaching of calculating the number of the workgroups (i.e., polygons) for later rendering would have provided Ringseth and YE’s system with the advantage and capability to provide a divided-area-based rendering device that performs area division and reduces an amount of intermediate data used for rendering which improving the system performance and efficiency (see YUDA, [0008]).
As per claim 35, it is a method claim of claim 26 above. Therefore, it is rejected for the same reason as claim 26 above.
Claims 27 and 36 are rejected under 35 U.S.C. 103 as being unpatentable over Ringseth and YE, as applied to claims 21 and 30 respectively above, and further in view of KYO (US Pub. 2013/0024667 A1).
KYO was cited in the IDS filed on 08/21/2025.
As per claim 27, Ringseth and YE teach the invention according to claim 21 above. Ringseth and YE fail to specifically teach wherein the current sub-kernel includes a limited number of workgroups in the second dimension that is smaller than the number of workgroups in the compute kernel in the second dimension.
However, KYO teaches wherein the current sub-kernel includes a limited number of workgroups in the second dimension that is smaller than the number of workgroups in the compute kernel in the second dimension (KYO, Fig. 7, 170 WGs, (1, a, j, s), (5, e, n t) (as current sub-kernel includes a limited number of workgroups in the second dimension, i.e., “a” and “e”, which is smaller than the number of workgroups (i.e., middle group of Lx and Ly which including a, b, c, d and e, f, g, h etc.,).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth and YE with KYO because KYO’s teaching of dividing the dividing the data block with each other according to the contents of the arithmetic processing and the configuration of the OpenCL devices would have provided Ringseth and YE’s system with the advantage and capability to allow the system to processing the different sub-blocks of data among different devices in order to improving the system performance and efficiency (see KYO, [0028]).
As per claim 36, it is a method claim of claim 27 above. Therefore, it is rejected for the same reason as claim 27 above.
Claims 28 and 37 are rejected under 35 U.S.C. 103 as being unpatentable over Ringseth and YE, as applied to claims 21 and 30 respectively above, and further in view of Abramson et al. (US Patent. 5,860,154).
As per claim 28, Ringseth and YE teach the invention according to claim 21 above. Ringseth further teaches wherein the sub-kernel control circuitry is configured to receive the sub-kernel limit value and adjusted coordinates(Ringseth, [0013] decomposes a computational space (represented by indexable_type <N, T> in the embodiment of FIG. 1) into sub-spaces (i.e., tiles) defined by an extent (represented by _Extent in the embodiment of FIG. 1). The tiles may be mapped to execution structures (e.g., thread groups (DirectX), thread blocks (CUDA); [0031] a tile may be set to be a multiple of the local structure constants so that the number of local structures that fit into a tile determines the unrolling factor of the implicit loop dimensions of the execution structure. This relationship may be seen in the local view decomposition: [0036] Assume there is a local view structure determined by 16.times.16 thread group dimensions (as sub-kernel limit value). Analogously, tile the matrices A, B and C into 16.times.16 tiles; [0032] _index=_tileIndex*thread_group_dimensions+localIndex, where _tileIndex represents the coordinates of the tile containing _index (i.e., input indexable type 14) and _localIndex is the offset within that tile (as adjusted coordinates); Fig. 2, 12 tile communication operator).
Ringseth and YE fail to specifically teach wherein the sub-kernel control circuitry is configured to receive the sub-kernel limit value and transmit the adjusted coordinates in a single clock cycle.
However, Abramson teaches wherein the sub-kernel control circuitry is configured to receive the sub-kernel limit value and transmit the adjusted coordinates in a single clock cycle (Abramson, Abstract, A macro instruction is provided for a microprocessor which allows a programmer to specify a base value, index, scale factor and displacement value for calculating an effective address and returning that result in a single clock cycle; also see claims 1-4; ).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth and YE with Abramson because Abramson’s teaching of returning that result in a single clock cycle would have provided Ringseth and YE’s system with the advantage and capability to allow the system to processing the tasks more efficiently which improving the system performance (see Abramson, Col 1, lines 54-62).
As per claim 37, it is a method claim of claim 28 above. Therefore, it is rejected for the same reason as claim 28 above.
Claim 29 is rejected under 35 U.S.C. 103 as being unpatentable over Ringseth and YE, as applied to claim 21 above, and further in view of SATHE et al. (US Pub. 2015/0091913 A1).
SATHE was cited in the IDS filed on 08/21/2025.
As per claim 29, Ringseth and YE teach the invention according to claim 21 above. Ringseth further teaches wherein the apparatus is a computing device that includes (Ringseth, Fig. 5, 100, [0055] Computer system 100 represents any suitable processing device configured for a general purpose or a specific purpose. Examples of computer system 100 include a server, a personal computer, a laptop computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a mobile telephone, and an audio/video device):
a graphics processor (Ringseth, Fig. 5, 122; [0068] where the set of PEs 122 includes one or more GPUs);
a display (Ringseth, Fig. 5, 108);
a central processing unit (Ringseth, Fig. 5, 102; [0059] execution on one or more general purpose processing elements 102 (e.g., central processing units (CPUs)); and
network interface circuitry (Ringseth, Fig. 5, 112; [0066] Network devices 112 include any suitable type, number, and configuration of network devices configured to allow computer system 100 to communicate across one or more networks (not shown). Network devices 112 may operate according to any suitable networking protocol and/or configuration to allow information to be transmitted by computer system 100 to a network or received by computer system 100 from a network).
Ringseth and YE fail to specifically teach a graphics processor that includes the graphics processor circuitry and the sub-kernel control circuitry.
However, SATHE teaches a graphics processor that includes the graphics processor circuitry and the sub-kernel control circuitry (SATHE, Fig. 1, 102 graphic processor, 106 Vertex manager (as sub-kernel control circuitry); Fig. 2, 102 that including vertex shader; [0020] lines 1-5, GPU 102 includes an input assembler 204 and vertex shader 206. In particular, the input assembler 204 may read vertex related data and may assemble the vertex related data into primitives that may be used at subsequent stages of a graphics pipeline).
It would have been obvious to one having ordinary skill in the art before the effective filling date of the claimed invention to have combined the teaching of Ringseth and YE with SATHE because SATHE’s teaching of the GPU processor that including different components/circuity for processing the operations would have provided Ringseth and YE’s system with the advantage and capability to allow the system to improving the graphic operation speed which improving the system efficiency and performance.
Allowable Subject Matter
Claims 22, 31 and 39 are objected to as being dependent upon a rejected base claim, but would be allowable if overcomes the rejections under 35 U.S.C. 101 and 112 (b) and rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Reasons of Allowable Subject Matter:
The closest prior arts of record Ringseth (US Pub. 2012/0089961 A1) teaches a computing system that provides a tile communication operator that decomposes a computational space into sub-spaces (i.e., tiles) that may be mapped to execution structures (e.g., thread groups) of data parallel compute nodes. In addition, the computing system will adjusting the received tile coordinates based on the determined offset values to ensure optimal processing (see Ringseth, Figs. 1, 2, 5; [0013], [0026], [0031-0036]; [00079]).
YE (US Pub. 2022/0076487 A1) teaches a mechanism that allow the CPU to transmitting the coordinate data to the GPU for processing/rendering (see YE, claims 9 and 17).
KYO (US Pub. 2013/0024667 A1) teaches a distribution system that executing kernel in an N-dimensional index space, determining workgroups from the kernel in different dimensions, generating multiple-sub-kernels based on iterate through the multiple dimensions with utilizing the transfer order for designating the order of the respective sub-write blocks/workgroups (i.e., the set of the workgroups (i.e., Fig. 1, each set of workgroups (i.e., respective 170, 160, 164 and 162)) are generated based on the transfer order and executing the workgroups/kernel by the different processing elements (PEs).
The feature “first circuitry configured to determine, based on an increment amount and the sub-kernel limit value for the first dimension, a next position in the first dimension and an increment amount for the second dimension; second circuitry configured to determine, at least partially in parallel with the determination of the next position in the first dimension, next positions in the second dimension for multiple possible increment amounts in the second dimension; and select circuitry configured to: select one of the next positions generated by the second circuitry based on the determined increment amount for the second dimension from the first circuitry; and provide the coordinates of the batch of workgroups to the sub-kernel control circuitry based on the next position in the first dimension and the select next position in the second dimension” when taken in the context of the claims as a whole, were not found in the prior art teachings.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZUJIA XU whose telephone number is (571)272-0954. The examiner can normally be reached M-F 9:30-5:30 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee J Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZUJIA XU/Primary Examiner, Art Unit 2195