Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1- 20 are presented for the examination.
Abstract Objected2. The abstract of the disclosure is objected to because the abstract exceed more than 150 words in length. Correction is required. See MPEP § 608.01(b).
§ 101 2. 35 U.S.C. 101 reads as follows
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
3.Claims 1, 19, 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
As to Claims 1, 19, 20 have been rejected under 35 USC 101 for abstract idea without significantly more. Under Step 2A, Prong 1, the “ dividing the two-dimensional array of values into a plurality of two-dimensional sub-arrays of values ” recite a mental process since “ dividing” is function. This limitation thus describes a “mathematical relationship,” which is specifically identified as an example in the “mathematical concepts” grouping of abstract ideas. Moreover, the recited conversion can be practically performed in the human mind, and so it also falls into the “mental process” group of abstract ideas. Thus, limitation (c) recites a concept that falls into the “mathematical concept” and “mental process” groups of abstract ideas.
Under Prong 2, the additional element “ performing, using a plurality of threads, an initial phase of the separable operation for said sub-array of values in order to generate a respective processed value for each value of said sub-array of values; each of the plurality of threads writing a respective first plurality of processed values to the memory over a plurality of writing steps, said first plurality of processed values corresponding to a one-dimensional sequence of values of said sub-array of values; each of the plurality of threads reading a respective second plurality of processed values from the memory over a plurality of reading steps, said second plurality of processed values corresponding to a perpendicular one-dimensional sequence of values of a sub-array of values in a transposed position within the array of values relative to said sub-array of values; and performing, using the plurality of threads, a subsequent phase of the separable operation for the plurality of processed values read by the plurality of threads in order to generate a respective output value for each value of the sub-array of values in the transposed position; wherein a respective processed value is written into each of the memory banks of the memory in at least one of the plurality of writing steps, and a respective processed value is read from each of the memory banks of the memory in at least one of the plurality of reading steps.” are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component, or merely a generic computer or generic computer components to perform the judicial exception, Accordingly, the additional elements do not integrate the recited judicial exception into a practical application, and the claim is therefore directed to the judicial exception. See MPEP 2106.05(f).
Under Step 2B, the additional elements “ performing, using a plurality of threads, an initial phase of the separable operation for said sub-array of values in order to generate a respective processed value for each value of said sub-array of values; each of the plurality of threads writing a respective first plurality of processed values to the memory over a plurality of writing steps, said first plurality of processed values corresponding to a one-dimensional sequence of values of said sub-array of values; each of the plurality of threads reading a respective second plurality of processed values from the memory over a plurality of reading steps, said second plurality of processed values corresponding to a perpendicular one-dimensional sequence of values of a sub-array of values in a transposed position within the array of values relative to said sub-array of values;” - this generally have been a mental process although the sub-array could be a generic computer component if the spec describes it as actual computer hardware.
“performing, using the plurality of threads, a subsequent phase of the separable operation for the plurality of processed values read by the plurality of threads in order to generate a respective output value for each value of the sub-array of values in the transposed position; wherein a respective processed value is written into each of the memory banks of the memory in at least one of the plurality of writing steps, and a respective processed value is read from each of the memory banks of the memory in at least one of the plurality of reading steps” - this is mere instructions to apply the mental process under mpep 2106.05(f), amounts to merely generally linking the use of the judicial exception to a particular technological environment or field or use, and is merely applying the judicial exception, therefore, does not amount to significantly more, hence, cannot provide an inventive concept.
6. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application. See MPEP 2106.05(d). Thus, the claim is not patent eligible.
DETAILED ACTION
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 19, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) and further in view of Kondo( US 20050180221 A1).
As to claim 1, Li teaches the memory comprising a plurality of memory banks, wherein in each writing or reading step each memory bank can be written into or read from by only one respective thread( each thread to access a different memory bank in the memory unit, para[0007], ln 8-10/ All threads may access anywhere in a common area in the memory unit. In one embodiment, the common area may be a shared memory space that includes all memory banks, para[0077]/ the address vector, read data vector and write data vector may be simply split into each memory bank with a one-to-one mapping so that the data of different threads may be mapped into different memory banks. For example, the i-th address in the vector address may be for thread i , para[0064], ln 3-11)
the method comprising: dividing the two-dimensional array of values into a plurality of two-dimensional sub-arrays of values for each of the plurality of sub-arrays ( the PE array 214 may be a two-dimensional array that may comprise multiple rows of processing elements[two-dimensional array of values] (e.g., one or more rows of PEs may be positioned underneath the PEs 218.1-218.N). It should be noted that the PE array 214 may be a composite of MPs, SBs, ICSBs and PEs for illustration purpose and used to refer to these components collectively, para[0044], ln 12-18/ 2-D data path (including a processing element (PE) array and interconnections) to process massive parallel data. Each data path may be segmented into sections[a plurality of two-dimensional sub-arrays of values]. In a 1-D data path, a section may include a memory port, a switch box, a PE and an ICSB[a plurality of two-dimensional sub-arrays of values] in one column; and in a 2-D data path, a section may include a memory port, two or more switch boxes, two or more PEs and an ICSB[a plurality of two-dimensional sub-arrays of values] in one column. Para[0100], ln 14-21/ a physical data path (PDP) may be made to have repetitive structure. For example, each column may be identical and each PDP may comprise same amount of repetitive columns. As shown in FIG. 9C, the VDP of FIG. 9B may be divided into three PDPs (e.g., PDP1, PDP2 and PDP3) for a 2×2 PE array and thus, the three PDPs may have the same structure. The 2×2 PE array may be the whole PE array of an embodiment of a RPP, or may be part of a N×N (e.g., N being 32 as an example) PE array of another embodiment of a RPP. There may be many connections between PEs in one PDP (e.g., A to B and C, B to D, C to D, etc.), between two consecutive PDPs (e.g., D to E and F, G to J, H to J, F to I, etc.) and between non-consecutive PDPs (e.g., B to K), para[0101], ln 5-20).
Nickolls teaches performing, using a plurality of threads, an initial phase of the separable operation for said sub-array of values in order to generate a respective processed value for each value of said sub-array of values( an array of data values (e.g., pixels) can be filtered using a 2-D kernel-based filter algorithm, in which the filtered value of each pixel is determined based on the pixel and its neighbors. In some instances the filter is separable and can be implemented by computing a first pass along the rows of the array to produce an intermediate array, then computing a second pass along the columns of the intermediate array. In one CTA implementation of a separable 2-D filter, the threads of the CTA load the input data set (or a portion thereof) into shared memory 306, then synchronize. Each thread performs the row-filter for one point of the data set and writes the intermediate result to shared memory 306, para[0054], ln 1-18).
a separable operation on a two-dimensional array of values at a processing unit comprising a memory( an array of data values (e.g., pixels) can be filtered using a 2-D kernel-based filter algorithm, in which the filtered value of each pixel is determined based on the pixel and its neighbors. In some instances the filter is separable and can be implemented by computing a first pass along the rows of the array to produce an intermediate array, then computing a second pass along the columns of the intermediate array, para[0054], ln 1-10);
each of the plurality of threads writing a respective first plurality of processed values to the memory over a plurality of writing steps( In one CTA implementation of a separable 2-D filter, the threads of the CTA load the input data set (or a portion thereof) into shared memory 306, then synchronize. Each thread performs the row-filter for one point of the data set[values] and writes the intermediate result to shared memory 306, para[0054], ln 8-14)
each of the plurality of threads reading a respective second plurality of processed values from the memory over a plurality of reading steps( and a thread may read row-filter results[values] that were written by any thread of the CTA, para[0054], ln 18-21)
said second plurality of processed values corresponding to a perpendicular one-dimensional sequence of values of a sub-array of values in a transposed position within the array of values relative to said sub-array of values( After all threads have written their row-filter results to shared memory 306 and have synchronized at that point, each thread performs the column filter for one point of the data set. In the course of performing the column filter, each thread reads the appropriate row-filter results from shared memory 306, and a thread may read row-filter results that were written by any thread of the CTA, para[0054], ln 10-20);
and performing, using the plurality of threads, a subsequent phase of the separable operation for the plurality of processed values read by the plurality of threads in order to generate a respective output value for each value of the sub-array of values in the transposed position( in one CTA implementation of a separable 2-D filter, the threads of the CTA load the input data set (or a portion thereof) into shared memory 306, then synchronize. Each thread performs the row-filter for one point of the data set and writes the intermediate result to shared memory 306. After all threads have written their row-filter results to shared memory 306 and have synchronized at that point, each thread performs the column filter for one point of the data set. In the course of performing the column filter, each thread reads the appropriate row-filter results from shared memory 306, and a thread may read row-filter results that were written by any thread of the CTA. The threads write their column-filter results to shared memory 306, para[0054], ln 7-20).
It would have been obvious to one of the ordinary skill in the art before the effective filling of claimed invention was made to modify the teaching of Li with Nickolls
to incorporate above feature because this provides methods for scalably exploiting parallelism in a parallel processing subsystem. A problem to be solved is hierarchically decomposed into at least two levels of sub-problems .
Kondo teaches said first plurality of processed values corresponding to a one-dimensional sequence of values of said sub-array of values( divides a two-dimensional array of data into vertical one-dimensional or horizontal one-dimensions data groups, and changes the sequence of storage into each bank of the memory for each data group, right col 17, ln 20-25/ Two-dimensionally arrayed image data ID shown in FIG. 57A is divided into vertical stripe-like groups as shown in FIG. 57B, para[0304], ln 1-3/ FIG. 58A shows a two-dimensional reference area and reference data distribution in the reference area by way of example. It is assumed as in FIG. 57 that the reference data is divided into vertical stripes as shown in FIG. 58B. The stripe is vertically compressed as shown in FIG. 58C. If even one reference data exists in the stripe, a flag is set, which means that the data can be stored as in the method of storing one-dimensionally arrayed data, para[0306]),
wherein a respective processed value is written into each of the memory banks of the memory in at least one of the plurality of writing steps, and a respective processed value is read from each of the memory banks of the memory in at least one of the plurality of reading steps(Two-dimensionally arrayed image data ID shown in FIG. 57A is divided into vertical stripe-like groups as shown in FIG. 57B, and data groups on one stripe are stored onto one word line in a memory bank as shown in FIG. 57C. By selecting a memory bank for each stripe-like data group according to the distribution of reference data at this time, it is possible to read two-dimensionally arrayed data simultaneously, para[0304]/ the above object can be attained by providing a data storage controller which stores data into a memory including a plurality of memory banks and reads a plurality of desired data simultaneously from the memory, para[0035]/ the data being divided among the plurality of memory banks of the memory, whether the data going to be stored are ones at locations corresponding to the access pattern, para[0039], ln 3-7/ a judging means for judging, based on an access pattern representing a plurality of desired data to be read simultaneously when storing data sequentially into the memory with the data being divided among the plurality of memory banks of the memory, whether the data going to be stored are ones at locations corresponding to the access patterns; and a memory controlling means for storing all data at the locations corresponding to the access pattern into different memory banks by skipping a memory bank in which the data are to be stored, by incrementing its address, when the data to be read simultaneously are ones at the locations corresponding to the access pattern, and storing data at the locations corresponding to the access pattern into the memory bank having the bank address thereof incremented , para[0033], ln 5-9 to para[0034], ln 1-9/ hanging the sequence of storage into each bank of the memory according to a storage pattern corresponding to the distribution of the data to be read simultaneously, thereby storing the data to be read simultaneously into different banks, respectively, para[0077]).
It would have been obvious to one of the ordinary skill in the art before the effective filling of claimed invention was made to modify the teaching of Li and Nickolls with Kondo
to incorporate above feature because this is necessary to simultaneously read a plurality of desired data meeting a purpose from image data stored in the semiconductor memory.
As to claims 19, 20, they are rejected for the same reason as to claim 1 above. In additional, Li teaches memory, non-transitory computer readable storage medium(computer-executable instructions stored on one or more computer-readable storage media. The one or more computer-readable storage media may include non-transitory computer-readable media (such as removable or non-removable magnetic disks, magnetic tapes or cassettes, solid state drives (SSDs), hybrid hard drives, CD-ROMs, CD-RWs, DVDs, or any other tangible storage medium), volatile memory components (such as DRAM or SRAM), or nonvolatile memory components (such as hard drives)), para[0142], ln 3-20).
Claim(s) 2, 3, 4, 8, 13, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) in view of Kondo( US 20050180221 A1) and further in view of Kapsenberg(US 20220182583 A1).
As to claim 2, Kapsenberg teaches the memory bank into which a processed value is to be written is determined in dependence on a write buffer array having a number of elements greater than the number of elements in the two-dimensional array of values ( As the memory buffer 453 receives the data elements of the image frame, the memory buffer 453 may write the data elements to a memory array in the memory buffer 453. The memory array may include memory locations or cells configured to store a single bit or multiple bits of image data (e.g., image or pixel data). The memory buffer 453 may include any suitable number of memory arrays configured to store the data elements of image frames. The memory locations of the array may be arranged as a two-dimensional matrix or array of random access memory (RAM) or other addressable memory. The number of memory locations of the array (e.g., image buffer size) may be equal to the number of data elements in each image frame in the sequence of image frames. In exemplary implementations, the number of memory locations may be greater than the number of data elements in each image frame in the sequence, para[0089]/ The image frame may represent a two-dimensional image having rows of data elements and columns of data elements., para[0119], ln 24-27).
It would have been obvious to one of the ordinary skill in the art before the effective filling of claimed invention was made to modify the teaching of Li, Nickolls and Kondo with Kapsenberg to incorporate above feature because this enables the transfer of the data elements of an image frame (e.g., image data) between the various components and systems.
As to claim 3, Kapsenberg teaches wherein the write buffer array comprises: value elements corresponding to values of the two-dimensional array; and padding elements( para[0085], ln 1-15/ para[0087], ln 1-15/ para[0089], ln 1-10/para[0088], ln 1-20) for the same reason as to claim 2 above.
As to claim 4, Kapsenberg teaches the padding elements correspond to memory padding( para[0088], ln 3-14) for the same reason as to claim 2 above.
As to claim 8, Kapsenberg teaches the write buffer array is mapped to the memory such that processed values corresponding to values of the two-dimensional array are written into the memory in memory locations to which the value elements of the write buffer array are mapped, and processed values corresponding to values of the two-dimensional array are not written to the memory in memory locations to which the padding elements of the write buffer array are mapped( para[0089]/ para[0087], ln 1-15) for the same reason as to claim 2 above.
As to claim 9, Kapsenberg teaches the memory location into which a thread writes a processed value is determined in dependence on a base memory address, a writing offset amount and a writing padding amount, the writing offset amount and the writing padding amount being dependent on the position of the value within the array of values to which that processed value corresponds( para[0090], ln 1-20).
As to claim 13, Kapsenberg teaches the memory location from which a thread reads a processed value is determined in dependence on a base memory address, a reading offset amount and a reading padding amount, the reading offset amount and the reading padding amount being dependent on the position of the value within the array of values to which that processed value corresponds( para[0098], ln 1-14) for the same reason as to claim 2 above.
As to claim 18, Kapsenberg teaches the one dimensional sequence of values of said sub-array of values is a row of values of said sub-array of values and the perpendicular one dimensional sequence of values of the sub-array of values in the transposed position is a column of values of the sub-array of values in the transposed position; or the one dimensional sequence of values of said sub-array of values is a column of values of said sub-array of values and the perpendicular one dimensional sequence of values of the sub-array of values in the transposed position is a row of values of the sub-array of values in the transposed position( para[0085], ln 5-18/ para[0090], ln 2-18) for the same reason as to claim 2 above.
Claim(s) 15 is rejected under 35 U.S.C. 103 as being unpatentable over Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) in view of Kondo( US 20050180221 A1) and further in view of Crinon( US 4800443 A),
As to claim 15, Crinon teaches wherein the plurality of two-dimensional sub-arrays of values are non-overlapping( wherein the two-dimensional array of threshold values is composed of a plurality of identical two-dimensional sub-arrays in non-overlapping relationship, col 13, ln 50-61 to col 14, ln 30-31).
It would have been obvious to one of the ordinary skill in the art before the effective filling of claimed invention was made to modify the teaching of Li, Nickolls and Kondo with Kapsenberg to incorporate above feature because this applies read/write enable signals to the buffer, such that during each active line interval one row of the buffer is in the write state and the other five rows are in the read state.
Claim(s) 16 is rejected under 35 U.S.C. 103 as being unpatentable over Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) in view of Kondo( US 20050180221 A1) and further in view of Briongos( US 20230168954 A1).
As to claim 16, Briongos teaches the plurality of threads are processed by processing logic comprised by a core of the processing unit, the processing logic being implemented on a chip and the memory being physically located on the same chip as the processing logic( According to the terminology of the present disclosure, “cache memory associated with a core” means a cache memory comprised in a processor comprising the core and that the cache memory shares the same physical chip with the core (e.g., the cache memory is physically on the same substrate as the core of the processor executing at least one thread of each of the enclaves), para[0089], ln 15-21).
It would have been obvious to one of the ordinary skill in the art before the effective filling of claimed invention was made to modify the teaching of Li, Nickolls and Kondo with Kapsenberg to incorporate above feature because this enables fast communication between two enclaves on the same platform.
Claim(s) 17 is rejected under 35 U.S.C. 103 as being unpatentable over Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) in view of Kondo( US 20050180221 A1) and further in view of Jackson(US 20070101072 A1).
As to claim 17, Jackson teaches the two-dimensional array of values is a two-dimensional array of pixel values( FIG. 5 is a flowchart showing the operations performed by one embodiment of the present invention in writing to variable-resolution memory (VRM). As previously described, image data can be written to the VRM as a two-dimensional array of pixel values at full-resolution and read in accordance with implementations of the present invention, para[0049], ln 1-7).
It would have been obvious to one of the ordinary skill in the art before the effective filling of claimed invention was made to modify the teaching of Li, Nickolls and Kondo with Kapsenberg to incorporate above feature because this reduces the processor cycles available to perform other tasks or requires more powerful processors with increased device complexity, power consumption and cost.
Allowable Subject Matter
Claims 5-8, 10, 11, 12, 14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Reasons for Allowance
As to claims 5-8, Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) in view of Kondo( US 20050180221 A1) do not teach the write buffer array comprises groups of contiguous value elements corresponding to values of the two-dimensional array, and padding elements interspersed between said groups. The method of claim 5, wherein the number of value elements in each group is (i) equal to, (ii) a multiple of, or (iii) a factor of the number of memory banks comprised by the memory , The method of claim 5, wherein the number of value elements in each group is equal to or less than the number of threads used to perform the separable operation for the array of values.
As to claim 8, Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) in view of Kondo( US 20050180221 A1) do not teach the write buffer array is mapped to the memory such that processed values corresponding to values of the two-dimensional array are written into the memory in memory locations to which the value elements of the write buffer array are mapped, and processed values corresponding to values of the two-dimensional array are not written to the memory in memory locations to which the padding elements of the write buffer array are mapped.
As to claim 10-12, 14, Li( US 20180267931 A1) in view of Nickolls( US 20110238955 A1) and in view of Kondo( US 20050180221 A1) do not teach the two-dimensional array of values divided into a plurality of two-dimensional sub-arrays of values is represented by a multidimensional array [I][J][K][M], where I and J represent the number of sub-arrays of values within the array of values in each of the two dimensions and K and M represent the number of values within each of the sub-arrays of values in each of the two dimensions, and wherein each value in the array of values has a coordinate [i][j][k][m] which defines its position within the multidimensional array [l][J][K][M], and wherein:the writing offset amount is equal to (i xJxKxM)+(jxKxM)+(kxM)+mor(ixJxKxM)+ (j x Kx M) + (mxK)+ k; and the writing padding amount is equal to the writing offset amount divided by a padding frequency. The method of claim 10, wherein the padding frequency is (i) equal to, (ii) a multiple of, or (iii) a factor of the number of memory banks comprised by the memory. wherein the padding frequency is equal to or less than the number of threads used to perform the separable operation for the array of values.
Response to the argument:
A. Applicant amendment filed on 07/09/2026 has been considered but they are not persuasive:
Applicant argued in substance that :
“Each data path may be segmented into sections. In a 2-D data path, a section may include a memory port, two or more switch boxes, two or more PEs and an ICSB in one column. Segmenting a data path comprising a processing element (PE) array and interconnections into sections is not the same as dividing a two-dimensional array of values into a plurality of two-dimensional sub-arrays of values, as claimed in claim 1. Thus, the mapping of Li onto this claim feature is incorrect.”
“ each thread performs the row-filter for one point of the data set and writes the intermediate result [i.e. singular, not plural] to shared memory 306. That is, in Nickolls, each thread writes one, singular, intermediate result to shared memory, and not a respective first plurality of processed values, as claimed in claim 1. Thus, the mapping of Nickolls onto this claim feature is incorrect.”
“There is no disclosure in Nickolls that the row-filter results read by a thread correspond to a perpendicular one- dimensional sequence of values of a sub-array of values in a transposed position within the array of values relative to said sub-array of values, as claimed in claim 1. Nickolls in fact is silent with respect to this feature. Thus, the mapping of Nickolls onto this claim feature is also incorrect.”
“It cannot be said that Kondo (filed in 2005) is further defining the intermediate result of Nickolls (filed in 2011). In fact, it does not even make any technological sense to assert that the one, singular, intermediate result written to shared memory by each thread in paragraph [0054] of Nickolls corresponds to a vertical stripe-like group (i.e. plurality) of data as introduced in paragraphs [0304] and [0306]”
“Kondo discloses simultaneously reading desired data from a memory. However, Kondo is silent with respect to simultaneously writing desired data into a plurality of memory banks of that memory. Instead, Kondo describes storing (e.g. writing) data sequentially into the memory with the data being divided among the plurality of memory banks of the memory (see paragraph [0039]) “.
“ Claims 1, 19 and 20 have been rejected under 35 U.S.C. 101 as allegedly being directed the abstract idea of a mental process. Applicant respectfully traverses this rejection”.
B. Examiner respectfully disagreed with Applicant's remarks:
As to the point(1), There were no explanation how the segmenting a data path comprising a processing element (PE) array and interconnections into sections is not the same as dividing a two-dimensional array of values into a plurality of two-dimensional sub-arrays of values.
Li teaches the PE array 214 may be a two-dimensional array that may comprise multiple rows of processing elements[two-dimensional array of values] (e.g., one or more rows of PEs may be positioned underneath the PEs 218.1-218.N). It should be noted that the PE array 214 may be a composite of MPs, SBs, ICSBs and PEs for illustration purpose and used to refer to these components collectively, para[0044], ln 12-18/ 2-D data path (including a processing element (PE) array and interconnections) to process massive parallel data. Each data path may be segmented into sections[a plurality of two-dimensional sub-arrays of values]. In a 1-D data path, a section may include a memory port, a switch box, a PE and an ICSB[a plurality of two-dimensional sub-arrays of values] in one column; and in a 2-D data path, a section may include a memory port, two or more switch boxes, two or more PEs and an ICSB[a plurality of two-dimensional sub-arrays of values] in one column. Para[0100], ln 14-21/ a physical data path (PDP) may be made to have repetitive structure. For example, each column may be identical and each PDP may comprise same amount of repetitive columns. As shown in FIG. 9C, the VDP of FIG. 9B may be divided into three PDPs (e.g., PDP1, PDP2 and PDP3) for a 2×2 PE array and thus, the three PDPs may have the same structure. The 2×2 PE array may be the whole PE array of an embodiment of a RPP, or may be part of a N×N (e.g., N being 32 as an example) PE array of another embodiment of a RPP. There may be many connections between PEs in one PDP (e.g., A to B and C, B to D, C to D, etc.), between two consecutive PDPs (e.g., D to E and F, G to J, H to J, F to I, etc.) and between non-consecutive PDPs (e.g., B to K), para[0101], ln 5-20).
As to the point(2), Nickolls teaches an array of data values (e.g., pixels) can be filtered using a 2-D kernel-based filter algorithm, in which the filtered value of each pixel is determined based on the pixel and its neighbors. In some instances the filter is separable and can be implemented by computing a first pass along the rows of the array to produce an intermediate array, then computing a second pass along the columns of the intermediate array, para[0054], ln 1-10/ In one CTA implementation of a separable 2-D filter, the threads of the CTA load the input data set (or a portion thereof) into shared memory 306, then synchronize. Each thread performs the row-filter for one point of the data set[values] and writes the intermediate result to shared memory 306, para[0054], ln 8-14).
As to the point(3), Nickolls teaches After all threads have written their row-filter results to shared memory 306 and have synchronized at that point, each thread performs the column[transposed position] filter for one point of the data set. In the course of performing the column filter, each thread reads the appropriate row-filter results from shared memory 306, and a thread may read row-filter results that were written by any thread of the CTA. The threads write their column-filter results to shared memory 306. The resulting data array can be stored to global memory or retained in shared memory 306 for further processing, para[0054], ln 10-27).
As to the point(4), Both references of Nickolls and Kondo teach the same field of separate data related to 2 dimension to write and read data to the memory, Mickolls teaches thread write and read values to and from the memory ( in one CTA implementation of a separable 2-D filter, the threads of the CTA load the input data set (or a portion thereof) into shared memory 306, then synchronize. Each thread performs the row-filter for one point of the data set and writes the intermediate result to shared memory 306. After all threads have written their row-filter results to shared memory 306 and have synchronized at that point, each thread performs the column filter for one point of the data set. In the course of performing the column filter, each thread reads the appropriate row-filter results from shared memory 306, and a thread may read row-filter results that were written by any thread of the CTA. The threads write their column-filter results to shared memory 306, para[0054], ln 7-20).
Kondo teaches value is write and read to and from memory bank( Two-dimensionally arrayed image data ID shown in FIG. 57A is divided into vertical stripe-like groups as shown in FIG. 57B, and data groups on one stripe are stored onto one word line in a memory bank as shown in FIG. 57C. By selecting a memory bank for each stripe-like data group according to the distribution of reference data at this time, it is possible to read two-dimensionally arrayed data simultaneously, para[0304]/ the above object can be attained by providing a data storage controller which stores data into a memory including a plurality of memory banks and reads a plurality of desired data simultaneously from the memory, para[0035]/ the data being divided among the plurality of memory banks of the memory, whether the data going to be stored are ones at locations corresponding to the access pattern, para[0039], ln 3-7/ a judging means for judging, based on an access pattern representing a plurality of desired data to be read simultaneously when storing data sequentially into the memory with the data being divided among the plurality of memory banks of the memory, whether the data going to be stored are ones at locations corresponding to the access patterns; and a memory controlling means for storing all data at the locations corresponding to the access pattern into different memory banks by skipping a memory bank in which the data are to be stored, by incrementing its address, when the data to be read simultaneously are ones at the locations corresponding to the access pattern, and storing data at the locations corresponding to the access pattern into the memory bank having the bank address thereof incremented , para[0033], ln 5-9 to para[0034], ln 1-9/ hanging the sequence of storage into each bank of the memory according to a storage pattern corresponding to the distribution of the data to be read simultaneously, thereby storing the data to be read simultaneously into different banks, respectively, para[0077]).
As to the point 5, The claims recite the data is read and write to memory but do not require whether the data is simultaneously reading or sequentially writing into the memory.
As to the point 6, As to Claims 1, 19, 20 have been rejected under 35 USC 101 for abstract idea without significantly more. Under Step 2A, Prong 1, the “ dividing the two-dimensional array of values into a plurality of two-dimensional sub-arrays of values ” recite a mental process since “ dividing” is function. This limitation thus describes a “mathematical relationship,” which is specifically identified as an example in the “mathematical concepts” grouping of abstract ideas. Moreover, the recited conversion can be practically performed in the human mind, and so it also falls into the “mental process” group of abstract ideas. Thus, limitation (c) recites a concept that falls into the “mathematical concept” and “mental process” groups of abstract ideas.
Under Prong 2, the additional element “ performing, using a plurality of threads, an initial phase of the separable operation for said sub-array of values in order to generate a respective processed value for each value of said sub-array of values; each of the plurality of threads writing a respective first plurality of processed values to the memory over a plurality of writing steps, said first plurality of processed values corresponding to a one-dimensional sequence of values of said sub-array of values; each of the plurality of threads reading a respective second plurality of processed values from the memory over a plurality of reading steps, said second plurality of processed values corresponding to a perpendicular one-dimensional sequence of values of a sub-array of values in a transposed position within the array of values relative to said sub-array of values; and performing, using the plurality of threads, a subsequent phase of the separable operation for the plurality of processed values read by the plurality of threads in order to generate a respective output value for each value of the sub-array of values in the transposed position; wherein a respective processed value is written into each of the memory banks of the memory in at least one of the plurality of writing steps, and a respective processed value is read from each of the memory banks of the memory in at least one of the plurality of reading steps.” are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component, or merely a generic computer or generic computer components to perform the judicial exception, Accordingly, the additional elements do not integrate the recited judicial exception into a practical application, and the claim is therefore directed to the judicial exception. See MPEP 2106.05(f).
Under Step 2B, the additional elements “ performing, using a plurality of threads, an initial phase of the separable operation for said sub-array of values in order to generate a respective processed value for each value of said sub-array of values; each of the plurality of threads writing a respective first plurality of processed values to the memory over a plurality of writing steps, said first plurality of processed values corresponding to a one-dimensional sequence of values of said sub-array of values; each of the plurality of threads reading a respective second plurality of processed values from the memory over a plurality of reading steps, said second plurality of processed values corresponding to a perpendicular one-dimensional sequence of values of a sub-array of values in a transposed position within the array of values relative to said sub-array of values;” - this generally have been a mental process although the sub-array could be a generic computer component if the spec describes it as actual computer hardware.
“performing, using the plurality of threads, a subsequent phase of the separable operation for the plurality of processed values read by the plurality of threads in order to generate a respective output value for each value of the sub-array of values in the transposed position; wherein a respective processed value is written into each of the memory banks of the memory in at least one of the plurality of writing steps, and a respective processed value is read from each of the memory banks of the memory in at least one of the plurality of reading steps” - this is mere instructions to apply the mental process under mpep 2106.05(f), amounts to merely generally linking the use of the judicial exception to a particular technological environment or field or use, and is merely applying the judicial exception, therefore, does not amount to significantly more, hence, cannot provide an inventive concept.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Conclusion
US 20110238955 A1 teaches In one CTA implementation of a separable 2-D filter, the threads of the CTA load the input data set (or a portion thereof) into shared memory , then synchronize. Each thread performs the row-filter for one point of the data set and writes the intermediate result to shared memory .
US 7584342 B1 A teaches two-dimensional (N.times.N) separable filter is a convolution filter in which the filter can be expressed in terms of a "row" vector and a "column" vector. The row vector is applied to each row as a convolution filter with a kernel that is one row high and N columns wide, and the column vector is applied to the row-filter results as a convolution filter with a kernel that is one column wide and N rows high.
CN 104077233 A teaches (Graphic Processing Unit, GPU) thread into two-dimensional format X dimension is divided by number of all image, Y-dimension is divided by number of all filter, each graphics processor thread is responsible for calculating the plurality of filters to the convolution of multiple image, but only the one data point corresponding to the convolution kernel.
CN 103139591 B teaches using the calculated DirextX11 Shader main idea for mean and variance statistics of the image, adopting dimension reduction finally, all the pixels in the one-dimensional image counting operation can be finished by one thread.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LECHI TRUONG whose telephone number is (571)272-3767. The examiner can normally be reached 10-8 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Young Kevin can be reached on (571)270-3180. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LECHI TRUONG/Primary Examiner, Art Unit 2194