DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1-7, and 9-20 are currently pending in the present application, with claims 1, 7, 11, and 20 being independent.
Response to Amendments / Arguments
Applicant’s arguments, see Pg. 8, filed 04/23/2026 with respect to claims 1, 4-6, 11, 14-15, and 20 have been fully considered and are persuasive. The claim objections of claims 1, 4-6, 11, 14-15, and 20 has been withdrawn.
Applicant’s arguments, see Pg. 8-10, filed 04/23/2026, with respect to claims 1-6, and 9-20 have been fully considered and are persuasive. The 35 U.S.C. §112(b) rejections of claims 1-6 and 9-20 has been withdrawn.
Applicant's arguments filed 04/23/2026 have been fully considered but they are not persuasive.
Applicant argues: The applied references, NVIDIA “NVIDIA PTX ISA: NVIDIA CUDA Programming Guide”, Version 8.4 published on March 2, 2024, by NVIDIA Corporation, and The Dude et al. (2017, March 27). Parallel ray tracing in 16x16 chunks. Stack Overflow. https://stackoverflow.com/questions/43056609/parallel-ray-tracing-in-16x16-chunks, hereinafter referred to as “The Dude”, do not teach the claimed thread-group allocation and per-thread subregion processing in combination with Jacobson et al., "Spatial Adaptive Sampling in Real Time Ray Tracing", Department of Computer Science, Lund University, (2021), pages 1-67, hereinafter referred to as “Jacobson”
Examiner replies: In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, Jacobsen is relied upon for the ray-tracing and adaptive sampling process, where Jacobsen expressly discloses a ray tracer that traces rays through a scene and uses an adaptive sampler that performs repeated “schedule-shoot-interpolate” passes, including scheduling new rays, shooting rays, and interpolating pixels until the complete image is generated (Pg. 34-36, Section 4.2 and Fig. 4.6). Jacobsen further discloses that different pixels are ray traced in different stages (Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes…The figure starts out with a 9x9 grid, then 5x5 grid, and lastly a 3x3 grid).
NVIDIA is then incorporated for the conventional GPU thread-group organization used to implement parallel graphics workloads. NVIDIA expressly discloses that a “CTA is an array of threads that execute a kernel concurrently or in parallel,” and that each CTA thread uses a thread identifier to determine its assigned role, assigned input/output positions, compute addresses, and select work to perform (Pg. 11, Section 2.2.1; CTA, is an array of threads that execute a kernel concurrently or in parallel…Each CTA thread uses its thread identifier to determine its assigned role, assign specific input and output positions, compute addresses, and select work to perform. The thread identifier is a three-element vector tid, (with elemenets tid.x, tid.y, and tid.z) that specifies the thread's position within a 1D, 2D, or 3D CTA…). Thus, NVIDIA is applied to organize Jacobsen’s ray-tracing work into groups of parallel threads.
The Dude is incorporated after Jacobsen’s ray-tracing process has been placed into the NVIDIA-style parallel thread framework, expressly disclosing a known task-assignment feature ofassigning portions of an image to worker threads. The Dude teaches that a manager object returns “a chunk of work” each time a thread calls GetTask(), and explains that the task may be a row or a 16x16 block. Thus, The Dude is applied to assign each thread a subregion/chunk of the render output and to have the thread execute the ray-tracing task for that assigned subregion (Response from Adrian McCarthy's; Manager::GetTask()…t = std::make_unique<Task>(next_row);…return t;…the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like.) When all the tasks have been issued, it just returns an empty pointer, which essentially tells the calling thread that there's nothing left to do, and the calling thread will then exit). Thus, The Dude is used for the per-thread subregion/task assignment missing from Jacobsen and NVIDIA
Accordingly, combining these teachings would have been a routine use of known parallel ray-tracing techniques to improve load balancing, thread utilization, and rendering efficiency by implementing Jacobsen’s ray-tracing process in a parallel thread-group architecture, as taught by NVIDIA, with per-thread work distribution, as taught by The Dude.
Applicant argues: It is improper to base a conclusion of obviousness upon facts gleaned only through hindsight. "To draw on hindsight knowledge of the patented invention, when the prior art does not contain or suggest that knowledge, is to use the invention as a template for its own reconstruction- an illogical and inappropriate process by which to determine patentability." Sensonics, Inc. v. Aerosonic Corp., 81 F.3d 1566, 1570 (Fed. Cir. 1996) (citing W.L. Gore & Assoc, v. Garlock, Inc., 721 F.2d 1540, 1553 (Fed. Cir. 1983)).
Examiner replies: In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971).
Applicant argues: Jacobsen does not disclose “determining, for different regions of the render output, a relative number of rays to be traced for the region of the render output… and “determining the number of rays to be traced…based on the relative number of rays to be traced for the region of the render output and the ray tracing budget number (B) for the render output,” wherein the ray tracing budget (B) corresponds “to a total number of rays that can be traced when generating the render output,” and further asserts that that the office appears to conflate the two different “modes” discussed in Jacobsen. Jacobsen’s “1 spp budget” relates only to the “basic mode,” while Jacobsen’s adaptive sampling mode merely reduces the number of traced pixels by interpolation.
Examiner replies: Jacobsen expressly discloses that the real-time ray tracer operates under a sample/ray limit. Jacobsen states that the “basic mode” of the ray tracer “uses 1 spp” (Pg. 21, Section 3.3.1; default mode of the ray tracer…uses 1 spp), and further states that “the budget for a real time tracer is limited to that of ca 1 spp, and explains that in the ray tracing system, each pixel ray “bounces seven times and samples the light source” (Pg. 30, Section 4.1.2; The budget for a real time tracer is limited to that of ca 1 spp, it will generate a high frequency noisy ray traced image…where each pixel ray bounces seven times and samples the light source…), confirming that the spp value corresponds to rays/samples traced in generating the render output. Under broadest reasonable, the claimed “ray tracing budget number (B)” is not limited to a particular stored variable named B, nor a preferred budgeting algorithm. Rather, it broadly encompasses a numerical limit or target amount of rays/samples available for generating the frame. Therefore, Jacobsen’s approximately 1 spp budget defines the available ray count as approximately one ray/sample for each sampling position of the output.
Applicant’s distinction between Jacobsen’s basic mode and adaptive mode does not negate Jacobsen’s express disclosure of a ray tracing budget. The rejection relies on Jacobsen’s 1 spp disclosure to teach the claimed budget number (B) i.e., total available ray/sample capacity for generating a real-time render output. Therefore, Jacobsen’s disclosure expressly teaches the claimed “ray tracing budget number (B)”.
Applicant argues: Jacobsen does not disclose “determining, for different regions of the render output, a relative number of rays to be traced for the region of the render output…” and further asserts that Jacobsen merely determines whether individual pixels are interpolatable, rather than assigning relative regional ray counts.
Examiner replies: The argument is not persuasive because the claims do not require the “relative number” to be a specific stored ratio or the exact values as shown in Fig. 9 of applicant’s originally filed disclosure. The claim broadly recites determining, for different regions, a relative number of rays for the region. Jacobsen expressly discloses that different portions of the image are sampled differently according to the adaptive sampling process.
For example, Jacobsen states that the adaptive sampler performs “scheduling the new rays, shooting the rays…and trying to interpolate as many pixels as possibly,” and that the cycle is then performed a certain number of times until the complete image is generated (Section 4.2 and Fig. 4.6; schedule rays…adaptive sampling…each iteration of the schedule-shoot-interpolate passes, we try to interpolate by sampling every X'th pixel on the vertical and horizontal axis. That is, we divide the image into mxn grids of size X x X). Jacobsen therefore determines, relative to other image regions/pixels, which regions or sampling positions receive ray tracing and which are interpolated. Under broadest reasonable interpretation, this encompasses “determining, for different regions of the render output, a relative number of rays to be traced for the region of the render output…”.
Applicant argues: Jacobsen does not disclose “each thread tracing a different number of rays for one or more sampling positions of the subregion compared to other sampling positions of the subregion, based on the order that the thread cycles over the sampling positions” and further states that in Jacobsen, some sampling positions are ray traced and some are not, but this is not based on an “order” in which the sampling positions are “cycled over”, but rather, it is determined according to the predetermined algorithm, i.e. with pixels at corners/midpoints of the grids having a ray traced for them, with other pixels instead being interpolated.
Examiner replies: Applicant’s argument is not persuasive because the claim does not require a particular remainder-allocation embodiment, such as allocating three rays among four sampling positions by giving the first three positions one ray and the last position none, nor does the claim require a specific type of order. Under broadest reasonable interpretation, “Based on the order that the thread cycles over the sampling positions” reasonably encompasses an ordered traversal or ordered sampling schedule.
Jacobsen expressly discloses such an ordered sampling schedule. Section 4.2 and Fig. 4.6 expressly states that the adaptive sample includes “scheduling the new rays, shooting the rays…and trying to interpolate as many pixels as possibly,” and that “this cycle is then performed a certain number of times until the complete image is generated”. Jacobsen further discloses that “each iteration” of the schedule-shoot-interpolate passes samples every X’th pixel and divides the image into grids of size XxX (Pg. 34, Section 4.2). Figure 4.7 expressly shows ordered stages of ray tracing, stating that “orange pixels are ray traced in the first step, blue in the second step, and green in the third” (Pg. 35, Section 4.2). Jacobsen also states that once the corners of grid size Xi have been sampled, the grid size is halved and grid size Xi-1 is used, thereby dividing each box into four overlapping sub-boxes (Pg. 35, Section 4.2). Thus, Jacobsen’s adaptive sampler cycles through sampling positions in a predetermined order defined by schedule-shoot-interpolate passes and grid refinement. Positions selected earlier in the ordered schedule, such as grid corners or center positions, are queued for ray tracing, while other positions are interpolated or delayed until later stages. Therefore, different sampling positions receive different ray tracing stages based on their positions in the ordered adaptive sampling cycle. Jacobsen’s order being predetermined by an algorithm does not remove it from the scope of “order,” and when this ordered sampling schedule is applied by each worker thread to its assigned row or 16x16 block, as taught by The Dude, the thread cycles over the sampling positions of its subregion and traces rays for different positions based on that ordered schedule.
Regarding the remaining arguments: Applicant argues with respect to the amended claim language, which is fully addressed in the prior art rejections set forth below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6, 11-17, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jacobson et al., "Spatial Adaptive Sampling in Real Time Ray Tracing", Department of Computer Science, Lund University, (2021), pages 1-67, hereinafter referred to as “Jacobson”, in view of NVIDIA “NVIDIA PTX ISA: NVIDIA CUDA Programming Guide”, Version 8.4 published on March 2, 2024, by NVIDIA Corporation, and in further view of The Dude et al. (2017, March 27). Parallel ray tracing in 16x16 chunks. Stack Overflow. https://stackoverflow.com/questions/43056609/parallel-ray-tracing-in-16x16-chunks, hereinafter referred to as “The Dude”.
Regarding claim 1, Jacobson discloses a method of operating a graphics processor to generate a render output made up of a plurality of sampling positions by performing a ray tracing process in which rays are traced through a scene to be rendered (Pg. 21, Section 3.1; The ray tracer, a backwards traced path tracer that supports diffuse, specular, and transmissive materials…also uses importance sampling, defined as sampling of rays that affect the estimation. Section 4; implementation of the ray tracer…pipeline stages of the basic mode, then explains how a single ray is sampled (including how importance sampling is performed), and ends with a more thorough look at the de-noising filter implementation. Pg. 25, Section 3.2.3; For the case of real time ray tracers…video footage was captured to compare the rendered scene…The path was chosen on a per scene basis, with the idea that the path should show various parts of the scene, and expose the ray tracing code for different challenges)
wherein the ray tracing budget number (B) for the render output corresponding to a total number of rays that can be traced when generating the render output (Pg. 21, Section 3.3.1; default mode of the ray tracer…uses 1 spp. Pg. 30, Section 4.1.2; The budget for a real time tracer is limited to that of ca 1 spp, it will generate a high frequency noisy ray traced image…where each pixel ray bounces seven times and samples the light source…), and wherein different numbers of rays can be traced for different regions of the render output (Pg. 23, Section 3.1.3; adaptive sampling mode…Figure 3.3 shows that we have sampled points A to E in a 9x9 image. From those values, we see that A, B, C are close enough for us to estimate the values in-between points (red) without having to ray trace them. The next step would then be to trace intersecting points…white pixels are not calculated, red pixels are interpolated pixels, and the rest are ray traced. Pg. 34, Section 4.2; sampling every X'th pixel on the vertical and horizontal axis. That is, we divide the image into mxn grids of size XxX), the method comprising:
determining, for different regions of the render output, a relative number of rays to be traced for the region of the render output (Pg. 23, Fig. 3.3 and Pg. 34-36, Section 4.2 and Fig. 4.6; schedule rays…adaptive sampling…each iteration of the schedule-shoot-interpolate passes, we try to interpolate by sampling every X'th pixel on the vertical and horizontal axis. That is, we divide the image into mxn grids of size X x X……Section 4.2.1; To communicate between the ray shooter and the ray scheduler, we work with a texture that says which pixels should be ray traced (queued), and which should not be. That is, for the same size as the application window resolution, we have a texture which operates as a mask…For each time we run the schedule-shoot-interpolate passes, we launch the ray scheduler and the ray tracer program on each pixel),
(Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes…The figure starts out with a 9x9 grid, then 5x5 grid, and lastly a 3x3 grid),
(Pg. 23, Section 3.1.3; adaptive sampling mode…Pg. 34-36, Section 4.2; each iteration of the schedule-shoot-interpolate passes, we try to interpolate by sampling every X'th pixel on the vertical and horizontal axis. That is, we divide the image into mxn grids of size X x X…to communicate between the ray shooter and the ray scheduler, we work with a texture that says which pixels should be ray traced (queued), and which should not be…For each time we run the schedule-shoot-interpolate passes, we launch the ray scheduler and the ray tracer program on each pixel. Fig. 3.3 and Fig. 4.7) and the ray tracing budget number (B) for the render output (Pg. 21, Section 3.3.1; default mode of the ray tracer…uses 1 spp. Pg. 30, Section 4.1.2; The budget for a real time tracer is limited to that of ca 1 spp, it will generate a high frequency noisy ray traced image…where each pixel ray bounces seven times and samples the light source… Examiner's note: Jacobsen 3.1.1 and 4.1.2 describe a fixed real-time sampling budget, while Jacobsen 3.1.3 and 4.2 show an adaptive sampler that gives some regions more samples than others, effectively allocating the global budget among regions according to their relative importance)
and performing ray tracing for the region (Pg. 27-28, Section 4.1; The core system. Pg. 28-29 and Fig. 4.2; ray tracing system is a backwards traced path tracer with importance sampling, we cast rays from the cameras point of view and let it bounce several times…Pg. 30-31, Section 4.1.2; raw colour values generated per pixel can be seen in figure 4.5, where each pixel ray bounces seven times and samples the light source in the same way illustrated in figure 4.3).
Jacobson does appear to explicitly disclose allocating groups of threads to a region of the render output.
In the same art of GPU parallel computing, NVIDIA discloses allocating groups of threads to a region of the render output (Pg. 11, Section 2.2.1; CTA, is an array of threads that execute a kernel concurrently or in parallel…Each CTA thread uses its thread identifier to determine its assigned role, assign specific input and output positions, compute addresses, and select work to perform. The thread identifier is a three-element vector tid, (with elemenets tid.x, tid.y, and tid.z) that specifies the thread's position within a 1D, 2D, or 3D CTA…Examiner's note: CTAs have a 1D/2D/3D shape (ntid.x, ntid.y, ntid.z) and each thread's index selects a position within that array, which can map to a block of pixels or a region of the render output (CTAs are launched over 2D data like images)).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to implement the GPU thread-group organization taught by NVDIA into the ray-tracing pipeline of Jacobson. Jacobsen already recites that a large portion of the computation is performed on a massively parallel GPU with multiple cores and threads (Jacobson Pg. 13, Section 2.2), so a person of ordinary skill in the art would see the NVDIA PTX programming guide for conventional techniques on how to organize GPU threads in groups (CTAs) over 2D image regions to execute the ray-tracing techniques efficiently. Combining Jacobson’s parallel GPU workload into standard cooperative thread arrays yield predictable results in providing efficient thread allocation and improve real-time ray-tracing performance.
Jacobson in view of NVDIA does not disclose to perform the ray tracing for the region of the render output, each thread of the groups of threads being allocated to a subregion of the region; determining the number of rays to be traced by each thread of the groups of threads when performing ray tracing for the subregion to which they have been allocated; and including each of the threads tracing the determined number of rays for the subregion to which they have been allocated.
In the same art of parallel ray tracing, The Dude discloses to perform the ray tracing for the region of the render output, each thread of the groups of threads being allocated to a subregion of the region (Response from Adrian McCarthy; you create a Manager object that returns a chunk of work (in the form of a Task) each time a thread calls its GetTask method…
std::unique_ptr<Task> Manager::GetTask() {
std::lock_guard guard(mutex);
std::unique_ptr<Task> t;
if (next_row < HEIGHT) {
t = std::make_unique<Task>(next_row);
++next_row;
}
return t;
}
…the manager creates a new task to ray trace the next row. (You could use 16x16 bocks instead of rows if you like)…Examiner's note: shows a Manager that, each time a worker thread calls GetTask(), returns a new chunk of work (row or 16x16 block) derived from a shared "next_row" counter. This divides the total rows/pixels (and thus rays) across the available threads, so each thread is assigned to a specific number of pixels/rays for its subregion (each worker thread processes a subregion of the image)),
determining the number of rays to be traced by each thread of the groups of threads when performing ray tracing for the subregion to which they have been allocated (Response from Adrian McCarthy's; Manager::GetTask()…t = std::make_unique<Task>(next_row);…return t;…the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like.) When all the tasks have been issued, it just returns an empty pointer, which essentially tells the calling thread that there's nothing left to do, and the calling thread will then exit),
and including each of the threads tracing the determined number of rays for the subregion to which they have been allocated (Response from Adrian McCarthy; void WorkerThread(Manager *manager) {
while (auto task = manager->GetTask()) {
task->Execute();
}
}
Examiner's note: each worker thread runs a loop that repeatedly calls manager to get a Task (row/16x16 block) and executes ray tracing for that particular subregion it has been allocated).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to further incorporate the region-assignment technique of The Dude into the combined system of Jacobsen and NVDIA. The Dude addresses how different image regions take different amounts of time to ray trace, so by applying this known scheduling technique to Jacobsen’s ray-tracing threads, organized as groups as taught by NVDIA, would have been a routine use of familiar load-balancing. The motivation lies in the advantage of improving thread utilization, faster rendering, and improving overall CPU/GPU utilization.
Regarding claim 2, Jacobsen in view of NVDIA and in further view of The Dude discloses the method of claim 1, and further discloses wherein the relative number of rays to be traced for different regions of the render output is determined based on data indicating the presence of sampling positions in one or more different regions of the render output that could particularly benefit from receiving more ray tracing samples (Jacobson Pg. 23, Section 3.1.3; the adaptive sampling mode aims to lower the average sampling per pixel count by skipping ray tracing some pixels…From those values, we see that A, B, and C are close enough for us to estimate the values in-between these points (red) without having to ray trace them. Pg. 36-37, Section 4.2.2; When we try to interpolate pixels (that is, estimating new data points from already known data) …we demand that certain parameters need to be similar enough. If they are not similar enough, we schedule new rays…Examiner's note: uses per-pixel data (Section 4.1.2-4.1.4; ray tracing geometry buffer and history buffer) to decide which pixels can be interpolated and which must be ray traced again).
The motivation to combine would’ve been the same as that set forth above with respect to claim 1.
Regarding claim 3, Jacobsen in view of NVDIA and in further view of The Dude discloses the method of claim 2, and further discloses wherein the data indicating the presence of sampling positions in one or more different regions of the render output that could particularly benefit from receiving more ray tracing samples (Jacobson Pg. 23, Section 3.1.3; the adaptive sampling mode aims to lower the average sampling per pixel count by skipping ray tracing some pixels…From those values, we see that A, B, and C are close enough for us to estimate the values in-between these points (red) without having to ray trace them. Pg. 36-37, Section 4.2.2; When we try to interpolate pixels (that is, estimating new data points from already known data) …we demand that certain parameters need to be similar enough. If they are not similar enough, we schedule new rays…Examiner's note: uses per-pixel data (Section 4.1.2-4.1.4; ray tracing geometry buffer and history buffer) to decide which pixels can be interpolated and which must be ray traced again) comprises one or more of:
data indicating an area of disocclusion for the region (Jacobson Pg. 33, Section 4.1.4; We now need to check if we actually see the same thing from the previous point of view and the current point of view. For each pixel, we compare the history buffer and ray tracing buffers' world positions, world normals, and their object IDs…If all of these criteria are met, we have a successful re-projection…If the re-projection was not successful, we set the history length of that pixel to one. Examiner's note: pixels failing re-projection tests are disoccluded or otherwise newly visible (the implementation flags them by resetting history)),
data indicating an area of specular highlights for the region; data indicating an area of high spatiotemporal variance and/or soft shadows (Jacobson Pg. 21, Section 3.1; The ray tracer, a backwards traced path tracer that supports diffuse, specular, and transmissive materials…Pg. 33, Section 4.1.4; use the variance to process the colour representation of the pixels, we use the RGB luminance representation as the dataset to calculate the variance on…these two values are what we define as the moments of the pixels…due to the rough and noisy data that is the raw ray tracing input, we first estimate the moments spatially, and when the said history length is long enough, we calculate the variance temporally),
and data from a learned algorithm or neural network, optionally feedback data from a denoiser (Jacobson Pg. 32, Section 4.1.4; Spatio-Temporal Variance Guided Filtering denoising filter… (σz, σl, σn) and a projection limit θp. These can be thought of as the parameters for the filtering process).
The motivation to combine would’ve been the same as that set forth above with respect to claim 1.
Regarding claim 4, Jacobsen in view of NVDIA and in further view of The Dude discloses the method of claim 1, Jacobson further teaches is determined based on a determined number of rays that are to be traced for the region (Jacobsen Section 3.1.3 and Pg. 22, Section 3.1.3; Figure 3.3 shows that we have sampled points A to E in a 9x9 image. From those values, we see that A, B, and C are close enough for us to estimate the values in-between these points…).
Jacobson does not disclose wherein the number of groups of threads allocated to a region of the render output.
In the same art of parallel GPU computing, NVIDIA discloses wherein the number of groups of threads allocated to a region of the render output (NVIDIA Pg. 11, Section 2.2.1; CTA, is an array of threads that execute a kernel concurrently or in parallel…Each CTA thread uses its thread identifier to determine its assigned role, assign specific input and output positions, compute addresses, and select work to perform. The thread identifier is a three-element vector tid, (with elemenets tid.x, tid.y, and tid.z) that specifies the thread's position within a 1D, 2D, or 3D CTA…Examiner's note: CTAs have a 1D/2D/3D shape (ntid.x, ntid.y, ntid.z) and each thread's index selects a position within that array, which can map to a block of pixels or a region of the render output (CTAs are launched over 2D data like images))
The motivation to combine would’ve been the same as that set forth above with respect to claim 1.
Regarding claim 5, Jacobsen in view of NVDIA and in further view of The Dude discloses the method of claim 1, and further discloses wherein determining the number of rays to be traced by each thread of the groups of threads when performing ray tracing for the subregion to which they have been allocated comprises:
determining an approximate number of rays to be traced for the region of the render output by multiplying the relative number of rays to be traced for the region by the ray tracing budget number (B), divided by the sum of all the relative numbers of rays to be traced for all of the regions of the render output (Jacobson Pg. 21, Section 3.3.1; default mode of the ray tracer…uses 1 spp. Pg. 30, Section 4.1.2; The budget for a real time tracer is limited to that of ca 1 spp, it will generate a high frequency noisy ray traced image…where each pixel ray bounces seven times and samples the light source…Pg. 23, Section 3.1.3; the adaptive sampling mode aims to lower the average sampling per pixel count by skipping ray tracing some pixels. Examiner's note: states a fixed sampling budget, then uses the adaptive sampling rules to give some regions more sampling and others none, showing importance sampling and distributing the global budget across region in correspondence to that importance).
The motivation to combine would’ve been the same as that set forth above with respect to claim 1.
Regarding claim 6, Jacobsen in view of NVIDIA and in further view of The Dude discloses the method of claim 5, Jacobson does not disclose wherein the number of groups of threads allocated to a region of the render output is M, and wherein each of the M groups of threads comprises a number of threads N.
In the same art of parallel GPU computing, NVIDIA discloses wherein the number of groups of threads allocated to a region of the render output is M, and wherein each of the M groups of threads comprises a number of threads N (NVIDIA Pg. 11, Section 2.2.1; The vector ntid specifies the number of threads in each CTA dimension).
The motivation to combine would’ve been the same as that set forth above with respect to claim 1.
Jacobson in view of NVIDIA does not disclose and the method further comprises rounding the determined approximate number of rays to be traced for the region to a nearest multiple of M*N to thereby give a rounded value, and dividing this rounded value by M*N to give the number of rays to be traced by each of the M*N threads when performing ray tracing for the subregion to which they have been allocated.
In the same art of parallel ray tracing, The Dude discloses and the method further comprises rounding the determined approximate number of rays to be traced for the region to a nearest multiple of M*N to thereby give a rounded value (The Dude's Code snippet; auto size = WIDTH*HEIGHT;), and dividing this rounded value by M*N (The Dude's Code snippet; auto chunk = size / nThreads;) to give the number of rays to be traced by each of the M*N threads when performing ray tracing for the subregion to which they have been allocated (The Dude's post; dividing the image into as many chunks as the system has and rendering them parallel).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine Jacobson’s per-region ray tracing budgets with the GPU execution model as taught by NVIDIA and the per-thread division of work by rounding each region’s ray count to a multiple of the total thread MxN and then dividing that amount evenly among the thead as taught by The Dude. Matching work counts to warp/group sizes and balancing per-thread workload is a standard optimization that avoids idle threads, maximizes occupancy, and overall improves throughput.
Regarding claim 11, claim 11 is the system claim (Jacobson Section 2.2 Technical Details) of method claim 1 and is accordingly rejected using substantially similar rationale as to that which is set for with respect to claim 1.
Regarding claim 12, claim 12 has similar limitations as of claim 2, except it is a system claim (Jacobson Section 2.2 Technical Details), therefore it is rejected under the same rationale as claim 2.
Regarding claim 13, claim 13 has similar limitations as of claim 3, except it is a system claim (Jacobson Section 2.2 Technical Details), therefore it is rejected under the same rationale as claim 3.
Regarding claim 14, claim 14 has similar limitations as of claim 4, except it is a system claim (Jacobson Section 2.2 Technical Details), therefore it is rejected under the same rationale as claim 4.
Regarding claim 15, claim 15 has similar limitations as of claim 5, except it is a system claim (Jacobson Section 2.2 Technical Details), therefore it is rejected under the same rationale as claim 5.
Regarding claim 16, Jacobson in view of NVDIA and in further view of The Dude discloses the graphics processor of claim 11, Jacobson further discloses wherein each of the subregions of the region comprises a plurality of sampling positions (Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes…The figure starts out with a 9x9 grid, then 5x5 grid, and lastly a 3x3 grid),
Jacobson in view of NVDIA does not disclose and each thread traces the determined number of rays for the subregion by cycling over sampling positions of its allocated subregion in turn to trace one or more rays for one or more of the sampling positions of the subregion.
In the same art of parallel ray tracing, The Dude discloses and each thread traces the determined number of rays for the subregion by cycling over sampling positions of its allocated subregion in turn to trace one or more rays for one or more of the sampling positions of the subregion (The Dude Response from Adrian McCarthy; void WorkerThread(Manager *manager) {
while (auto task = manager->GetTask()) {
task->Execute();
}
}
Examiner's note: each worker thread runs a loop that repeatedly calls manager to get a Task (row/16x16 block) and executes ray tracing for that particular subregion it has been allocated).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to implement The Dude’s multi-threaded work-distribution technique into Jacobson’s and NVDIA’s combined system. Doing so allows Jacobson’s scene-dependent ray-tracing to be dynamically fed to worker threads, yielding predictable results in improving parallel CPU/GPU utilization, balance the load among threads, and reducing render time.
Regarding claim 17, Jacobson in view of NVDIA and in further view of The Dude discloses the graphics processor of claim 16, Jacobson further discloses is not a multiple of the number of sampling positions that each subregion comprises (Pg. 35, Section 4.2.1; Figure 4.7 shows which pixels are ray traced for the adaptive sampler with initial grid size of 9x9 at different stages. Orange pixels are ray traced in the first step, blue in the second step, and green in the third. Red pixels shows the pixel we are trying to calculate. Red box shows which area the red pixel looks at in the different stages (but does not use the white pixels). Examiner's note: within a 9x9 subregion, only some pixels are traced in each pass, and the rest are white/interpolated. When the adaptive sampling technique is executed by a thread for its subregion, the total number of rays that thread traces for that subregion equals the number of traced pixels, not all sampling positions. For example, tracing 5 pixels in a 9x9 subregion yields 5 rays, which is not a multiple of 81);
and each thread traces a different number of rays for one or more sampling positions of the subregion compared to other sampling positions of the subregion, based on the order that the thread cycles over the sampling positions (Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes… Figure 4.7 shows which pixels are ray traced for the adaptive sampler with initial grid size of 9x9 at different stages. Orange pixels are ray traced in the first step, blue in the second step, and green in the third. Red pixels shows the pixel we are trying to calculate. Red box shows which area the red pixel looks at in the different stages (but does not use the white pixels).
Jacobson in view of NVIDIA does not disclose wherein the number of rays to be traced by each of the threads for the subregion to which they have been allocated.
In the same art of parallel ray tracing, The Dude discloses wherein the number of rays to be traced by each of the threads for the subregion to which they have been allocated (The Dude Response from Adrian McCarthy; the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like.))
The motivation to combine would’ve been the same as that set forth above with respect to claim 1.
Regarding claim 19, Jacobson in view of NVDIA and in further view of The Dude discloses the graphics processor of claim 11, and Jacobson further discloses wherein the graphics processor is configured to (Jacobson Section 2.2 Technical Details), when successively generating one or more plural render outputs having corresponding regions (Pg. 25, Section 3.2.3; For the case of real time ray tracers…video footage was captured to compare the rendered scene…The path was chosen on a per scene basis, with the idea that the path should show various parts of the scene, and expose the ray tracing code for different challenges. Pg. 25, Section 3.2.4; The same paths for when the camera was moving was used in all measurements, that is, the path used for frame generation measurement is the same path used when capturing video footage of the rendered scene. Examiner's note: renders the same scenes along fixed camera paths and uses history buffers to align pixels frame-toframe, corresponding regions and subregions across successive frames), each corresponding region comprising a corresponding set of subregions (Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes…The figure starts out with a 9x9 grid, then 5x5 grid, and lastly a 3x3 grid), and when generating successive render outputs (Pg. 25, Section 3.2.3; For the case of real time ray tracers…video footage was captured to compare the rendered scene…The path was chosen on a per scene basis, with the idea that the path should show various parts of the scene, and expose the ray tracing code for different challenges. Pg. 25, Section 3.2.4; The same paths for when the camera was moving was used in all measurements, that is, the path used for frame generation measurement is the same path used when capturing video footage of the rendered scene. Examiner's note: renders the same scenes along fixed camera paths and uses history buffers to align pixels frame-toframe, corresponding regions and subregions across successive frames).
Jacobson in view of NVIDIA does not disclose each corresponding subregion being allocated to a same thread, and cycle the sampling position at which each thread starts cycling over the sampling positions of each corresponding subregion to which it is allocated.
In the same art of parallel ray tracing, The Dude discloses each corresponding subregion being allocated to a same thread (Response from Adrian McCarthy; t = std::make_unique<Task>(next_row); ++next_row; … the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like. Examiner’s note: worker thread based on per row or per block), and cycle the sampling position at which each thread starts cycling over the sampling positions of each corresponding subregion to which it is allocated (The Dude Response from Adrian McCarthy; void WorkerThread(Manager *manager) {
while (auto task = manager->GetTask()) {
task->Execute();
}
}
…the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like. Examiner's note: each worker thread runs a loop that repeatedly calls manager to get a Task (row/16x16 block) and executes ray tracing for that particular subregion it has been allocated),
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to implement The Dude’s multi-threaded work-distribution technique into Jacobson’s and NVIDIA’s combined system. Doing so allows uniform sampling across frames, improving the uniformity and consistency of the sequence of render outputs without increasing ray-tracing workload.
Regarding claim 20, claim 11 is the CRM claim (Jacobson Section 2.2 Technical Details) of method claim 1 and is accordingly rejected using substantially similar rationale as to that which is set for with respect to claim 1.
Claim(s) 7 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jacobson et al., "Spatial Adaptive Sampling in Real Time Ray Tracing", Department of Computer Science, Lund University, (2021), pages 1-67, hereinafter referred to as “Jacobson”, in view of The Dude et al. (2017, March 27). Parallel ray tracing in 16x16 chunks. Stack Overflow. https://stackoverflow.com/questions/43056609/parallel-ray-tracing-in-16x16-chunks, hereinafter referred to as “The Dude”.
Regarding claim 7, Jacobson discloses a method of operating a graphics processor to generate a render output made up of a plurality of sampling positions by performing a ray tracing process in which rays are traced through a scene to be rendered (Pg. 21, Section 3.1; The ray tracer, a backwards traced path tracer that supports diffuse, specular, and transmissive materials…also uses importance sampling, defined as sampling of rays that affect the estimation. Section 4; implementation of the ray tracer…pipeline stages of the basic mode, then explains how a single ray is sampled (including how importance sampling is performed), and ends with a more thorough look at the de-noising filter implementation. Pg. 25, Section 3.2.3; For the case of real time ray tracers…video footage was captured to compare the rendered scene…The path was chosen on a per scene basis, with the idea that the path should show various parts of the scene, and expose the ray tracing code for different challenges), the method comprising:
(Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes…The figure starts out with a 9x9 grid, then 5x5 grid, and lastly a 3x3 grid),
performing ray tracing for the region (Pg. 27-28, Section 4.1; The core system. Pg. 28-29 and Fig. 4.2; ray tracing system is a backwards traced path tracer with importance sampling, we cast rays from the cameras point of view and let it bounce several times…Pg. 30-31, Section 4.1.2; raw colour values generated per pixel can be seen in figure 4.5, where each pixel ray bounces seven times and samples the light source in the same way illustrated in figure 4.3), .
(Pg. 35, Section 4.2.1; Figure 4.7 shows which pixels are ray traced for the adaptive sampler with initial grid size of 9x9 at different stages. Orange pixels are ray traced in the first step, blue in the second step, and green in the third. Red pixels shows the pixel we are trying to calculate. Red box shows which area the red pixel looks at in the different stages (but does not use the white pixels). Examiner's note: within a 9x9 subregion, only some pixels are traced in each pass, and the rest are white/interpolated. When the adaptive sampling technique is executed by a thread for its subregion, the total number of rays that thread traces for that subregion equals the number of traced pixels, not all sampling positions. For example, tracing 5 pixels in a 9x9 subregion yields 5 rays, which is not a multiple of 81); the method comprises: each thread tracing a different number of rays for one or more sampling positions of the subregion compared to other sampling positions of the subregion, based on the order that the thread cycles over the sampling positions (Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes… Figure 4.7 shows which pixels are ray traced for the adaptive sampler with initial grid size of 9x9 at different stages. Orange pixels are ray traced in the first step, blue in the second step, and green in the third. Red pixels shows the pixel we are trying to calculate. Red box shows which area the red pixel looks at in the different stages (but does not use the white pixels).
Jacobson fails to explicitly disclose allocating a plurality of threads to a region of the render output, each thread being allocated to a subregion of the region to perform ray tracing for the subregion, and including each of the threads tracing rays for its allocated subregion by cycling over sampling positions of its allocated subregion in turn to trace one or more rays for one or more of the sampling positions of the subregion, and wherein the number of rays to be traced by each of the threads for the subregion to which they have been allocated.
In the same art of parallel ray tracing, The Dude discloses allocating a plurality of threads to a region of the render output, each thread being allocated to a subregion of the region to perform ray tracing for the subregion,
(Response from Adrian McCarthy; you create a Manager object that returns a chunk of work (in the form of a Task) each time a thread calls its GetTask method…
std::unique_ptr<Task> Manager::GetTask() {
std::lock_guard guard(mutex);
std::unique_ptr<Task> t;
if (next_row < HEIGHT) {
t = std::make_unique<Task>(next_row);
++next_row;
}
return t;
}
the manager creates a new task to ray trace the next row. (You could use 16x16 bocks instead of rows if you like)…Examiner's note: shows a Manager that, each time a worker thread calls GetTask(), returns a new chunk of work (row or 16x16 block) derived from a shared "next_row" counter. This divides the total rows/pixels (and thus rays) across the available threads, so each thread is assigned to a specific number of pixels/rays for its subregion (each worker thread processes a subregion of the image)),
including each of the threads tracing rays for its allocated subregion by cycling over sampling positions of its allocated subregion in turn to trace one or more rays for one or more of the sampling positions of the subregion (Response from Adrian McCarthy; void WorkerThread(Manager *manager) {
while (auto task = manager->GetTask()) {
task->Execute();
}
}
Examiner's note: each worker thread runs a loop that repeatedly calls manager to get a Task (row/16x16 block) and executes ray tracing for that particular subregion it has been allocated).
wherein the number of rays to be traced by each of the threads for the subregion to which they have been allocated (The Dude Response from Adrian McCarthy; the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like.))
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to implement The Dude’s multi-threaded work-distribution technique into Jacobson’s ray tracing system. Doing so allows Jacobson’s scene-dependent ray-tracing to be dynamically fed to worker threads, yielding predictable results in improving parallel CPU/GPU utilization, balance the load among threads, and reducing render time.
Examiner’s note: as stated above in examiner’s response to arguments, the claim does not require a particular remainder-allocation embodiment, such as allocating three rays among four sampling positions by giving the first three positions one ray and the last position none, nor does the claim require a specific type of order. Under broadest reasonable interpretation, “Based on the order that the thread cycles over the sampling positions” reasonably encompasses an ordered traversal or ordered sampling schedule. Jacobsen expressly discloses such an ordered sampling schedule (Fig. 4.6 and Pg. 34, Section 4.2; scheduling new rays, shooting the rays using the old program, and trying to interpolate as many pixels as possible. This cycle is then performed a certain number of times until the complete image is generated…each iteration of the schedule-shoot-interpolate passes…sampling every X’th pixel on the vertical and horizontal axis…Once we have sampled the corners of grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes…orange pixels are rasy traced in the first step, blue in the second, and green in the third. Fig. 4.7 and Pg. 36, Section 4.2.1; says which pixels should be traced (queued), and which should not be…).
Thus, Jacobsen’s adaptive sampler cycles through sampling positions in a predetermined order defined by schedule-shoot-interpolate passes and grid refinement. Positions selected earlier in the ordered schedule, such as grid corners or center positions, are queued for ray tracing, while other positions are interpolated or delayed until later stages. Therefore, different sampling positions receive different ray tracing stages based on their positions in the ordered adaptive sampling cycle. Jacobsen’s order being predetermined by an algorithm does not remove it from the scope of “order,” and when this ordered sampling schedule is applied by each worker thread to its assigned row or 16x16 block, as taught by The Dude, the thread cycles over the sampling positions of its subregion and traces rays for different positions based on that ordered schedule.
Regarding claim 10, Jacobson in view of The Dude discloses the method of claim 7, and Jacobson further discloses wherein the graphics processor is configured to (Jacobson Section 2.2 Technical Details), repeating the method steps of claim 7 to successively generate one or more further render outputs having corresponding regions (Pg. 25, Section 3.2.3; For the case of real time ray tracers…video footage was captured to compare the rendered scene…The path was chosen on a per scene basis, with the idea that the path should show various parts of the scene, and expose the ray tracing code for different challenges. Pg. 25, Section 3.2.4; The same paths for when the camera was moving was used in all measurements, that is, the path used for frame generation measurement is the same path used when capturing video footage of the rendered scene. Examiner's note: renders the same scenes along fixed camera paths and uses history buffers to align pixels frame-to-frame, corresponding regions and subregions across successive frames), each corresponding region comprising a corresponding set of subregions (Pg. 35, Section 4.2 and Fig. 4.7; Once we have sampled the corners of the grid size Xi, we half the size and use grid size Xi-1. In other words, we divide each box of the grid into four new overlapping sub-boxes…The figure starts out with a 9x9 grid, then 5x5 grid, and lastly a 3x3 grid), and when generating successive render outputs (Pg. 25, Section 3.2.3; For the case of real time ray tracers…video footage was captured to compare the rendered scene…The path was chosen on a per scene basis, with the idea that the path should show various parts of the scene, and expose the ray tracing code for different challenges. Pg. 25, Section 3.2.4; The same paths for when the camera was moving was used in all measurements, that is, the path used for frame generation measurement is the same path used when capturing video footage of the rendered scene. Examiner's note: renders the same scenes along fixed camera paths and uses history buffers to align pixels frame-toframe, corresponding regions and subregions across successive frames).
Jacobson does not disclose each corresponding subregion being allocated to a same thread, and each thread cycling the sampling position at which each thread starts cycling over the sampling positions of each corresponding subregion to which it is allocated.
In the same art of parallel ray tracing, The Dude discloses each corresponding subregion being allocated to a same thread (Response from Adrian McCarthy; t = std::make_unique<Task>(next_row); ++next_row; … the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like. Examiner’s note: worker thread based on per row or per block), and each thread cycling the sampling position at which each thread starts cycling over the sampling positions of each corresponding subregion to which it is allocated (The Dude Response from Adrian McCarthy; void WorkerThread(Manager *manager) {
while (auto task = manager->GetTask()) {
task->Execute();
}
}
…the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like. Examiner's note: each worker thread runs a loop that repeatedly calls manager to get a Task (row/16x16 block) and executes ray tracing for that particular subregion it has been allocated),
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to implement The Dude’s multi-threaded work-distribution technique into Jacobson’s ray tracing system. Doing so allows uniform sampling across frames, improving the uniformity and consistency of the sequence of render outputs without increasing ray-tracing workload.
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jacobson et al., "Spatial Adaptive Sampling in Real Time Ray Tracing", Department of Computer Science, Lund University, (2021), pages 1-67, hereinafter referred to as “Jacobson”, in view of The Dude et al. (2017, March 27). Parallel ray tracing in 16x16 chunks. Stack Overflow. https://stackoverflow.com/questions/43056609/parallel-ray-tracing-in-16x16-chunks, hereinafter referred to as “The Dude”, and in further view of Shirley et al. Ray tracing in one weekend, December 2020. https://raytracing. github.io/books/RayTracingInOneWeekend.html [Online; Sourced 2021-07- 12], hereinafter referred to as “Shirley”.
Regarding claim 9, Jacobson in view of The Dude disclose the method of claim 7, and further discloses comprising each thread starting the cycling over the sampling positions of its allocated subregion (The Dude Response from Adrian McCarthy; void WorkerThread(Manager *manager) {
while (auto task = manager->GetTask()) {
task->Execute();
}
}
…the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like. Examiner's note: each worker thread runs a loop that repeatedly calls manager to get a Task (row/16x16 block) and executes ray tracing for that particular subregion it has been allocated).
Jacobson and The Dude are combined for the reason set forth above with respect to claim 7.
Jacobson in view of The Dude does not disclose at a random sampling position of the subregion relative to the sampling position at which each other thread starts its cycle over the sampling positions of its allocated subregion.
In the same art of ray tracing, Shirley discloses at a random sampling position of the subregion relative to the sampling position at which each other thread starts its cycle over the sampling positions of its allocated subregion (Section 8.2; pixel_sample_square() that generates a random sample point within the unit square centered at the origin.…
vec3 pixel_sample_square() const {
// Returns a random point in the square surrounding a pixel at the origin.
auto px = -0.5 + random_double();
auto py = -0.5 + random_double();
return (px * pixel_delta_u) + (py * pixel_delta_v);
}).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the multi-threaded subregion processing of Jacobson in view of The Dude with Shirley’s random sampling positions. Randomized sampling or Monte Carlo methods is a known technique in adaptive sampling, and yields predictable results in effective anti-aliasing, reduced bias, and provide a more realistic output.
Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jacobson et al., "Spatial Adaptive Sampling in Real Time Ray Tracing", Department of Computer Science, Lund University, (2021), pages 1-67, hereinafter referred to as “Jacobson”, in view of NVIDIA “NVIDIA PTX ISA: NVIDIA CUDA Programming Guide”, Version 8.4 published on March 2, 2024, by NVIDIA Corporation, in further view of The Dude et al. (2017, March 27). Parallel ray tracing in 16x16 chunks. Stack Overflow. https://stackoverflow.com/questions/43056609/parallel-ray-tracing-in-16x16-chunks, hereinafter referred to as “The Dude”, and in further view of Shirley et al. Ray tracing in one weekend, December 2020. https://raytracing. github.io/books/RayTracingInOneWeekend.html [Online; Sourced 2021-07- 12], hereinafter referred to as “Shirley”.
Regarding claim 18, Jacobson in view of NVDIA and in further view of The Dude discloses the graphics processor of claim 17, and further discloses wherein each thread starts the cycling over the sampling positions of its allocated subregion (The Dude Response from Adrian McCarthy; void WorkerThread(Manager *manager) {
while (auto task = manager->GetTask()) {
task->Execute();
}
}
…the manager creates a new task to ray trace the next row. (You could use 16x16 blocks instead of rows if you like. Examiner's note: each worker thread runs a loop that repeatedly calls manager to get a Task (row/16x16 block) and executes ray tracing for that particular subregion it has been allocated).
Jacobson, NVDIA, and The Dude are combined for the reason set forth above with respect to claim 1.
Jacobson in view of NVIDIA and in further view of The Dude does not disclose at a random sampling position of the subregion relative to the sampling position at which each other thread starts its cycle over the sampling positions of its allocated subregion.
In the same art of ray tracing, Shirley discloses at a random sampling position of the subregion relative to the sampling position at which each other thread starts its cycle over the sampling positions of its allocated subregion (Section 8.2; pixel_sample_square() that generates a random sample point within the unit square centered at the origin.…
vec3 pixel_sample_square() const {
// Returns a random point in the square surrounding a pixel at the origin.
auto px = -0.5 + random_double();
auto py = -0.5 + random_double();
return (px * pixel_delta_u) + (py * pixel_delta_v);
}).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the multi-threaded subregion processing of Jacobson in view of NVIDIA in further view of The Dude with Shirley’s random sampling positions. Randomized sampling or Monte Carlo methods is a known technique in adaptive sampling, and yields predictable results in effective anti-aliasing, reduced bias, and provide a more realistic output.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JENNY NGAN TRAN whose telephone number is (571)272-6888. The examiner can normally be reached Mon-Thurs 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at (571) 272-2330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JENNY N TRAN/Examiner, Art Unit 2615
/ALICIA M HARRINGTON/Supervisory Patent Examiner, Art Unit 2615