DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to the Applicants' communication filed on January 03, 2025. In virtue of this communication, claims 1-29 are currently presented in the instant application.
Drawings
The drawings were submitted on January 03, 2025. These drawings are reviewed and accepted by the examiner.
Information Disclosure Statement
The information Disclosure Statement (IDS) Form PTO-1449, filed on February 12, 2025, follow the provisions of 37 CFR 1.97. Accordingly, the information disclosed therein was considered by the examiner.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1 rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of U.S. Patent No. 11816500 and 12223353. Although the claim at issue are not identical, they are not patentably distinct from each other because claim 1 of the Presen Applications is generic all that is recited in claim 1 of the Patent; in other word, the claim 1 of the Presen Application is anticipated by claim 1 of the Patent Application because it contains all the limitation of claim 1 of the Presen Applications, which is therefore an obvious variant thereof.
An illustration of the claim correspondence is as follows:
Present Applicant
Patent No. 11816500
Patent No. 12223353
Claim 1,
An apparatus comprising:
Claim 1,
A graphics multiprocessor, comprising:
Claim 1,
A graphics multiprocessor, comprising:
graphics processing circuitry coupled to a memory, the graphics processing circuitry having
a queue having an initial state of groups with having a first group having threads of a first instruction type and a second instruction type first and second instruction types and a second group having the threads of the first and second instruction types; and
a queue having an initial state of groups with a first group having threads of first and second instruction types and a second group having threads of the first and second instruction types; and
a queue having an initial
state of groups with
first group having
threads of
first and second
instruction types and a
second group having
threads of the first and
second instruction
types;
a regroup engine to regroup the threads into a third group having first threads of the first instruction type and a fourth group having second threads of the second instruction type.
a regroup circuitry to regroup threads from the initial state of groups of first and second groups into a regrouped state of groups based on instruction type for the queue including a third group having threads of the first instruction type and a fourth group having threads of the second instruction type, wherein the regroup circuitry to cause the third group to replace the first group in the regrouped state of groups for the queue and the fourth group to replace the second group in the regrouped state of groups based on instruction type for the queue.
a regroup circuitry to regroup threads from the initial state of groups of first and second groups into a regrouped state of groups including a third group having threads of the first instruction type and a fourth group having threads of the second instruction type based on an instruction type and to determine an order of inserting the third group and the fourth group into the queue to minimize divergence between threads.
The following table illustrates the conflicting claim pairs:
Present Application
1
3
4
5
7
Patent No. 11816500
1
3
4
5
7
Patent No. 12223353
24
26,
27
29
30
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-4 are rejected under 35 U.S.C. 102(1)(a) as being anticipated by
Abdel-Majeed, Mohammed et al.: "Warped gates: Gating aware scheduling and power gating for GPGPUs", 2013 46th Annual IEEE/ACM International Symposium on Microarchitecture, 12/7/13, (hereinafter “Mohammad”).
Regarding claim 1. An apparatus (Mohammad, see at least page 1, column 2, first par. and lines 1-3, “Graphics processing units (GPUs) are massively parallel processors that are designed to run workloads with thousands of concurrent threads”) comprising:
graphics processing circuitry coupled to a memory, the graphics processing circuitry (Mohammad, see section 2.1 and section 2.2) having
a queue having an initial state of groups with a first group having threads of first and second instruction types and a second group having threads of the first and second instruction types (Mohammad, see page 4, column 2 and the first 5 lines of the last paragraph, Figure 4 illustrates the shortcomings of current warp scheduling and its implication on power gating techniques. In this simplified illustration, the active warps set contains 10 warps with a mix of integer and floating point instructions”); and
PNG
media_image1.png
538
743
media_image1.png
Greyscale
a regroup engine to regroup threads into a third group having threads of the first instruction type and a fourth group having threads of the second instruction type (Mohammad, page 2, column 1, third paragraph, lines 1-6“Gating Aware Two-level warp scheduler: To address the inefficiencies of GPGPU scheduler in extracting idle periods, we present a gating-aware Two-level warp scheduler (GATES). GATES prioritizes issuing clusters of instructions that require the same type of execution unit for longer intervals before switching to a new instruction type”).
Regarding claim 2. Mohammad further discloses wherein the regroup engine to cause the third group to replace the first group in the queue and the fourth group to replace the second group in the queue having a regrouped state (Mohammad, see page 5, column 1, paragraph 2 and lines 7-17, Two-level scheduler (GATES) which takes into account previously issued instruction types in determining which ready warp to issue next. GATES prioritizes issuing the same instruction type as was issued in prior issue cycle to coalesce the utilization and idle periods of integer and floating point units. GATES will keep issuing instructions from the same type as long as there are ready warps in the active warps set. GATES switches to a warp with different instruction type when there are no more ready warps in the active warps set with the same instruction type as the one issued in the previous issue cycle).
Regarding claim 3. Mohammad further discloses the apparatus of claim 1 (as rejected above), Mohammad further discloses wherein the first instruction type or the second instruction type comprise one or more of a load/store instruction, an integer instruction, a floating point instruction, an integer mac instruction, an integer add instruction, a floating point add instruction, a floating point fma instruction, a floating point sine instruction, or a floating point cosine instruction (Mohammad, see page 3, column 1, first paragraph, lines 3-7, “The decoded instruction field includes the instruction type that determines which execution unit type that instruction requires for execution, namely an integer unit (INT), floating point unit (FP), special purpose functional unit (SFU), or load/store unit (LD/ST)”).
Regarding claim 4. Mohammad discloses the apparatus of claim 1 (as rejected above), Mohammad further discloses wherein the graphics processing circuitry further comprises: a thread scheduler coupled to the queue; and a plurality of execution units coupled to the thread scheduler (Mohammad, see page 3, column 1 and the last 4 lines of the second paragraph, “In GTX480, two schedulers are integrated in each SM and each scheduler can issue one ready warp per cycle as long as there are no structural hazards”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Mohammad Abdel-Majeed et al. (“Warped Gates: Gating Aware Scheduling and Power Gating for GPGPUs”, hereinafter “Mohammad”), as applied claim 4 above, and further in view of Shen et al. (US 20090172362 A1).
Regarding claim 5. Mohammad discloses claim 4 (as rejected above), but Mohammad does not disclose wherein the thread scheduler is configured to schedule the first instruction type of the third group for execution on a first processing resource with full utilization of the first processing resource, and wherein the thread scheduler is further configured to schedule the second instruction type of the fourth group for execution on a second processing resource with full utilization of the second processing resource.. However,
Shen discloses:
wherein the thread scheduler is configured to schedule the first instruction type of the third group for execution on a first processing resource with full utilization of the first processing resource, and wherein the thread scheduler is further configured to schedule the second instruction type of the fourth group for execution on a second processing resource with full utilization of the second processing resource. (Shen, see at least par. [0022-0023], In one embodiment, the integer execution unit 212 includes at least one data arithmetic logic unit (ALU) 220 configured to perform arithmetic operations based on the integer instruction operation being executed, at least one address generation unit (AGU) 222 configured to generate addresses for accessing data from cache/memory for the integer instruction operation being executed, a scheduler (not shown), a load/store unit (LSU) 224 to control the loading of data from memory/store data to memory, and a thread retirement module 226 configured to maintain intermediate results and to commit the results of the integer instruction operation to architectural state. In one embodiment, the ALU 220 and the AGU 222 are implemented as the same unit. The integer execution unit 212 further can include an input to receive data from the FPU 216 upon which depends one or more integer instruction operations being processed by the integer execution unit 212. The integer execution unit 214 can be similarly configured. [0023] In operation, the integer execution units 212 and 214 and the FPU 216 operate in parallel while sharing the resources of the front-end unit 202. Instructions associated with one or more threads are fetched by the instruction fetch module 206 and decoded by the instruction decode module 208. The instruction dispatch module 210 then can dispatch instruction operations represented by the decoded instructions to a select one of the integer execution unit 212, the integer execution unit 214, or the FPU 216 based on a variety of factors, such as operation type (e.g., integer or floating point), associated thread, loading, resource availability, architecture limitations, and the like. The instruction operations, thus dispatched, can be executed by their respective execution units during the same execution cycle. For floating point operations represented by buffered decoded instructions, the instruction dispatch module 210 determines the dispatch order to the FPU 216 based on thread priority, forward progress requirements, and the like. For integer instruction operations represented by buffered decoded instructions, the instruction dispatch module 210 determines both the dispatch order and which integer execution unit is to execute which integer instruction operation based on any of a variety of dispatch criteria, such as thread association, priority, loading, etc.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to combine the instruction types of Mohammad to include wherein the thread scheduler is configured to schedule the first instruction type of the third group for execution on a first processing resource with full utilization of the first processing resource, and wherein the thread scheduler is further configured to schedule the second instruction type of the fourth group for execution on a second processing resource with full utilization of the second processing resource, as taught by Shen, thereby to decrease overall instruction execution bandwidth. As an alternative, some multithreaded processing devices implement finer multithreading whereby instructions from multiple threads can be multiplexed at the beginning of the processing pipeline. However, the order in which the threads are selected for processing at the beginning of the processing pipeline typically is maintained for all subsequent stages of the pipeline. This can lead to processing inefficiencies in the event that a particular stage of the processing pipeline is idled by an instruction operation while waiting for some external event (e.g., the return of data from memory). Accordingly, a more flexible thread selection technique in a processing pipeline would be advantageous (Shen, see par. [0004]).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Abdel-Majeed, Mohammed et al.: "Warped gates: Gating aware scheduling and power gating for GPGPUs", 2013 46th Annual IEEE/ACM International Symposium on Microarchitecture, 12/7/13, (hereinafter “Mohammad”) as applied claim 1 and further in view of Johnson et al. (US 20200081748 A1).
Regarding claim 7. Mohammad further discloses wherein the regroup engine utilizes regrouping policies and an order that a new regrouped group is inserted in the queue is optimized depending on latencies, wherein the regroup engine utilizes regrouping policies to minimize divergence between the threads (Mohammad, see page 4, paragraph 2 and lines 2-8, “The order of instructions in the set is shown in the figure at the top. We assume each instruction is a simple add instruction,
each instruction has a latency of four cycles, and initiation interval is one cycle. These are the default parameters in GPGPUSim’s configuration file for Fermi [6]. The Two-level”),
Mohammad does not disclose wherein the regroup engine utilizes regrouping policies and an order that a new regrouped group is inserted in the queue is optimized depending on latencies, wherein the regroup engine utilizes regrouping policies to minimize divergence between the threads. However,
Johnson discloses:
wherein the regroup engine utilizes regrouping policies and an order that a new regrouped group is inserted in the queue is optimized depending on latencies, wherein the regroup engine utilizes regrouping policies to minimize divergence between the threads (Johnson, see pars. [0073, 0108, 0112 and 0114]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to combine the instruction types of Mohammad to include wherein the regroup engine utilizes regrouping policies and an order that a new regrouped group is inserted in the queue is optimized depending on latencies, wherein the regroup engine utilizes regrouping policies to minimize divergence between the threads”, as taught by Johnson. The modification provides an improved system and method for synchronization of multi-thread lanes, thereby to improve execution efficiency by promoting thread convergence. Threads executing common code sections (e.g., inner loop bodies) are urged to converge using instructions inserted at strategic locations in computer code sections. In support of this various instructions and/or code markers (e.g., compiler directives, ISA extensions) are introduced: Predict, Join, Wait, Confirm, Rejoin, and Cancel. Intelligently (e.g., based on profiles and/or policies) inserting these instructions into thread bodies enables the threads in a warp or other group to cooperate with the thread scheduler (e.g., of a streaming multiprocessor in a graphics processing unit) to promote thread convergence. (Johnson, see par. [0005]).
Claims 21-23 and 26-28 are rejected under 35 U.S.C. 103 as being unpatentable over Abdel-Majeed, Mohammed et al.: "Warped gates: Gating aware scheduling and power gating for GPGPUs", 2013 46th Annual IEEE/ACM International Symposium on Microarchitecture, 12/7/13, (hereinafter “Mohammad”), and further in view of Smith (US 20120089812 A1).
Regarding claim 21. (New) Mohammad discloses a method comprising:
maintaining, by graphics processing circuitry of a computing device, a queue having an initial state of groups having a first group having threads of a first instruction type and a second instruction type and a second group having the threads of the first and second instruction types (Mohammad, see page 4, column 2 and the first 5 lines of the last paragraph, Figure 4 illustrates the shortcomings of current warp scheduling and its implication on power gating techniques. In this simplified illustration, the active warps set contains 10 warps with a mix of integer and floating point instructions”); and
PNG
media_image1.png
538
743
media_image1.png
Greyscale
regrouping the threads into a third group having first threads of the first instruction type and a fourth group having second threads of the second instruction type (Mohammad, page 2, column 1, third paragraph, lines 1-6“Gating Aware Two-level warp scheduler: To address the inefficiencies of GPGPU scheduler in extracting idle periods, we present a gating-aware Two-level warp scheduler (GATES). GATES prioritizes issuing clusters of instructions that require the same type of execution unit for longer intervals before switching to a new instruction type”).
Mohammad does not disclose maintaining, by graphics processing circuitry of a computing device, a queue having an initial state of groups having a first group having threads of a first instruction type and a second instruction type and a second group having the threads of the first and second instruction types. However, Smith discloses:
maintaining, by graphics processing circuitry of a computing device, a queue having an initial state of groups having a first group having threads of a first instruction type and a second instruction type and a second group having the threads of the first and second instruction types (Smith, see at least par. [0079]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to combine the instruction types of Mohammad, to include maintaining, by graphics processing circuitry of a computing device, a queue having an initial state of groups having a first group having threads of a first instruction type and a second instruction type and a second group having the threads of the first and second instruction types, as taught by Smith. The modification provides an improved system and method for synchronization of multi-thread lanes, therefore provides a high degree of hardware parallelism and allow both dependent and independent threads to be executed concurrently. However, dependent threads (where the execution of one or more threads relies on the results of another thread) need to be synchronised to maintain error free processing (Smith, see par. [0013]).
Regarding claim 22. Mohammad in view of Smith discloses a method of claim 21 (as rejected above), and Mohammad in view of Smith further discloses further comprising causing the third group to replace the first group in the queue and the fourth group to replace the second group in the queue having a regrouped state (Mohammad, see page 5, column 1, paragraph 2 and lines 7-17, Two-level scheduler (GATES) which takes into account previously issued instruction types in determining which ready warp to issue next. GATES prioritizes issuing the same instruction type as was issued in prior issue cycle to coalesce the utilization and idle periods of integer and floating point units. GATES will keep issuing instructions from the same type as long as there are ready warps in the active warps set. GATES switches to a warp with different instruction type when there are no more ready warps in the active warps set with the same instruction type as the one issued in the previous issue cycle).
Regarding claim 23. Mohammad in view of Smith discloses a method of claim 21 (as rejected above), Mohammad in view of Smith further discloses wherein the first instruction type or the second instruction type comprises one or more of a load/store instruction, an integer instruction, a floating point instruction, an integer mac instruction, an integer add instruction, a floating point add instruction, a floating point fma instruction, a floating point sine instruction, or a floating point cosine instruction (Mohammad, see page 3, column 1, first paragraph, lines 3-7, “The decoded instruction field includes the instruction type that determines which execution unit type that instruction requires for execution, namely an integer unit (INT), floating point unit (FP), special purpose functional unit (SFU), or load/store unit (LD/ST)”).
Regarding claim 26. At least one computer-readable medium having stored thereon instructions which, when executed, cause a computing device to perform the method of claim 21. Therefore, claim 26 is further rejected based on the same rationale as claim 21 set forth above and incorporated herein.
Regarding claim 27. The computer-readable medium of claim 27 performs the method of claim 22. Therefore, claim 27 is further rejected based on the same rationale as claim 22 set forth above and incorporated herein.
Regarding claim 28. (New) The computer-readable medium of claim 28 performs same step of claim 23. Therefore, claim 28 is further rejected based on the same rationale as claim 23 set forth above and incorporated herein.
Claim 24 and 29 are rejected under 35 U.S.C. 103 as being unpatentable over Mohammad Abdel-Majeed et al. (“Warped Gates: Gating Aware Scheduling and Power Gating for GPGPUs”, hereinafter “Mohammad”) in view of Smith (US 20120089812 A1), as applied claim 21 above, and further in view of Shen et al. (US 20090172362 A1).
Regarding claim 24. Mohammad in view of Smith discloses claim 21 (as rejected above), but Mohammad in view of Smith does not disclose further comprising scheduling the first instruction type of the third group for execution on a first processing resource with full utilization of the first processing resource, and scheduling the second instruction type of the fourth group for execution on a second processing resource with full utilization of the second processing resource. However,
Shen discloses:
further comprising scheduling the first instruction type of the third group for execution on a first processing resource with full utilization of the first processing resource, and scheduling the second instruction type of the fourth group for execution on a second processing resource with full utilization of the second processing resource (Shen, see at least par. [0022-0023], [0022] In one embodiment, the integer execution unit 212 includes at least one data arithmetic logic unit (ALU) 220 configured to perform arithmetic operations based on the integer instruction operation being executed, at least one address generation unit (AGU) 222 configured to generate addresses for accessing data from cache/memory for the integer instruction operation being executed, a scheduler (not shown), a load/store unit (LSU) 224 to control the loading of data from memory/store data to memory, and a thread retirement module 226 configured to maintain intermediate results and to commit the results of the integer instruction operation to architectural state. In one embodiment, the ALU 220 and the AGU 222 are implemented as the same unit. The integer execution unit 212 further can include an input to receive data from the FPU 216 upon which depends one or more integer instruction operations being processed by the integer execution unit 212. The integer execution unit 214 can be similarly configured. [0023] In operation, the integer execution units 212 and 214 and the FPU 216 operate in parallel while sharing the resources of the front-end unit 202. Instructions associated with one or more threads are fetched by the instruction fetch module 206 and decoded by the instruction decode module 208. The instruction dispatch module 210 then can dispatch instruction operations represented by the decoded instructions to a select one of the integer execution unit 212, the integer execution unit 214, or the FPU 216 based on a variety of factors, such as operation type (e.g., integer or floating point), associated thread, loading, resource availability, architecture limitations, and the like. The instruction operations, thus dispatched, can be executed by their respective execution units during the same execution cycle. For floating point operations represented by buffered decoded instructions, the instruction dispatch module 210 determines the dispatch order to the FPU 216 based on thread priority, forward progress requirements, and the like. For integer instruction operations represented by buffered decoded instructions, the instruction dispatch module 210 determines both the dispatch order and which integer execution unit is to execute which integer instruction operation based on any of a variety of dispatch criteria, such as thread association, priority, loading, etc.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to combine the instruction types of Mohammad to include further comprising scheduling the first instruction type of the third group for execution on a first processing resource with full utilization of the first processing resource, and scheduling the second instruction type of the fourth group for execution on a second processing resource with full utilization of the second processing resource, as taught by Shen, thereby to decrease overall instruction execution bandwidth. As an alternative, some multithreaded processing devices implement finer multithreading whereby instructions from multiple threads can be multiplexed at the beginning of the processing pipeline. However, the order in which the threads are selected for processing at the beginning of the processing pipeline typically is maintained for all subsequent stages of the pipeline. This can lead to processing inefficiencies in the event that a particular stage of the processing pipeline is idled by an instruction operation while waiting for some external event (e.g., the return of data from memory). Accordingly, a more flexible thread selection technique in a processing pipeline would be advantageous (Shen, see par. [0004]).
Regarding claim 29. The computer-readable medium of claim 29 performs same step of claim 24. Therefore, claim 24 is further rejected based on the same rationale as claim 24 set forth above and incorporated herein.
Claims 25 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Abdel-Majeed, Mohammed et al.: "Warped gates: Gating aware scheduling and power gating for GPGPUs", 2013 46th Annual IEEE/ACM International Symposium on Microarchitecture, 12/7/13, (hereinafter “Mohammad”), and further in view of Smith (US 20120089812 A1) as applied claim 21 above, and further in view Johnson et al. (US 20200081748 A1).
Regarding claim 25. Mohammed in view of Smith discloses the method of claim 21 (as rejected above), but Mohammed in view of Smith does not disclose wherein the regroup engine utilizes regrouping policies and an order that a new regrouped group is inserted in the queue is optimized depending on latencies, wherein the regroup engine utilizes regrouping policies to minimize divergence between the threads.
Johnson discloses:
wherein the regroup engine utilizes regrouping policies and an order that a new regrouped group is inserted in the queue is optimized depending on latencies, wherein the regroup engine utilizes regrouping policies to minimize divergence between the threads (Johnson, see pars. [0073, 0108, 0112 and 0114]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to combine the instruction types of Mohammad to include wherein the regroup engine utilizes regrouping policies and an order that a new regrouped group is inserted in the queue is optimized depending on latencies, wherein the regroup engine utilizes regrouping policies to minimize divergence between the threads, as taught by Johnson. The modification provides an improved system and method for synchronization of multi-thread lanes, thereby to improve execution efficiency by promoting thread convergence. Threads executing common code sections (e.g., inner loop bodies) are urged to converge using instructions inserted at strategic locations in computer code sections. In support of this various instructions and/or code markers (e.g., compiler directives, ISA extensions) are introduced: Predict, Join, Wait, Confirm, Rejoin, and Cancel. Intelligently (e.g., based on profiles and/or policies) inserting these instructions into thread bodies enables the threads in a warp or other group to cooperate with the thread scheduler (e.g., of a streaming multiprocessor in a graphics processing unit) to promote thread convergence. (Johnson, see par. [0005]).
Regarding claim 30. The computer-readable medium of claim 30 performs same step of claim 25. Therefore, claim 30 is further rejected based on the same rationale as claim 25 set forth above and incorporated herein.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KIM THANH THI TRAN whose telephone number is (571)270-1408. The examiner can normally be reached Monday-Friday 8:00am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ALICIA HARRINGTON can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KIM THANH T TRAN/ Examiner, Art Unit 2615
/JAMES A THOMPSON/ Primary Examiner, Art Unit 2615