DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
The Office Action is in response to claims filed 02/11/2024.
Claims 1-20 are pending.
Claim Objections
On line 4 of claim 7, the claim recites “determined based a number …” Examiner believes the underlined portion is a typo and believes that the quotation should read as “determined based on a number.” Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a thread scheduler configured to” in claims 1-10. See MPEP § 2181(l)(A) "The following is a list of non structural generic placeholders that may invoke 35 U.S.C. 112(f): "mechanism for," "module for," "device for," "unit for," "component for," "element for," "member for," "apparatus for," "machine for," or "system for.'"'. Examiner notes that “thread scheduler” is the generic placeholder that invokes 35 U.S.C. 112(f).
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. A review of the disclosure as originally files, hereafter “disclosure”, reveals that the corresponding structure of “task scheduler” is generic computing hardware (see at least ¶ [024] which states “the thread scheduler 106 is a hardware component comprising a plurality of sub-components each associated with a stage of the graph structure”). In accordance with MPEP § 2181 (II)(B), when the corresponding structure of computer implemented mean plus function limitations corresponds to a general purpose computer, an algorithm is required to transform the general purpose computer into a special purpose computer to be sufficient as corresponding structure. Upon further review of the disclosure, applicant has failed to define the algorithm for each of the claimed functions and has instead only provided either verbatim support for the claimed function (which is insufficient as a steps of steps of a corresponding algorithm) or exemplary language that does not make clear the metes and bounds of the algorithm. As such, see rejections under 35 U.S.C. § 112(a) and (b) below.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 3, 4, 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 2, 3, 4, 6, 7, 8, 9, 10, 11, 12, 13, 14, 16, 17, 18, 19, and 20, respectively of patented application 18/210,706 (hereafter ‘706) in view of Agarwal et al. Pub. No. US 20190171497 A1 (hereafter Agarwal) as exemplified in the table below.
Instant Application
18/210,706
1. A graph streaming processing system comprising:
a first processor array;
a second processor; and
a thread scheduler configured to:
dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor;
determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the at least one thread of the second node is identified to be dependent on the output data generated by execution of the at least one thread associated with of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer;
dispatch the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node.
1. A graph streaming processing system comprising:
a first processor array;
a second processor;
and a thread scheduler configured to:
dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in a private data buffer of the second processor;
determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node;
Agarwal ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”
dispatch the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer.
Agarwal ¶ [0081] states “For an embodiment, the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread,”
3. The graph streaming processing system as claimed in claim 1,
wherein the thread scheduler dispatches the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
2. The graph streaming processing system as claimed in claim 1,
wherein the thread scheduler dispatches the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
4. The graph streaming processing system as claimed in claim 1, wherein the first processor array is configured to:
receive the at least one thread of the second node dispatched by the thread scheduler;
retrieve input data required for execution of the at least one thread of the second node from a shared data buffer shared between the first processor array and the second processor
execute the at least one thread of the second node to generate the output data;
and perform at least one of:
writing the output data to the shared data buffer, upon determination that the at least one subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array, wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor;
and writing the output data into the activation data buffer, upon determination that the at least one subsequent thread is being dispatched to the second processor.
3. The graph streaming processing system as claimed in claim 1, wherein the first processor array is configured to:
receive the at least one thread dispatched by the thread scheduler;
retrieve input data required for execution of the at least one thread from a data buffer shared between the first processor array and the second processor;
execute the at least one thread to generate the output data;
and perform at least one of:
writing the output data to a shared data buffer, upon determination that the subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array, wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor;
and writing the output data into the private data buffer, upon determination that the subsequent thread is being dispatched to the second processor.
5. The graph streaming processing system as claimed in claim 1,
wherein the second processor comprises a write unit to enable the first processor array to write the output data into the activation data buffer
4. The graph streaming processing system as claimed in claim 1,
wherein the second processor comprises a write unit to enable the first processor array to write the output data into the private data buffer.
7. The graph streaming processing system as claimed in claim 1,
wherein the activation data buffer is configured to store a predetermined number of data units segregated from a plurality of data units, wherein the data unit corresponds to a slice of the output data and wherein the predetermined number of data units is determined based a number of data units required to execute the at least one thread of the second node.
6. The graph streaming processing system as claimed in claim 1,
wherein the private data buffer is configured to store a predetermined number of data units segregated from a plurality of data units, wherein the data unit corresponds to a slice of the output data and wherein the predetermined number of data units is determined based a number of data units required to execute the at least one thread of the second node.
8. The graph streaming processing system as claimed in claim 1, wherein the thread scheduler determines the sufficiency of the at least one data unit for execution of the at least one thread of the second node by:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread of the first node;
detecting storing of the at least one data unit in the activation data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
7. The graph streaming processing system as claimed in claim 1, wherein the thread scheduler determines the sufficiency of the at least one data unit for execution of the at least one thread of the second node by:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the private data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
9. The graph streaming processing system as claimed in claim 8, wherein the thread scheduler is further configured to:
dispatch at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
8. The graph streaming processing system as claimed in claim 7, wherein the thread scheduler is further configured to:
dispatch at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
10. The graph streaming processing system as claimed in claim 1, wherein to determine to dispatch the at least one subsequent thread of the first node for execution, the thread scheduler is configured to:
detect execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer;
evaluate the availability of the predefined threshold buffer size on the activation data buffer;
and perform one of:
dispatch the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
9. The graph streaming processing system as claimed in claim 1, wherein to determine to dispatch the at least one subsequent thread of the first node for execution, the thread scheduler is configured to:
detect execution of the at least one thread of the second node by consuming the at least one data unit stored in the private data buffer;
evaluate the availability of the predefined threshold buffer size on the private data buffer;
and perform one of:
dispatch the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
11. A method comprising:
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer;
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node.
10. A method comprising:
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in a private data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node;
Agarwal ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer.
Agarwal ¶ [0081] states “For an embodiment, the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread,”
12. The method as claimed in claim 11,
wherein dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the
second processor.
11. The method as claimed in claim 10,
wherein dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
13. The method as claimed in claim 11, wherein determining the sufficiency of the at least one data unit for execution of the at least one thread of the second node comprises:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least
one thread of the first node;
detecting storing of the at least one data unit in the activation data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
12. The method as claimed in claim 10, wherein determining the sufficiency of the at least one data unit for execution of the at least one thread of the second node comprises:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the private data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
14. The method as claimed in claim 11, further comprising:
dispatching the at least one subsequent thread of the first node to generate at least
one subsequent data unit, before dispatching the at least one thread of the second node for
execution, upon determining that the at least one data unit comprises insufficient data for
execution of the at least one thread of the second node.
13. The method as claimed in claim 10, further comprising:
dispatching at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
15. The method as claimed in claim 11, wherein determining to dispatch the at least one subsequent thread of the first node for execution further comprises:
detecting execution of the at least one thread of the second node by consuming the
at least one data unit stored in the activation data buffer;
evaluating the availability of the predefined threshold buffer size on the activation
data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining
that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining
that the predefined threshold buffer size is not available.
14. The method as claimed in claim 10, wherein determining to dispatch the at least one subsequent thread of the first node for execution further comprises:
detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the private data buffer;
evaluating the availability of the predefined threshold buffer size on the private data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
16. A non-transitory computer-readable medium having program instructions stored thereon, wherein the program instructions, when executed by a thread-scheduler of a graph streaming processing system, facilitate:
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least
one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer;
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of
the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node.
16. A non-transitory computer-readable medium having program instructions stored thereon, wherein the program instructions, when executed by a thread-scheduler of a graph streaming processing system, facilitate:
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in a private data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node;
Agarwal ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer.
Agarwal ¶ [0081] states “For an embodiment, the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread,”
17. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions further facilitate: dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
17. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions further facilitate: dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
18. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions configured to determine the sufficiency of the at least one data unit further facilitate:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the activation data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
18. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions configured to determine the sufficiency of the at least one data unit further facilitate:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the private data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
19. The non-transitory computer-readable medium as claimed in claim 16, wherein the
program instructions further facilitate:
dispatching at least one subsequent thread of the first node to generate at least one
subsequent data unit, before dispatching the at least one thread of the second node for
execution, upon determining that the at least one data unit comprises insufficient data for
execution of the at least one thread of the second node.
19. The non-transitory computer-readable medium as claimed in claim 16, wherein the program instructions further facilitate:
dispatching at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
20. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions configured to determine to dispatch the at least one subsequent thread of the first node further facilitate:
detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer;
evaluating the availability of the predefined threshold buffer size on the activation
data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining
that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
20. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions configured to determine to dispatch the at least one subsequent thread of the first node further facilitate:
detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the private data buffer;
evaluating the availability of the predefined threshold buffer size on the private data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
Although the claims at issue are not identical, they are not patentably distinct from each other because the instant application and ‘447 overlap in scope. For example, the claims of the instant application recite “activation data buffer”, and the claims of ‘447 recite “private data buffer.” Although “activation” and “private” are not identical, they are not patentably distinct. Regarding claims 1, 11, and 16, it would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine the storing of output data in the output command buffer, the child node checking whether parent nodes are completed, and allocating required memory of Agarwal with the determining of sufficiency of ‘447. A person having ordinary skill in the art would have been motivated to perform this combination “for improving the thread scheduling mechanisms during graph processing” (¶ [0037]). Resolving dependencies of parent nodes before dispatching child nodes provides benefits of such as ensuring no conflicts of child nodes accessing data that has not been output.
Claims 1, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 2, 3, 4, 3, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, and 18, respectively, of copending application 19/445,999 (hereafter ‘999) in view of Agarwal et al. Pub. No. US 20190171497 A1 (hereafter Agarwal) as exemplified in the table below.
Instant Application
19/445,999
1. A graph streaming processing system comprising:
a first processor array;
a second processor; and
a thread scheduler configured to:
dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor;
determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the at least one thread of the second node is identified to be dependent on the output data generated by execution of the at least one thread associated with of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer;
dispatch the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node.
1. A processing system comprising:
a first processor array;
a second processor;
and a thread scheduler configured to:
dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in a private data buffer of the second processor;
determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node;
Agarwal ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”
dispatch the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer;
Agarwal ¶ [0081] states “For an embodiment, the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread,”
wherein the second processor is configured to: receive the at least one thread dispatched by the thread scheduler; retrieve input data required for execution of the at least one thread from at least one of the shared data buffer and the private data buffer; execute the at least one thread to generate the output data; and perform at least one of: writing the output data into the private data buffer, upon determination that the subsequent thread is being dispatched to the second processor; and writing the output data into the shared data buffer, upon determination that the subsequent thread is being dispatched to the first processor array.
3. The graph streaming processing system as claimed in claim 1,
wherein the thread scheduler dispatches the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
2. The processing system as claimed in claim 1,
wherein the thread scheduler dispatches the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
4. The graph streaming processing system as claimed in claim 1, wherein the first processor array is configured to:
receive the at least one thread of the second node dispatched by the thread scheduler;
retrieve input data required for execution of the at least one thread of the second node from a shared data buffer shared between the first processor array and the second processor
execute the at least one thread of the second node to generate the output data;
and perform at least one of:
writing the output data to the shared data buffer, upon determination that the at least one subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array, wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor;
and writing the output data into the activation data buffer, upon determination that the at least one subsequent thread is being dispatched to the second processor.
3. The processing system as claimed in claim 1, wherein the first processor array is configured to:
receive the at least one thread dispatched by the thread scheduler;
retrieve input data required for execution of the at least one thread from a data buffer shared between the first processor array and the second processor;
execute the at least one thread to generate the output data;
and perform at least one of:
writing the output data to a shared data buffer, upon determination that the subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array, wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor;
and writing the output data into the private data buffer, upon determination that the subsequent thread is being dispatched to the second processor.
5. The graph streaming processing system as claimed in claim 1,
wherein the second processor comprises a write unit to enable the first processor array to write the output data into the activation data buffer
4. The processing system as claimed in claim 1,
wherein the second processor comprises a write unit to enable the first processor array to write the output data into the private data buffer.
6. The graph streaming processing system as claimed in claim 4,
wherein the second processor is configured to:
receive the at least one thread of the second node dispatched by the thread scheduler;
retrieve input data required for execution of the at least one thread of the second node from at least one of the shared data buffer and the activation data buffer;
execute the at least one thread to generate the output data; and
perform at least one of:
writing the output data into the activation data buffer, upon determination that the subsequent thread is being dispatched to the second processor;
And writing the output data into the shared data buffer, upon determination that the at least one subsequent thread is being dispatched to the first processor array.
3. The processing system as claimed in claim 1, wherein the first processor array is configured to: receive the at least one thread dispatched by the thread scheduler; retrieve input data required for execution of the at least one thread from a data buffer shared between the first processor array and the second processor; execute the at least one thread to generate the output data; and perform at least one of: writing the output data to a shared data buffer, upon determination that the subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array, wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor; and writing the output data into the private data buffer, upon determination that the subsequent thread is being dispatched to the second processor.
1. A processing system comprising: a first processor array; a second processor; and a thread scheduler configured to: dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in a private data buffer of the second processor; determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node; dispatch the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient; and determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer;
wherein the second processor is configured to:
receive the at least one thread dispatched by the thread scheduler;
retrieve input data required for execution of the at least one thread from at least one of the shared data buffer and the private data buffer;
execute the at least one thread to generate the output data;
and perform at least one of:
writing the output data into the private data buffer, upon determination that the subsequent thread is being dispatched to the second processor;
and writing the output data into the shared data buffer, upon determination that the subsequent thread is being dispatched to the first processor array.
7. The graph streaming processing system as claimed in claim 1,
wherein the activation data buffer is configured to store a predetermined number of data units segregated from a plurality of data units, wherein the data unit corresponds to a slice of the output data and wherein the predetermined number of data units is determined based a number of data units required to execute the at least one thread of the second node.
5. The processing system as claimed in claim 1,
wherein the private data buffer is configured to store a predetermined number of data units segregated from a plurality of data units, wherein the data unit corresponds to a slice of the output data and wherein the predetermined number of data units is determined based a number of data units required to execute the at least one thread of the second node.
8. The graph streaming processing system as claimed in claim 1, wherein the thread scheduler determines the sufficiency of the at least one data unit for execution of the at least one thread of the second node by:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread of the first node;
detecting storing of the at least one data unit in the activation data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
6. The processing system as claimed in claim 1, wherein the thread scheduler determines the sufficiency of the at least one data unit for execution of the at least one thread of the second node by:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the private data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
9. The graph streaming processing system as claimed in claim 8, wherein the thread scheduler is further configured to:
dispatch at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
7. The processing system as claimed in claim 6, wherein the thread scheduler is further configured to:
dispatch at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
10. The graph streaming processing system as claimed in claim 1, wherein to determine to dispatch the at least one subsequent thread of the first node for execution, the thread scheduler is configured to:
detect execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer;
evaluate the availability of the predefined threshold buffer size on the activation data buffer;
and perform one of:
dispatch the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
8. The processing system as claimed in claim 1, wherein to determine to dispatch the at least one subsequent thread of the first node for execution, the thread scheduler is configured to:
detect execution of the at least one thread of the second node by consuming the at least one data unit stored in the private data buffer;
evaluate the availability of the predefined threshold buffer size on the private data buffer;
and perform one of:
dispatch the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
11. A method comprising:
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer;
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node.
9. A method comprising:
dispatching, by a thread scheduler of a processing system, one thread associated with a first node, to one of a first processor array and a second processor of the processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in a private data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node;
Agarwal ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”
dispatching, by the thread scheduler, the at least one thread of the second node,to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer;
Agarwal ¶ [0081] states “For an embodiment, the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread,”
receiving, by the second processor, the at least one thread dispatched by the thread scheduler; retrieving, by the second processor, input data required for execution of the at least one thread from at least one of the shared data buffer and the private data buffer; executing, by the second processor, the at least one thread to generate the output data; and performing at least one of: writing the output data into the private data buffer, upon determination that the subsequent thread is being dispatched to the second processor; and writing the output data into the shared data buffer, upon determination that the subsequent thread is being dispatched to the first processor array.
12. The method as claimed in claim 11,
wherein dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the
second processor.
10. The method as claimed in claim 9,
wherein dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
13. The method as claimed in claim 11, wherein determining the sufficiency of the at least one data unit for execution of the at least one thread of the second node comprises:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least
one thread of the first node;
detecting storing of the at least one data unit in the activation data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
11. The method as claimed in claim 9, wherein determining the sufficiency of the at least one data unit for execution of the at least one thread of the second node comprises:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the private data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
14. The method as claimed in claim 11, further comprising:
dispatching the at least one subsequent thread of the first node to generate at least
one subsequent data unit, before dispatching the at least one thread of the second node for
execution, upon determining that the at least one data unit comprises insufficient data for
execution of the at least one thread of the second node.
12. The method as claimed in claim 9, further comprising:
dispatching at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
15. The method as claimed in claim 11, wherein determining to dispatch the at least one subsequent thread of the first node for execution further comprises:
detecting execution of the at least one thread of the second node by consuming the
at least one data unit stored in the activation data buffer;
evaluating the availability of the predefined threshold buffer size on the activation
data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining
that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining
that the predefined threshold buffer size is not available.
13. The method as claimed in claim 9, wherein determining to dispatch the at least one subsequent thread of the first node for execution further comprises:
detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the private data buffer;
evaluating the availability of the predefined threshold buffer size on the private data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
16. A non-transitory computer-readable medium having program instructions stored thereon, wherein the program instructions, when executed by a thread-scheduler of a graph streaming processing system, facilitate:
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least
one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer;
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of
the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node.
14. A non-transitory computer-readable medium having program instructions stored thereon, wherein the program instructions, when executed by a thread-scheduler of a processing system, facilitate:
dispatching, by a thread scheduler of a processing system, one thread associated with a first node, to one of a first processor array and a second processor of the processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in a private data buffer of the second processor;
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node;
Agarwal ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient;
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer;
Agarwal ¶ [0081] states “For an embodiment, the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread,”
receiving, by the second processor, the at least one thread dispatched by the thread scheduler; retrieving, by the second processor, input data required for execution of the at least one thread from at least one of the shared data buffer and the private data buffer; executing, by the second processor, the at least one thread to generate the output data; and performing at least one of: writing the output data into the private data buffer, upon determination that the subsequent thread is being dispatched to the second processor; and writing the output data into the shared data buffer, upon determination that the subsequent thread is being dispatched to the first processor array.
17. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions further facilitate: dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
15. The non-transitory computer-readable medium as claimed in claim 14,
wherein the program instructions further facilitate: dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.
18. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions configured to determine the sufficiency of the at least one data unit further facilitate:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the activation data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
16. The non-transitory computer-readable medium as claimed in claim 14,
wherein the program instructions configured to determine the sufficiency of the at least one data unit further facilitate:
detecting an execution of the at least one thread of the first node;
detecting generation of the at least one data unit from the execution of the at least one thread;
detecting storing of the at least one data unit in the private data buffer;
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.
19. The non-transitory computer-readable medium as claimed in claim 16, wherein the
program instructions further facilitate:
dispatching at least one subsequent thread of the first node to generate at least one
subsequent data unit, before dispatching the at least one thread of the second node for
execution, upon determining that the at least one data unit comprises insufficient data for
execution of the at least one thread of the second node.
17. The non-transitory computer-readable medium as claimed in claim 14, wherein the program instructions further facilitate:
dispatching at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.
20. The non-transitory computer-readable medium as claimed in claim 16,
wherein the program instructions configured to determine to dispatch the at least one subsequent thread of the first node further facilitate:
detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer;
evaluating the availability of the predefined threshold buffer size on the activation
data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining
that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
18. The non-transitory computer-readable medium as claimed in claim 14,
wherein the program instructions configured to determine to dispatch the at least one subsequent thread of the first node further facilitate:
detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the private data buffer;
evaluating the availability of the predefined threshold buffer size on the private data buffer;
and performing one of:
dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available;
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available.
Although the claims at issue are not identical, they are not patentably distinct from each other because the instant application and ‘999 overlap in scope. For example, the claims of the instant application recite “activation data buffer”, and the claims of ‘999 recite “private data buffer.” Although “activation” and “private” are not identical, they are not patentably distinct. Regarding claims 1, 11, and 16, it would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine the storing of output data in the output command buffer, the child node checking whether parent nodes are completed, and allocating required memory of Agarwal with the determining of sufficiency of ‘999. A person having ordinary skill in the art would have been motivated to perform this combination “for improving the thread scheduling mechanisms during graph processing” (¶ [0037]). Resolving dependencies of parent nodes before dispatching child nodes provides benefits of such as ensuring no conflicts of child nodes accessing data that has not been output.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-10 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claims 1-10 recite a "thread scheduler" which invokes 35 U.S.C. § 112(f), see claim interpretation above. The disclosure does not recite sufficient corresponding structure (in this instance computer + algorithm), again see claim interpretation above. As such, and in accordance with MPEP § 2181 (ll)(B), last paragraph "When a claim containing a computer-implemented 35 U.S.C. 112(f) claim limitation is found to be indefinite under 35 U.S.C. 112(b) for failure to disclose sufficient corresponding structure (e.g., the computer and the algorithm) in the specification that performs the entire claimed function, it will also lack written description under 35 U.S.C. 112(a)."
Claims 1-10 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim limitation "thread scheduler" invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The disclosure fails to disclose sufficient corresponding structure (in this instance computer + algorithm), see claim interpretation above. As such, and in accordance with MPEP § 2181 (ll)(B) "For a computer-implemented 35 U.S.C. 112(f) claim limitation, the specification must disclose an algorithm for performing the claimed specific computer function, or else the claim is indefinite under 35 U.S.C. 112(b).". Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph.
Applicant may:
(a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
(b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
(a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception, an abstract idea, and it has not been integrated into practical application and the claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below.
Step 1: Claims 1-10 are directed to a graph streaming system and falls within the statutory class of machine. Claims 11-15 are directed to a method and fall within the statutory class of process. Claims 16-20 are directed to a non-transitory computer-readable medium and fall within the statutory class of articles of manufacture. Therefore, “Are the claims to a process, machine, manufacture or composition of matter?” Yes.
Step 2A Prong 1:
Claims 1, 11, and 16: The limitations “determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the at least one thread of the second node is identified to be dependent on the output data generated by execution of the at least one thread associated with of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer,” “upon determining that the at least one data unit is sufficient,” and “determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node” are a mental process. Determining sufficiency and determining to dispatch a subsequent thread involve steps of observation, evaluation, and forming a judgement. It is understood that these limitations are to be performed within a computer environment, however, it can also entirely be performed in the mind.
Therefore, Yes, claims 1, 11, and 16 recite a judicial exception. Step 2A Prong 2 will evaluate whether the claims integrate the judicial exception into a practical application.
Step 2A Prong 2:
Claims 1, 11, and 16: The judicial exception is not integrated into a practical application. Claims 1, 11, and 16 recites the following additional elements – “dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor” and “dispatch the at least one thread of the second node, to one of the first processor array and the second processor.” These limitations are considered insignificant extra-solution activities of data transmission and data storage (MPEP § 2106.05(g)). Claims 1, 11, and 16 also recite “a first processor array,” “a second processor,” and “a thread scheduler.” In light of the 112(f) claim interpretation, the thread scheduler is generic computer hardware. Claim 16 also recites “a non-transitory computer readable medium having program instructions stored thereon, wherein the program instructions, when executed by a thread-scheduler of a graph streaming processing system, facilitate.” These claim limitations are generic computing components used as a means to apply an exception (MPEP § 2106.05(f)). These additional elements do not integrate the judicial exception into a practical application.
Therefore, “Do the claims recite additional elements that integrate the judicial exception in a practical application?” No, these additional elements do not integrate the abstract idea into a practical application and they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
After having evaluated the inquiries set forth in Steps 2A Prong 1 and 2, it has been concluded that claims 1, 11, and 16 not only recite a judicial exception but that the claims are directed to the judicial exception as the judicial exception has not been integrated into practical application.
Step 2B:
Claims 1, 11, and 16: The claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above, the additional elements only amount to insignificant extra-solution activity and means to apply an exception. When reevaluating the means to apply an exception, alone or in combination with the other limitations, no inventive concept that amounts to significantly more was found. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” and “Storing and retrieving information in memory” are well understood, routine, and conventional computer functions. Dispatching threads is transmitting data over a network. Storing data in an activation data buffer is storing information in memory. Therefore, the claims fail Step 2B.
Therefore, “Do the claims recite additional elements that amount to significantly more than the judicial exception? No, these additional elements, alone or in combination, do not amount to significantly more than the judicial exception.
Having concluded analysis with in the provided framework, claims 1, 11, and 16 do not recite eligible subject matter under 35 U.S.C. § 101.
With regard to claim 2 it recites “wherein the second processor writes and reads data from the activation data buffer, and the first processor array only writes data to the activation data buffer.” This limitation is considered insignificant extra-solution activity of data reading and data storage. It does not integrate the judicial exception into a practical application, so the claim fails Step 2A Prong 2. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Storing and retrieving information in memory” is a well understood, routine, and conventional computer function. Therefore, the claim fails Step 2B. Therefore claim 2 does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claims 3, 12, and 17 it recites “wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor.” This limitation is a mental process, so the claims recite a judicial exception and fail Step 2A Prong 1. The claims additionally recite “wherein the thread scheduler dispatches the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread.” This limitation is an insignificant extra-solution activity of data transmission (MPEP § 2106.05(g)). It does not integrate the judicial exception into practical application, so the claims fail Step 2A Prong 2. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” is a well understood, routine, and conventional computer function. Therefore, the claims fail Step 2B. Therefore claims 3, 12, and 17 do not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim 4 it recites “upon determination that the at least one subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array” and “upon determination that the at least one subsequent thread is being dispatched to the second processor.” These limitations are a mental process, so the claim recites a judicial exception and fails Step 2A Prong 1. The claim additionally recites “receive the at least one thread of the second node dispatched by the thread scheduler,” “retrieve input data required for execution of the at least one thread of the second node from a shared data buffer shared between the first processor array and the second processor,” “writing the output data to the shared data buffer,” and “writing the output data into the activation data buffer.” These limitations are insignificant extra-solution activity of data gathering, data transmission, and data storage (MPEP § 2106.05(g)). The claim also recites “execute the at least one thread of the second node to generate the output data.” This limitation is means to apply an exception (MPEP § 2106.05(f)). The claim also recites “wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor.” This limitation is considered a field of use/technological environment (MPEP § 2106.05(h)). These additional elements do not integrate the judicial exception into a practical application, so the claims fail Step 2A Prong 2. When reevaluating the additional elements, alone or in combination, no inventive concept that amounts to significantly more was found. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” and “Storing and retrieving information in memory” are well understood, routine, and conventional computer functions. Therefore, the claims fail Step 2B. Therefore claim 4 does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim 5, it recites “wherein the second processor comprises a write unit to enable the first processor array to write the output data into the activation data buffer.” This limitation further limits the generic computing component in claim 1, so the claim is also considered generic computing components used as a means to apply an exception (MPEP § 2106.05(f)). It does not integrate the judicial exception into a practical application, so the claim fails Step 2A Prong 2. When reevaluating the means to apply an exception, alone or in combination with other claim limitations, no inventive concept that amounts to significantly more was found. Therefore, the claim fails Step 2B. Therefore claim 5 does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claims 6, it recites “upon determination that the subsequent thread is being dispatched to the second processor” and “upon determination that the at least one subsequent thread is being dispatched to the first processor array.” These limitations are a mental process, so the claim recites a judicial exception and fails Step 2A Prong 1. The claim also recites “receive the at least one thread of the second node dispatched by the thread scheduler,” “retrieve input data required for execution of the at least one thread of the second node from at least one of the shared data buffer and the activation data buffer,” “writing the output data into the activation data buffer,” and “writing the output data into the shared data buffer.” These limitations are insignificant extra-solution activities of data gathering, data transmission, and data storage (MPEP § 2106.05(g)). The claim also recites “execute the at least one thread to generate the output data.” This limitation is means to apply an exception (MPEP § 2106.05(f)). These additional elements do not integrate the judicial exception into a practical application, so the claim fails Step 2A Prong 2. When reevaluating the additional elements, alone or in combination, no inventive concept that amounts to significantly more was found. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” and “Storing and retrieving information in memory” are well understood, routine, and conventional computer functions. Therefore, the claims fail Step 2B. Therefore claim 6 does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim 7, it recites “wherein the predetermined number of data units is determined based a number of data units required to execute the at least one thread of the second node.” This limitation is considered a mental process, so the claims recite a judicial exception and fail Step 2A Prong 1. The claim also recites “wherein the activation data buffer is configured to store a predetermined number of data units segregated from a plurality of data units, wherein the data unit corresponds to a slice of the output data.” This limitation is considered field of use/technological environment (MPEP § 2106.05(h)). It does not integrate the judicial exception into a practical application, so the claim fails Step 2A Prong 2. When reevaluating the field of use/technological environment, alone or in combination with other claim limitations, no inventive concept that amounts to significantly more was found. Therefore, the claim fails Step 2B. Therefore claim 7 does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claims 8, 13, and 18 it recites “wherein the thread scheduler determines the sufficiency of the at least one data unit for execution of the at least one thread of the second node by,” “detecting an execution of the at least one thread of the first node,” “detecting generation of the at least one data unit from the execution of the at least one thread of the first node,” “detecting storing of the at least one data unit in the activation data buffer,” and “determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node.” These limitations further limit the mental process in claims 1, 11, and 16 respectively and are also considered as such. Therefore, the claims recite a judicial exception and fail Step 2A Prong 1. The claims do not include any additional elements that integrate the judicial exception into a practical application, so the claims fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claims fail Step 2B. Therefore claims 8, 13, and 18 do not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claims 9, 14, and 19 it recites “upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node.” This limitation is a mental process, so the claims recite a judicial exception and fail Step 2A Prong 1. The claims further recite “dispatch at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution.” This limitation is considered insignificant extra-solution activity of data transmission (MPEP § 2106.05(g)). It does not integrate the judicial exception into a practical application, so the claims fail Step 2A Prong 2. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” is a well understood, routine, and conventional computer function. Therefore, the claims fail Step 2B. Therefore claims 9, 14, and 19 do not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claims 10, 15, and 20 it recites “detect execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer,” “evaluate the availability of the predefined threshold buffer size on the activation data buffer,” “upon determining that the predefined threshold buffer size is available,” and “upon determining that the predefined threshold buffer size is not available.” These limitations are considered a mental process, so the claims recite a judicial exception and fail Step 2A Prong 1. The claims additionally recite “dispatch the at least one subsequent thread of the first node” and “dispatching at least one subsequent thread of the second node.” These limitations are considered insignificant extra-solution activity of data transmission (MPEP § 2106.05(g)). It does not integrate the judicial exception into a practical application, so the claims fail Step 2A Prong 2. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” is a well understood, routine, and conventional computer function. Therefore, the claims fail Step 2B. Therefore claims 10, 15, and 20 do not recite patent eligible subject matter under 35 U.S.C. 101.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5, 7-8, 10-11, 13, 15-16, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Teng et al. Pat. No. US 20190114534 A1 (hereafter Teng) in view of Agarwal et al. Pat. No. US 20190171497 A1 (hereafter Agarwal).
With regard to claim 1, Teng teaches a graph streaming processing system comprising (¶ [0055] states “FIG. 6 shows a neural network processing system, along with data flow and control signaling between a first processor element 602, a neural network accelerator 238, and a second processor element 604.”):
a first processor array (¶ [0055] states “FIG. 6 shows a neural network processing system, along with data flow and control signaling between a first processor element 602, a neural network accelerator 238, and a second processor element 604.” ¶ [0026] states “As used herein, a “processor element” can be a processor core of a computer system, heterogeneous processor circuits, or threads executing on one or more processor cores or processor circuits.” Examiner’s Note: the first and second processor elements 602 and 604 are considered the first processor array);
a second processor (¶ [0055] states “FIG. 6 shows a neural network processing system, along with data flow and control signaling between a first processor element 602, a neural network accelerator 238, and a second processor element 604.” Examiner’s Note: the neural network accelerator is the second processor);
and a thread scheduler configured to: dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor (¶ [0055] states “the first processor element can proceed in staging (2) an input data set 610 to the RAM 608 for processing by the neural network accelerator 238. Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” ¶ [0056] states “The RAMs 608 and 612 can be a single RAM 226 such as shown in FIG. 5 or physically separate RAMs, depending on implementation requirements.” ¶ [0052] states “In processing the first per-layer instruction, the neural network accelerator 238 reads input data from a first portion of the B/C buffer 518 in the RAM 226 and writes output data to a second portion of the B/C buffer in the RAM.” Examiner’s Note: RAM 608 and RAM 226 are the same RAM. B/C buffer 518 is the activation data buffer of the second processor because the data stored there is for neural network accelerator. The first processor executing indicates the thread executing was dispatched);
determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the at least one thread of the second node is identified to be dependent on the output data generated by execution of the at least one thread associated with of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” Examiner’s Note: the determination of when the first processor element has written the input dataset into the RAM for the neural network accelerator is when determining sufficiency takes place);
dispatch the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” ¶ [0056] states “[0056] The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.” Examiner’s Note: signaling the neural network accelerator to start performing operations is dispatching the thread of the second processor);
and determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” ¶ [0058] states “In response to receiving the queue-full signal from the first processor element, the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612, and then signals (1) the first processor element that the shared memory queue is empty.” Examiner’s Note: when the neural network accelerator completes processing, its associated activation data buffer becomes available. The neural network accelerator then signals the processor element 602. After processor element 602 and 604 finish processing data, processor element 604 signals processor element 602 to then place more data in the RAM 608, or activation data buffer. Before placing more data in RAM 602, a determination takes place).
Teng does not explicitly teach a graph streaming processing system, a thread scheduler, and threads dependent on other threads based on a graph.
However, in an analogous art, Agarwal teaches a graph streaming processing system comprising (¶ [0022] states “The described embodiments are embodied in methods, apparatuses and systems for accelerating graph stream processing.”):
a first processor array (¶ [0041] states “the GSP 410 includes a thread manager 420 that manages dispatching of threads of a plurality of thread processors 430,”);
a second processor;
and a thread scheduler configured to: dispatch at least one thread associated with a first node, to one of the first processor array and the second processor, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor (¶ [0023] states “a node includes one or more threads with each thread running the same code-block but (possibly) on different data and producing (possibly) different output data.” ¶ [0043] states “This mode of operation includes resolving dependencies of a child thread after dispatching the child thread. The centralized dispatcher (thread manager 620) maintains the status of all running threads.” ¶ [0050] states “As shown in the flow chart of FIG. 7, a first step 710 includes dispatching of threads of a cousin (producing) node.” Examiner’s Note: the cousin node is the first node);
determine sufficiency of the at least one data unit for executing at least one thread of a second node, wherein the at least one thread of the second node is identified to be dependent on the output data generated by execution of the at least one thread associated with of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer (¶ [0039] states “The hardware scheduler tracks the status of the currently running threads in a scorecard.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.” ¶ [0024] states “A thread may be dependent on data generated by other threads of the same node, and/or data generated by threads of other nodes.” ¶ [0027] states “That is, the thread of a node can only be dependent on an older thread. For an embodiment, a thread can only be dependent on threads of an earlier stage, or threads of the same stage that have been dispatched earlier.” See FIG. 2. ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014 … Note that while the command buffer 1014 is the output command buffer for the stage 1012, the command buffer 1014 is the input command buffer for the stage 1015.” Examiner’s Note: child node 205 is dependent on the output data of cousin node 204. Child node 205 is analogous to the second node. Cousin node 204 is analogous to the first node. When the child node checks whether the Mth thread of the cousin node is completed, it is determining whether output data of the cousin node is sufficient);
dispatch the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient (¶ [0050] states “A second step 720 includes dispatching threads of a child node” and “Upon completion of the Mth thread of the cousin node, a fifth step 750 includes the Nth thread of the child node proceeding with further execution after receiving the response to the dependency from the cousin node.” ¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).”);
and determine to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node (¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.” ¶ [0081] states “the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread.” ¶ [0047] states “The thread scheduler keeps on dispatching while the thread scheduler has the required resources in the processing cores.” Examiner’s Note: the required amount of memory in the output command buffer is the predefined threshold buffer size).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the graph streaming processing system and thread scheduler of Agarwal with the processor elements and neural network accelerator of Teng. As a result, the system handles dependency resolution between threads. A person having ordinary skill in the art would have been motivated to make this combination for the purpose of reducing execution time and increasing performance (¶ [0047] states “The thread scheduler keeps on dispatching while the thread scheduler has the required resources in the processing cores. This fills up the thread slots in the multi-threaded execution cores and allows each of the threads to determine execution based on their own dependencies. The execution time reduces considerably which results in higher performance”). See also ¶ [0048] – [0049] for additional improvements.
With regard to claim 2, Teng and Agarwal teach the graph streaming processing system as claimed in claim 1. Teng additionally teaches wherein the second processor writes and reads data from the activation data buffer, and the first processor array only writes data to the activation data buffer (¶ [0052] states “In processing the first per-layer instruction, the neural network accelerator 238 reads input data from a first portion of the B/C buffer 518 in the RAM 226 and writes output data to a second portion of the B/C buffer in the RAM.” ¶ [0055] states “the first processor element can proceed in staging (2) an input data set 610 to the RAM 608 for processing by the neural network accelerator 238.” Examiner’s Note: the B/C buffer is part of the activation data buffer. The first processor element only writes input data set to B/C buffer 518 in RAM 608. RAM 608 is the RAM 226).
With regard to claim 5, Teng and Agarwal teach the graph streaming processing system as claimed in claim 1. Teng additionally teaches wherein the second processor comprises a write unit to enable the first processor array to write the output data into the activation data buffer (See FIG. 2. ¶ [0032] states “The support circuits 214 include various devices that cooperate with the microprocessor 212 to manage data flow between the microprocessor 212, the system memory 216, the storage 218, the hardware accelerator 116, or any other peripheral device.” ¶ [0036] states “The static region 234 includes support circuits 240 for providing an interface to the peripheral bus 215, the NVM 224, and the RAM 226.” ¶ [0039] states “The neural network accelerator 238 can access the RAM 226 through the memory controllers 310.” Examiner’s Note: the support circuits 240 are the write unit that allows the processing element 602 to write output data into the activation data buffer. RAM 226 is the activation data buffer).
With regard to claim 7, Teng and Agarwal teach the graph streaming processing system as claimed in claim 1. Teng additionally teaches wherein the activation data buffer is configured to store a predetermined number of data units segregated from a plurality of data units, wherein the data unit corresponds to a slice of the output data and wherein the predetermined number of data units is determined based a number of data units required to execute the at least one thread of the second node (¶ [0052] states “In processing the first per-layer instruction, the neural network accelerator 238 reads input data from a first portion of the B/C buffer 518 in the RAM 226 and writes output data to a second portion of the B/C buffer in the RAM” and “The neural network accelerator thereafter alternates between portions of the B/C buffer used for input and output data with each successive per-layer instruction.” Examiner’s Note: a portion of the B/C buffer is a predetermined number of data units that is segregated from a second portion of the B/C buffer).
Agarwal additionally teaches wherein the activation data buffer is configured to store a predetermined number of data units segregated from a plurality of data units, wherein the data unit corresponds to a slice of the output data and wherein the predetermined number of data units is determined based a number of data units required to execute the at least one thread of the second node (¶ [0081] states “For an embodiment, the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread.” ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014” and “Note that while the command buffer 1014 is the output command buffer for the stage 1012, the command buffer 1014 is the input command buffer for the stage 1015.” ¶ [0036] states “dependencies of child threads are resolved before the child thread is dispatched.” ¶ [0040] states “the execution of the child thread 1 is initiated or dispatched at a time 310 at which the sibling (identical twin or not) thread 1 and the cousin thread 1 have completed their processing.” Examiner’s Note: the write commands are the output data. A single write command is a data unit and a slice of the output data. The hardware management of the command buffer allocates the required amount of memory for storing the output write commands. The required amount of memory is allocated before output data is written to the output command buffer, so the buffer has a capacity for a predetermined number of data units. Since dependencies of child threads are resolved before the child thread is dispatched, the number of data units needed to execute at least one thread of the second node is equivalent to the capacity of data units in the output command buffer. Therefore, the predetermined number of data units that the buffer stores is based on a number of data units required to execute at least one thread of the second node).
With regard to claim 8, Teng and Agarwal teach the graph streaming processing system as claimed in claim 1. Teng additionally teaches detecting storing of the at least one data unit in the activation data buffer (¶ [0052] states “In processing the first per-layer instruction, the neural network accelerator 238 reads input data from a first portion of the B/C buffer 518 in the RAM 226 and writes output data to a second portion of the B/C buffer in the RAM.” ¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.”).
Agarwal additionally teaches wherein the thread scheduler determines the sufficiency of the at least one data unit for execution of the at least one thread of the second node by: detecting an execution of the at least one thread of the first node (¶ [0039] states “a hardware scheduler (also referred to as a thread manager) is responsible for issuing threads for execution.” ¶ [0040] states “the execution of the child thread 1 is initiated or dispatched at a time 310 at which the sibling (identical twin or not) thread 1 and the cousin thread 1 have completed their processing.” Examiner’s Note: by issues a thread for execution, the scheduler has detected an execution of a thread. Cousin thread is a thread that belongs to a cousin node, which is the first node);
detecting generation of the at least one data unit from the execution of the at least one thread of the first node (¶ [0039] states “The hardware scheduler tracks the status of the currently running threads in a scorecard.” ¶ [0055] states “each thread includes a set of instructions operating on the plurality of thread processors 1030 and operating on a set of data and producing output data.” ¶ [0057] states “each stage 1012, 1015 includes the input command buffer parser 1013, 1016, wherein the command buffer parser 1013, 1016 generates the threads of the stage 1012, 1015 based upon commands of a command buffer 1011, 1014 located between the stage and the previous stage.” ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014. Note that the stage 1012 includes a write pointer (WP) for the output command buffer 1014. For an embodiment, the write pointer (WP) updates in a dispatch order. That is, for example, the write pointer (WP) updates when the thread T0 spawned commands are written, even if the thread T0 spawned commands are written after the T1 spawned commands are written.” ¶ [0066] states “The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated.” See FIG. 10. Examiner’s Note: the threads generate output data. Output data is the data unit. In FIG. 10, output data is stored in command buffer 1014. When the thread manager updates the status of thread and based on which code portion is being executed, the thread manager detects the generation of the data unit);
detecting storing of the at least one data unit in the activation data buffer (¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014. Note that the stage 1012 includes a write pointer (WP) for the output command buffer 1014. For an embodiment, the write pointer (WP) updates in a dispatch order. That is, for example, the write pointer (WP) updates when the thread T0 spawned commands are written, even if the thread T0 spawned commands are written after the T1 spawned commands are written.” ¶ [0066] states “That is, for an embodiment, the thread manager 1020 maintains the status of the threads through utilization of the scorecard 1022. The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated. One method of implementing this is for the compiler to insert (in the producer thread block of code) an instruction right after the instruction/s that produce the data for the consumer thread to increment a counter. The incremented counter in the scorecard is indicative of the dependency being satisfied.” See FIG. 10. Examiner’s Note: as shown by FIG. 10, the thread scheduler includes the output command buffer 1014. The thread scheduler also maintains the status of the producer threads. By using the scorecard, the thread scheduler detects the storing of data in the output command buffer);
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node (¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).” ¶ [0066] states “That is, for an embodiment, the thread manager 1020 maintains the status of the threads through utilization of the scorecard 1022. The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated. One method of implementing this is for the compiler to insert (in the producer thread block of code) an instruction right after the instruction/s that produce the data for the consumer thread to increment a counter. The incremented counter in the scorecard is indicative of the dependency being satisfied.” Examiner’s Note: when producer thread(s) finish execution, the data is generated and stored in a buffer. The hardware scheduler then checks, or makes a determination, about if there is sufficient data for the child thread).
With regard to claim 10, Teng and Agarwal teach the graph streaming processing system as claimed in claim 1. Teng additionally teaches wherein to determine to dispatch the at least one subsequent thread of the first node for execution, the thread scheduler is configured to: detect execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer (¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations” and “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” Examiner’s Note: the neural network accelerator signals the first processor when processing is complete. Completed processing indicates execution of the thread by consuming the data in the activation data buffer);
and perform one of: dispatch the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” ¶ [0058] states “the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612, and then signals (1) the first processor element that the shared memory queue is empty.” ¶ [0055] states “When the first processor element 602 receives a queue-empty signal (1) from the second processor element 604, the first processor element can proceed in staging (2) an input data set 610 to the RAM 608 for processing by the neural network accelerator 238.” Examiner’s Note: when the neural network accelerator stores output data in the shared memory queue, it has finished processing and the neural network accelerators’ activation buffer (B/C buffer) is available. The second processor element further processes the shared memory queue and signals the first processor element. The first processor element then repeats the flow by having subsequent threads dispatched. A thread of the first processor element is dispatched to load input data into the RAM 608 for the neural network accelerator);
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” ¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.” Examiner’s Note: the first processor element writes input data to the RAM, which means that the predefined threshold buffer size is not available. It then signals this determination to the neural network accelerator. As explained above, this thread is subsequent because the flows of steps 1-8.5 are repeated).
Agarwal additionally teaches wherein to determine to dispatch the at least one subsequent thread of the first node for execution, the thread scheduler is configured to: detect execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer (¶ [0043] states “The centralized dispatcher (thread manager 620) maintains the status of all running threads.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”);
evaluate the availability of the predefined threshold buffer size on the activation data buffer (¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.”);
and perform one of: dispatch the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available (¶ [0050] states “a first step 710 includes dispatching of threads of a cousin (producing) node. A second step 720 includes dispatching threads of a child node.”);
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available (¶ [0050] states “a first step 710 includes dispatching of threads of a cousin (producing) node. A second step 720 includes dispatching threads of a child node.”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the scheduling of threads based on availability of resources and the scheduling of multiple threads of a first or second node of Agarwal with the dispatching of subsequent threads based on buffer availability of Teng. As a result, subsequent threads of the first and second node are dispatched based on the flow of a producer and consumer of Teng. The first node of Agarwal is similar to the first processor element of Teng because they prepare/generate data for a dependent node. The second node of Agarwal is similar to the neural network accelerator of Teng because both consume data and are dependent on the first node/first processor element. The combination uses the availability of RAM for the neural network accelerator to determine whether to dispatch another thread for the first processing element or the neural network accelerator.
A person having ordinary skill in the art would have been motivated to make this combination “for improving the thread scheduling mechanisms during graph processing” (¶ [0037]). Teng also teaches that this allows additional data sets to be processed concurrently, which also is an improvement in thread scheduling mechanisms (¶ [0058] states “the first processor element can input another data set to the neural network accelerator for processing while the second processor element is performing the operations of the designated subset of layers of the neural network on the output data generated by the neural network accelerator for the previous input data set.”).
With regard to claim 11, Teng teaches a method comprising (¶ [0007] states “A disclosed method includes providing input data to a neural network accelerator by a first processor element of a host computer system.”):
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor (¶ [0055] states “FIG. 6 shows a neural network processing system, along with data flow and control signaling between a first processor element 602, a neural network accelerator 238, and a second processor element 604.” ¶ [0026] states “As used herein, a “processor element” can be a processor core of a computer system, heterogeneous processor circuits, or threads executing on one or more processor cores or processor circuits.” Examiner’s Note: RAM 608 and RAM 226 are the same RAM. B/C buffer 518 is the activation data buffer of the second processor because the data stored there is for neural network accelerator. The first processor executing indicates the thread executing was dispatched);
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” Examiner’s Note: the determination of when the first processor element has written the input dataset into the RAM for the neural network accelerator is when determining sufficiency takes place);
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” ¶ [0056] states “[0056] The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.” Examiner’s Note: signaling the neural network accelerator to start performing operations is dispatching the thread of the second processor);
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” ¶ [0058] states “In response to receiving the queue-full signal from the first processor element, the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612, and then signals (1) the first processor element that the shared memory queue is empty.” Examiner’s Note: when the neural network accelerator completes processing, its associated activation data buffer becomes available. The neural network accelerator then signals the processor element 602. After processor element 602 and 604 finish processing data, processor element 604 signals processor element 602 to then place more data in the RAM 608, or activation data buffer. Before placing more data in RAM 602, a determination takes place).
Teng does not explicitly teach a graph streaming processing system, a thread scheduler, and threads dependent on other threads based on a graph.
However, in an analogous art, Agarwal teaches dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor (¶ [0022] states “The described embodiments are embodied in methods, apparatuses and systems for accelerating graph stream processing.” ¶ [0023] states “a node includes one or more threads with each thread running the same code-block but (possibly) on different data and producing (possibly) different output data.” ¶ [0043] states “This mode of operation includes resolving dependencies of a child thread after dispatching the child thread. The centralized dispatcher (thread manager 620) maintains the status of all running threads.” ¶ [0050] states “As shown in the flow chart of FIG. 7, a first step 710 includes dispatching of threads of a cousin (producing) node.” Examiner’s Note: the cousin node is the first node);
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer (¶ [0039] states “The hardware scheduler tracks the status of the currently running threads in a scorecard.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.” ¶ [0024] states “A thread may be dependent on data generated by other threads of the same node, and/or data generated by threads of other nodes.” ¶ [0027] states “That is, the thread of a node can only be dependent on an older thread. For an embodiment, a thread can only be dependent on threads of an earlier stage, or threads of the same stage that have been dispatched earlier.” See FIG. 2. ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014 … Note that while the command buffer 1014 is the output command buffer for the stage 1012, the command buffer 1014 is the input command buffer for the stage 1015.” Examiner’s Note: child node 205 is dependent on the output data of cousin node 204. Child node 205 is analogous to the second node. Cousin node 204 is analogous to the first node. When the child node checks whether the Mth thread of the cousin node is completed, it is determining whether output data of the cousin node is sufficient);
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient (¶ [0050] states “A second step 720 includes dispatching threads of a child node” and “Upon completion of the Mth thread of the cousin node, a fifth step 750 includes the Nth thread of the child node proceeding with further execution after receiving the response to the dependency from the cousin node.” ¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).”);
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node (¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.” ¶ [0081] states “the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread.” ¶ [0047] states “The thread scheduler keeps on dispatching while the thread scheduler has the required resources in the processing cores.” Examiner’s Note: the required amount of memory in the output command buffer is the predefined threshold buffer size).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the graph streaming processing system and thread scheduler of Agarwal with the processor elements and neural network accelerator of Teng. As a result, the system handles dependency resolution between threads. A person having ordinary skill in the art would have been motivated to make this combination for the purpose of reducing execution time and increasing performance (¶ [0047] states “The thread scheduler keeps on dispatching while the thread scheduler has the required resources in the processing cores. This fills up the thread slots in the multi-threaded execution cores and allows each of the threads to determine execution based on their own dependencies. The execution time reduces considerably which results in higher performance”). See also ¶ [0048] – [0049] for additional improvements.
With regard to claim 13, Teng and Agarwal teach the method as claimed in claim 11. Teng additionally teaches detecting storing of the at least one data unit in the activation data buffer (¶ [0052] states “In processing the first per-layer instruction, the neural network accelerator 238 reads input data from a first portion of the B/C buffer 518 in the RAM 226 and writes output data to a second portion of the B/C buffer in the RAM.” ¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.”);
Agarwal additionally teaches wherein determining the sufficiency of the at least one data unit for execution of the at least one thread of the second node comprises: detecting an execution of the at least one thread of the first node (¶ [0039] states “a hardware scheduler (also referred to as a thread manager) is responsible for issuing threads for execution.” ¶ [0040] states “the execution of the child thread 1 is initiated or dispatched at a time 310 at which the sibling (identical twin or not) thread 1 and the cousin thread 1 have completed their processing.” Examiner’s Note: by issues a thread for execution, the scheduler has detected an execution of a thread. Cousin thread is a thread that belongs to a cousin node, which is the first node);
detecting generation of the at least one data unit from the execution of the at least one thread of the first node (¶ [0039] states “The hardware scheduler tracks the status of the currently running threads in a scorecard.” ¶ [0055] states “each thread includes a set of instructions operating on the plurality of thread processors 1030 and operating on a set of data and producing output data.” ¶ [0057] states “each stage 1012, 1015 includes the input command buffer parser 1013, 1016, wherein the command buffer parser 1013, 1016 generates the threads of the stage 1012, 1015 based upon commands of a command buffer 1011, 1014 located between the stage and the previous stage.” ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014. Note that the stage 1012 includes a write pointer (WP) for the output command buffer 1014. For an embodiment, the write pointer (WP) updates in a dispatch order. That is, for example, the write pointer (WP) updates when the thread T0 spawned commands are written, even if the thread T0 spawned commands are written after the T1 spawned commands are written.” ¶ [0066] states “The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated.” See FIG. 10. Examiner’s Note: the threads generate output data. Output data is the data unit. In FIG. 10, output data is stored in command buffer 1014. When the thread manager updates the status of thread and based on which code portion is being executed, the thread manager detects the generation of the data unit);
detecting storing of the at least one data unit in the activation data buffer (¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014. Note that the stage 1012 includes a write pointer (WP) for the output command buffer 1014. For an embodiment, the write pointer (WP) updates in a dispatch order. That is, for example, the write pointer (WP) updates when the thread T0 spawned commands are written, even if the thread T0 spawned commands are written after the T1 spawned commands are written.” ¶ [0066] states “That is, for an embodiment, the thread manager 1020 maintains the status of the threads through utilization of the scorecard 1022. The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated. One method of implementing this is for the compiler to insert (in the producer thread block of code) an instruction right after the instruction/s that produce the data for the consumer thread to increment a counter. The incremented counter in the scorecard is indicative of the dependency being satisfied.” See FIG. 10. Examiner’s Note: as shown by FIG. 10, the thread scheduler includes the output command buffer 1014. The thread scheduler also maintains the status of the producer threads. By using the scorecard, the thread scheduler detects the storing of data in the output command buffer);
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node (¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).” ¶ [0066] states “That is, for an embodiment, the thread manager 1020 maintains the status of the threads through utilization of the scorecard 1022. The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated. One method of implementing this is for the compiler to insert (in the producer thread block of code) an instruction right after the instruction/s that produce the data for the consumer thread to increment a counter. The incremented counter in the scorecard is indicative of the dependency being satisfied.” Examiner’s Note: when producer thread(s) finish execution, the data is generated and stored in a buffer. The hardware scheduler then checks, or makes a determination, about if there is sufficient data for the child thread).
With regard to claim 15, Teng and Agarwal teach the method as claimed in claim 11. Teng additionally teaches wherein determining to dispatch the at least one subsequent thread of the first node for execution further comprises: detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer (¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations” and “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” Examiner’s Note: the neural network accelerator signals the first processor when processing is complete. Completed processing indicates execution of the thread by consuming the data in the activation data buffer);
and performing one of: dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” ¶ [0058] states “the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612, and then signals (1) the first processor element that the shared memory queue is empty.” ¶ [0055] states “When the first processor element 602 receives a queue-empty signal (1) from the second processor element 604, the first processor element can proceed in staging (2) an input data set 610 to the RAM 608 for processing by the neural network accelerator 238.” Examiner’s Note: when the neural network accelerator stores output data in the shared memory queue, it has finished processing and the neural network accelerators’ activation buffer (B/C buffer) is available. The second processor element further processes the shared memory queue and signals the first processor element. The first processor element then repeats the flow by having subsequent threads dispatched. A thread of the first processor element is dispatched to load input data into the RAM 608 for the neural network accelerator);
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” ¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.” Examiner’s Note: the first processor element writes input data to the RAM, which means that the predefined threshold buffer size is not available. It then signals this determination to the neural network accelerator. As explained above, this thread is subsequent because the flows of steps 1-8.5 are repeated).
Agarwal additionally teaches wherein determining to dispatch the at least one subsequent thread of the first node for execution further comprises: detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer (¶ [0043] states “The centralized dispatcher (thread manager 620) maintains the status of all running threads.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”)
evaluating the availability of the predefined threshold buffer size on the activation data buffer (¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.”);
and performing one of: dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available (¶ [0050] states “a first step 710 includes dispatching of threads of a cousin (producing) node. A second step 720 includes dispatching threads of a child node.”);
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available (¶ [0050] states “a first step 710 includes dispatching of threads of a cousin (producing) node. A second step 720 includes dispatching threads of a child node.”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the scheduling of threads based on availability of resources and the scheduling of multiple threads of a first or second node of Agarwal with the dispatching of subsequent threads based on buffer availability of Teng. As a result, subsequent threads of the first and second node are dispatched based on the flow of a producer and consumer of Teng. The first node of Agarwal is similar to the first processor element of Teng because they prepare/generate data for a dependent node. The second node of Agarwal is similar to the neural network accelerator of Teng because both consume data and are dependent on the first node/first processor element. The combination uses the availability of RAM for the neural network accelerator to determine whether to dispatch another thread for the first processing element or the neural network accelerator.
A person having ordinary skill in the art would have been motivated to make this combination “for improving the thread scheduling mechanisms during graph processing” (¶ [0037]). Teng also teaches that this allows additional data sets to be processed concurrently, which also is an improvement in thread scheduling mechanisms (¶ [0058] states “the first processor element can input another data set to the neural network accelerator for processing while the second processor element is performing the operations of the designated subset of layers of the neural network on the output data generated by the neural network accelerator for the previous input data set.”).
With regard to claim 16, Teng teaches a non-transitory computer-readable medium having program instructions stored thereon, wherein the program instructions, when executed by a thread-scheduler of a graph streaming processing system, facilitate (¶ [0033] states “The system memory 216 is a device allowing information, such as executable instructions and data, to be stored and retrieved. The system memory 216 can include, for example, one or more random access memory (RAM) modules, such as double-data rate (DDR) dynamic RAM (DRAM). The storage device 218 includes local storage devices (e.g., one or more hard disks, flash memory modules, solid state disks, and optical disks) and/or a storage interface that enables the computing system 108 to communicate with one or more network data storage systems.” ¶ [0078] states “the processes may be provided via a variety of computer-readable storage media or delivery channels such as magnetic or optical disks or tapes, electronic storage devices, or as application services over a network.”):
dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor (¶ [0055] states “FIG. 6 shows a neural network processing system, along with data flow and control signaling between a first processor element 602, a neural network accelerator 238, and a second processor element 604.” ¶ [0026] states “As used herein, a “processor element” can be a processor core of a computer system, heterogeneous processor circuits, or threads executing on one or more processor cores or processor circuits.” Examiner’s Note: RAM 608 and RAM 226 are the same RAM. B/C buffer 518 is the activation data buffer of the second processor because the data stored there is for neural network accelerator. The first processor executing indicates the thread executing was dispatched);
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” Examiner’s Note: the determination of when the first processor element has written the input dataset into the RAM for the neural network accelerator is when determining sufficiency takes place);
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” ¶ [0056] states “[0056] The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.” Examiner’s Note: signaling the neural network accelerator to start performing operations is dispatching the thread of the second processor);
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” ¶ [0058] states “In response to receiving the queue-full signal from the first processor element, the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612, and then signals (1) the first processor element that the shared memory queue is empty.” Examiner’s Note: when the neural network accelerator completes processing, its associated activation data buffer becomes available. The neural network accelerator then signals the processor element 602. After processor element 602 and 604 finish processing data, processor element 604 signals processor element 602 to then place more data in the RAM 608, or activation data buffer. Before placing more data in RAM 602, a determination takes place).
Teng does not explicitly teach a graph streaming processing system, a thread scheduler, and threads dependent on other threads based on a graph.
However, in an analogous art, Agarwal teaches dispatching, by a thread scheduler of a graph streaming processing system, one thread associated with a first node, to one of a first processor array and a second processor of the graph streaming processing system, to generate an output data comprising at least one data unit, wherein the at least one data unit is stored in an activation data buffer of the second processor (¶ [0022] states “The described embodiments are embodied in methods, apparatuses and systems for accelerating graph stream processing.” ¶ [0023] states “a node includes one or more threads with each thread running the same code-block but (possibly) on different data and producing (possibly) different output data.” ¶ [0043] states “This mode of operation includes resolving dependencies of a child thread after dispatching the child thread. The centralized dispatcher (thread manager 620) maintains the status of all running threads.” ¶ [0050] states “As shown in the flow chart of FIG. 7, a first step 710 includes dispatching of threads of a cousin (producing) node.” Examiner’s Note: the cousin node is the first node);
determining, by the thread scheduler, sufficiency of at least one data unit for executing at least one thread of the second node, wherein the second node is identified to be dependent on the output data generated by execution of a plurality of threads of the first node, wherein the at least one data unit is sufficient when input data required for executing at least one thread of the second node is generated and available in the activation data buffer (¶ [0039] states “The hardware scheduler tracks the status of the currently running threads in a scorecard.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.” ¶ [0024] states “A thread may be dependent on data generated by other threads of the same node, and/or data generated by threads of other nodes.” ¶ [0027] states “That is, the thread of a node can only be dependent on an older thread. For an embodiment, a thread can only be dependent on threads of an earlier stage, or threads of the same stage that have been dispatched earlier.” See FIG. 2. ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014 … Note that while the command buffer 1014 is the output command buffer for the stage 1012, the command buffer 1014 is the input command buffer for the stage 1015.” Examiner’s Note: child node 205 is dependent on the output data of cousin node 204. Child node 205 is analogous to the second node. Cousin node 204 is analogous to the first node. When the child node checks whether the Mth thread of the cousin node is completed, it is determining whether output data of the cousin node is sufficient);
dispatching, by the thread scheduler, the at least one thread of the second node, to one of the first processor array and the second processor, upon determining that the at least one data unit is sufficient (¶ [0050] states “A second step 720 includes dispatching threads of a child node” and “Upon completion of the Mth thread of the cousin node, a fifth step 750 includes the Nth thread of the child node proceeding with further execution after receiving the response to the dependency from the cousin node.” ¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).”);
and determining, by the thread scheduler, to dispatch at least one subsequent thread of the first node for execution when a predefined threshold buffer size is available on the activation data buffer, wherein the predefined threshold buffer size indicates a minimum memory size required to store output data generated by executing at least one thread of the first node or the second node (¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.” ¶ [0081] states “the hardware management of the command buffer in each stage includes the forwarding of information required by every stage from the input command buffer to the output command buffer, allocation of the required amount of memory (for the output thread-spawn commands) in the output command buffer before scheduling a thread.” ¶ [0047] states “The thread scheduler keeps on dispatching while the thread scheduler has the required resources in the processing cores.” Examiner’s Note: the required amount of memory in the output command buffer is the predefined threshold buffer size).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the graph streaming processing system and thread scheduler of Agarwal with the processor elements and neural network accelerator of Teng. As a result, the system handles dependency resolution between threads. A person having ordinary skill in the art would have been motivated to make this combination for the purpose of reducing execution time and increasing performance (¶ [0047] states “The thread scheduler keeps on dispatching while the thread scheduler has the required resources in the processing cores. This fills up the thread slots in the multi-threaded execution cores and allows each of the threads to determine execution based on their own dependencies. The execution time reduces considerably which results in higher performance”). See also ¶ [0048] – [0049] for additional improvements.
With regard to claim 18, Teng and Agarwal teach the non-transitory computer-readable medium as claimed in claim 16. Teng additionally teaches detecting storing of the at least one data unit in the activation data buffer (¶ [0052] states “In processing the first per-layer instruction, the neural network accelerator 238 reads input data from a first portion of the B/C buffer 518 in the RAM 226 and writes output data to a second portion of the B/C buffer in the RAM.” ¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.”);
Agarwal additionally teaches wherein the program instructions configured to determine the sufficiency of the at least one data unit further facilitate: detecting an execution of the at least one thread of the first node (¶ [0039] states “a hardware scheduler (also referred to as a thread manager) is responsible for issuing threads for execution.” ¶ [0040] states “the execution of the child thread 1 is initiated or dispatched at a time 310 at which the sibling (identical twin or not) thread 1 and the cousin thread 1 have completed their processing.” Examiner’s Note: by issues a thread for execution, the scheduler has detected an execution of a thread. Cousin thread is a thread that belongs to a cousin node, which is the first node);
detecting generation of the at least one data unit from the execution of the at least one thread of the first node (¶ [0039] states “The hardware scheduler tracks the status of the currently running threads in a scorecard.” ¶ [0055] states “each thread includes a set of instructions operating on the plurality of thread processors 1030 and operating on a set of data and producing output data.” ¶ [0057] states “each stage 1012, 1015 includes the input command buffer parser 1013, 1016, wherein the command buffer parser 1013, 1016 generates the threads of the stage 1012, 1015 based upon commands of a command buffer 1011, 1014 located between the stage and the previous stage.” ¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014. Note that the stage 1012 includes a write pointer (WP) for the output command buffer 1014. For an embodiment, the write pointer (WP) updates in a dispatch order. That is, for example, the write pointer (WP) updates when the thread T0 spawned commands are written, even if the thread T0 spawned commands are written after the T1 spawned commands are written.” ¶ [0066] states “The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated.” See FIG. 10. Examiner’s Note: the threads generate output data. Output data is the data unit. In FIG. 10, output data is stored in command buffer 1014. When the thread manager updates the status of thread and based on which code portion is being executed, the thread manager detects the generation of the data unit);
detecting storing of the at least one data unit in the activation data buffer (¶ [0058] states “During processing of the threads T0, T1, write command(s) are spawned which are written into the output command buffer 1014. Note that the stage 1012 includes a write pointer (WP) for the output command buffer 1014. For an embodiment, the write pointer (WP) updates in a dispatch order. That is, for example, the write pointer (WP) updates when the thread T0 spawned commands are written, even if the thread T0 spawned commands are written after the T1 spawned commands are written.” ¶ [0066] states “That is, for an embodiment, the thread manager 1020 maintains the status of the threads through utilization of the scorecard 1022. The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated. One method of implementing this is for the compiler to insert (in the producer thread block of code) an instruction right after the instruction/s that produce the data for the consumer thread to increment a counter. The incremented counter in the scorecard is indicative of the dependency being satisfied.” See FIG. 10. Examiner’s Note: as shown by FIG. 10, the thread scheduler includes the output command buffer 1014. The thread scheduler also maintains the status of the producer threads. By using the scorecard, the thread scheduler detects the storing of data in the output command buffer);
and determining that the at least one data unit comprises sufficient data for execution of the at least one thread of the second node (¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).” ¶ [0066] states “That is, for an embodiment, the thread manager 1020 maintains the status of the threads through utilization of the scorecard 1022. The producer thread updates his status (where in the code is the producer thread currently done with execution) back to the thread manager and the scorecard is updated. One method of implementing this is for the compiler to insert (in the producer thread block of code) an instruction right after the instruction/s that produce the data for the consumer thread to increment a counter. The incremented counter in the scorecard is indicative of the dependency being satisfied.” Examiner’s Note: when producer thread(s) finish execution, the data is generated and stored in a buffer. The hardware scheduler then checks, or makes a determination, about if there is sufficient data for the child thread).
With regard to claim 20, Teng and Agarwal teach the non-transitory computer-readable medium as claimed in claim 16. Teng additionally teaches wherein the program instructions configured to determine to dispatch the at least one subsequent thread of the first node further facilitate: detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer (¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations” and “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” Examiner’s Note: the neural network accelerator signals the first processor when processing is complete. Completed processing indicates execution of the thread by consuming the data in the activation data buffer);
and performing one of: dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602.” ¶ [0058] states “the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612, and then signals (1) the first processor element that the shared memory queue is empty.” ¶ [0055] states “When the first processor element 602 receives a queue-empty signal (1) from the second processor element 604, the first processor element can proceed in staging (2) an input data set 610 to the RAM 608 for processing by the neural network accelerator 238.” Examiner’s Note: when the neural network accelerator stores output data in the shared memory queue, it has finished processing and the neural network accelerators’ activation buffer (B/C buffer) is available. The second processor element further processes the shared memory queue and signals the first processor element. The first processor element then repeats the flow by having subsequent threads dispatched. A thread of the first processor element is dispatched to load input data into the RAM 608 for the neural network accelerator);
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.” ¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.” Examiner’s Note: the first processor element writes input data to the RAM, which means that the predefined threshold buffer size is not available. It then signals this determination to the neural network accelerator. As explained above, this thread is subsequent because the flows of steps 1-8.5 are repeated).
Agarwal additionally teaches wherein the program instructions configured to determine to dispatch the at least one subsequent thread of the first node further facilitate: detecting execution of the at least one thread of the second node by consuming the at least one data unit stored in the activation data buffer (¶ [0043] states “The centralized dispatcher (thread manager 620) maintains the status of all running threads.” ¶ [0050] states “A third step 730 includes an Nth thread of the child node checking whether an Mth thread of the cousin node is completed.”)
evaluating the availability of the predefined threshold buffer size on the activation data buffer (¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.”);
and performing one of: dispatching the at least one subsequent thread of the first node, upon determining that the predefined threshold buffer size is available (¶ [0050] states “a first step 710 includes dispatching of threads of a cousin (producing) node. A second step 720 includes dispatching threads of a child node.”);
and dispatching at least one subsequent thread of the second node upon determining that the predefined threshold buffer size is not available (¶ [0050] states “a first step 710 includes dispatching of threads of a cousin (producing) node. A second step 720 includes dispatching threads of a child node.”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the scheduling of threads based on availability of resources and the scheduling of multiple threads of a first or second node of Agarwal with the dispatching of subsequent threads based on buffer availability of Teng. As a result, subsequent threads of the first and second node are dispatched based on the flow of a producer and consumer of Teng. The first node of Agarwal is similar to the first processor element of Teng because they prepare/generate data for a dependent node. The second node of Agarwal is similar to the neural network accelerator of Teng because both consume data and are dependent on the first node/first processor element. The combination uses the availability of RAM for the neural network accelerator to determine whether to dispatch another thread for the first processing element or the neural network accelerator.
A person having ordinary skill in the art would have been motivated to make this combination “for improving the thread scheduling mechanisms during graph processing” (¶ [0037]). Teng also teaches that this allows additional data sets to be processed concurrently, which also is an improvement in thread scheduling mechanisms (¶ [0058] states “the first processor element can input another data set to the neural network accelerator for processing while the second processor element is performing the operations of the designated subset of layers of the neural network on the output data generated by the neural network accelerator for the previous input data set.”).
Claim(s) 3, 12, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Agarwal and further in view of Sankaran et al. Pat. No. US 20200401440 A1 (hereafter Sankaran).
With regard to claim 3, Teng and Agarwal teach the graph streaming processing system as claimed in claim 1. Teng and Agarwal do not explicitly teach a process of determining which processor to dispatch a thread to.
However, in an analogous art, Sankaran teaches wherein the thread scheduler dispatches the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor (¶ [0197] states “a software pattern matcher 11417 may be used that identifies “hot” code, accelerator code, and/or code that requests a PE allocation. For example, the software pattern matcher 11417 compares the code sequence to a predetermined set of patterns stored in memory.” ¶ [0198] states “These components feed a selector 11411 which makes a selection of a PE to execute a thread, what link protocol to use, what migration should occur if there is a thread already executing on that PE, etc.” ¶ [0153] states “Embodiments of systems, methods, and apparatuses for selecting a processing element (e.g., a core or an accelerator) to process a thread.” Examiner’s Note: the thread scheduler uses the pattern matcher to determine which processing element (PE) will execute the thread. The pattern matcher uses a predetermined set of patterns which is the predefined processing information).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the pattern matcher for selecting which processing element to execute a thread with the processing element and neural network accelerator of Teng and the graph streaming processing system of Agarwal. A person having ordinary skill in the art would have been motivated to make this combination to “enable programmers to develop software targeted for a specific architecture, or an equivalent abstraction, while facilitating continuous improvements to the underlying hardware without requiring corresponding changes to the developed software” (¶ [0154]). Further, by using the pattern matcher, the most suitable processing element is used to execute the thread. This provides a performance benefit (¶ [0159] states “when an accelerator is available, and it is desirable to use the accelerator (use less power, for example)”).
With regard to claim 12, Teng and Agarwal teach the method as claimed in claim 11. Teng and Agarwal do not explicitly teach a process of determining which processor to dispatch a thread to.
However, in an analogous art, Sankaran teaches wherein dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor (¶ [0197] states “a software pattern matcher 11417 may be used that identifies “hot” code, accelerator code, and/or code that requests a PE allocation. For example, the software pattern matcher 11417 compares the code sequence to a predetermined set of patterns stored in memory.” ¶ [0198] states “These components feed a selector 11411 which makes a selection of a PE to execute a thread, what link protocol to use, what migration should occur if there is a thread already executing on that PE, etc.” ¶ [0153] states “Embodiments of systems, methods, and apparatuses for selecting a processing element (e.g., a core or an accelerator) to process a thread.” Examiner’s Note: the thread scheduler uses the pattern matcher to determine which processing element (PE) will execute the thread. The pattern matcher uses a predetermined set of patterns which is the predefined processing information).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the pattern matcher for selecting which processing element to execute a thread with the processing element and neural network accelerator of Teng and the graph streaming processing system of Agarwal. A person having ordinary skill in the art would have been motivated to make this combination to “enable programmers to develop software targeted for a specific architecture, or an equivalent abstraction, while facilitating continuous improvements to the underlying hardware without requiring corresponding changes to the developed software” (¶ [0154]). Further, by using the pattern matcher, the most suitable processing element is used to execute the thread. This provides a performance benefit (¶ [0159] states “when an accelerator is available, and it is desirable to use the accelerator (use less power, for example)”).
With regard to claim 17, Teng and Agarwal teach the non-transitory computer-readable medium as claimed in claim 16. Teng and Agarwal do not explicitly teach a process of determining which processor to dispatch a thread to.
However, in an analogous art, Sankaran teaches wherein the program instructions further facilitate: dispatching the at least one thread to one of the first processor array and the second processor based on a type of processing required for execution of the at least one thread, wherein the type of processing is determined based on predefined processing information associated with the first processor array and the second processor (¶ [0197] states “a software pattern matcher 11417 may be used that identifies “hot” code, accelerator code, and/or code that requests a PE allocation. For example, the software pattern matcher 11417 compares the code sequence to a predetermined set of patterns stored in memory.” ¶ [0198] states “These components feed a selector 11411 which makes a selection of a PE to execute a thread, what link protocol to use, what migration should occur if there is a thread already executing on that PE, etc.” ¶ [0153] states “Embodiments of systems, methods, and apparatuses for selecting a processing element (e.g., a core or an accelerator) to process a thread.” Examiner’s Note: the thread scheduler uses the pattern matcher to determine which processing element (PE) will execute the thread. The pattern matcher uses a predetermined set of patterns which is the predefined processing information).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the pattern matcher for selecting which processing element to execute a thread with the processing element and neural network accelerator of Teng and the graph streaming processing system of Agarwal. A person having ordinary skill in the art would have been motivated to make this combination to “enable programmers to develop software targeted for a specific architecture, or an equivalent abstraction, while facilitating continuous improvements to the underlying hardware without requiring corresponding changes to the developed software” (¶ [0154]). Further, by using the pattern matcher, the most suitable processing element is used to execute the thread. This provides a performance benefit (¶ [0159] states “when an accelerator is available, and it is desirable to use the accelerator (use less power, for example)”).
Claim(s) 4 and 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Agarwal and further in view of Bass et al. Pat. No. US 20130304990 A1 (hereafter Bass).
With regard to claim 4, Teng and Agarwal teach the graph streaming processing system as claimed in claim 1. Teng additionally teaches wherein the first processor array is configured to: receive the at least one thread of the second node dispatched by the thread scheduler (¶ [0050] states “A second step 720 includes dispatching threads of a child node.” ¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.”)
retrieve input data required for execution of the at least one thread of the second node from a shared data buffer shared between the first processor array and the second processor (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612 and when processing is complete signals (6) completion to the first processor element 602. The output data from the neural network accelerator can be viewed as an intermediate data set in implementations in which the second processor element 604 further processes the output data.” ¶ [0058] states “If the user configured the first processor element 602 to perform data conversion, the first processor element converts (6.5) the output data in the shared memory queue 614 and then signals (7) the second processor element 604 that the queue is full. In response to receiving the queue-full signal from the first processor element, the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612.” See FIG. 6. Examiner’s Note: the first processing element 602 retrieves input data from the shared memory queue 614 in RAM 612. Shared memory queue is accessible to the first and second processing elements and the neural network accelerator, so the shared memory queue is the shared data buffer);
execute the at least one thread of the second node to generate the output data (¶ [0058] states “If the user configured the first processor element 602 to perform data conversion, the first processor element converts (6.5) the output data in the shared memory queue 614.” Examiner’s Note: the conversion process is the execution of the at least one thread);
and perform at least one of: writing the output data to the shared data buffer, upon determination that the at least one subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array, wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor (¶ [0058] states “If the user configured the first processor element 602 to perform data conversion, the first processor element converts (6.5) the output data in the shared memory queue 614 and then signals (7) the second processor element 604 that the queue is full. In response to receiving the queue-full signal from the first processor element, the second processor element copies (8) the contents of the shared memory queue to another workspace in the RAM 612.” Examiner’s Note: the converted data is stored in the shared data buffer. Therefore, it is the first processor performing writing to the shared data buffer. The subsequent thread is processing element 604 that will copy the contents of the shared memory queue to another workspace);
and writing the output data into the activation data buffer, upon determination that the at least one subsequent thread is being dispatched to the second processor (¶ [0055] states "the first processor element can proceed in staging (2) an input data set 610 to the RAM 608 for processing by the neural network accelerator 238.” Examiner’s Note: the first processing element 602, which is part of the first processor array, writes to the RAM 608, or activation data buffer).
Agarwal additionally teaches wherein the first processor array is configured to: receive the at least one thread of the second node dispatched by the thread scheduler (¶ [0050] states “A second step 720 includes dispatching threads of a child node.” ¶ [0043] states “This mode of operation includes resolving dependencies of a child thread after dispatching the child thread.”);
execute the at least one thread of the second node to generate the output data (¶ [0050] states “Upon completion of the Mth thread of the cousin node, a fifth step 750 includes the Nth thread of the child node proceeding with further execution after receiving the response to the dependency from the cousin node.” ¶ [0005] states “each thread includes a set of instructions operating on the plurality of thread processors and operating on a set of data and producing output data”);
Although Teng teaches writing data into either the shared memory queue or the RAM for the neural network accelerator, Teng and Agarwal do not explicitly teach writing data into a specific buffer based on which processor is to subsequently execute a thread.
However, in an analogous art, Bass teaches and perform at least one of: writing the output data to the shared data buffer, upon determination that the at least one subsequent thread, dependent on the at least one thread, is being dispatched to the first processor array, wherein the shared data buffer is a memory subsystem shared by the first processor array and the second processor (¶ [0024] states “Based on its write configurations, the PBI bridge controller 103 either rejects the request to Cache Inject and the write data from the coprocessor is written to main memory via DMA transfer, or, if the request is granted, the write data is written to the local cache of the core processor on the primary bus expected to next use the write data.” Examiner’s Note: one of the processor cores represents the first processor array);
and writing the output data into the activation data buffer, upon determination that the at least one subsequent thread is being dispatched to the second processor (¶ [0024] states “Based on its write configurations, the PBI bridge controller 103 either rejects the request to Cache Inject and the write data from the coprocessor is written to main memory via DMA transfer, or, if the request is granted, the write data is written to the local cache of the core processor on the primary bus expected to next use the write data.” Examiner’s Note: one of the processor cores represents the second processor).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine writing data into the cache of a core processor that is expected to use the data next of Bass with the RAM for neural network accelerator and shared memory queue of Teng and the threads of Agarwal. As a result, the thread scheduler of Agarwal writes data into either the shared memory queue or the activation data buffer of Teng (See FIG. 6 of Teng. RAM 612 stores the shared memory queue and RAM 608 is the activation data buffer. RAM 608 has the B/C buffer) based on Bass’ teaching of writing data into the cache of the processor that will next use the data.
A person having ordinary skill in the art would have been motivated to make this combination for the purpose of efficiently transferring data from one processor to another processor (¶ [0009] states “Cache injection is a technique in which data is transferred into a cache during a DMA transfer into system memory, thus reducing or eliminating the delay associated with subsequently loading the data into cache for use by the processor.”). See also ¶ [0010] – [0012] for additional improvements.
With regard to claim 6, Teng, Agarwal, and Bass teach the graph streaming processing system as claimed in claim 4. Teng additionally teaches wherein the second processor is configured to: receive the at least one thread of the second node dispatched by the thread scheduler (¶ [0055] states “Once the first processor element has written the input data set to the RAM, the first processor element signals (3) to the neural network accelerator to commence performing the specified neural network operations on the input data set.”);
retrieve input data required for execution of the at least one thread of the second node from at least one of the shared data buffer and the activation data buffer (¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.”);
execute the at least one thread to generate the output data (¶ [0056] states “The neural network accelerator reads (4) the input data set 610 from the RAM 608 and performs the specified subset of neural network operations.” See FIG. 6 step (5) process and output);
and perform at least one of: writing the output data into the activation data buffer, upon determination that the subsequent thread is being dispatched to the second processor (¶ [0040] states “In the example, the neural network accelerator 238 uses interfaces 330, 331, and 332 to communicate with the interconnect 306” and “The read/write interface 332 can be used to read and write from the RAM 226 through a second one of the memory interfaces 312.”);
and writing the output data into the shared data buffer, upon determination that the at least one subsequent thread is being dispatched to the first processor array (¶ [0056] states “The neural network accelerator stores the output data in the shared memory queue 614 in the RAM 612”).
Agarwal additionally teaches wherein the second processor is configured to: receive the at least one thread of the second node dispatched by the thread scheduler (¶ [0050] states “A second step 720 includes dispatching threads of a child node.” ¶ [0043] states “This mode of operation includes resolving dependencies of a child thread after dispatching the child thread.”);
retrieve input data required for execution of the at least one thread of the second node from at least one of the shared data buffer and the activation data buffer (¶ [0057] states “For an embodiment, each stage 1012, 1015 includes the input command buffer parser 1013, 1016, wherein the command buffer parser 1013, 1016 generates the threads of the stage 1012, 1015 based upon commands of a command buffer 1011, 1014 located between the stage and the previous stage.” Examiner’s Note: input data for generating threads is retrieved from input command buffer 1011 and 1014. Input command buffer is a shared data buffer);
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the reading of input data from a shared data buffer of Agarwal with the neural network accelerator of Teng. As a result, the neural network accelerator reads input data from either the B/C buffer or a shared data buffer. A person having ordinary skill in the art would have been motivated to make this combination so that the neural network accelerator processes data that has been processed by either the first or second processing element (Teng ¶ [0058] describes a conversion process). Additionally, the neural network accelerator would be able to process data without needing the first processor element to load data into B/C buffer which provides a flexibility benefit. Neural network accelerators provide a performance benefit to some layers of a CNN (Teng ¶ [0025] states “In applications such as CNNs, the inventors have found that a performance benefit can be realized by implementing some layers of the CNN on a neural network accelerator, and implementing others of the layers on the host”).
Bass additionally teaches and perform at least one of: writing the output data into the activation data buffer, upon determination that the subsequent thread is being dispatched to the second processor (¶ [0024] states “Based on its write configurations, the PBI bridge controller 103 either rejects the request to Cache Inject and the write data from the coprocessor is written to main memory via DMA transfer, or, if the request is granted, the write data is written to the local cache of the core processor on the primary bus expected to next use the write data.” Examiner’s Note: one of the processor cores represents the first processor array);
and writing the output data into the shared data buffer, upon determination that the at least one subsequent thread is being dispatched to the first processor array (¶ [0024] states “Based on its write configurations, the PBI bridge controller 103 either rejects the request to Cache Inject and the write data from the coprocessor is written to main memory via DMA transfer, or, if the request is granted, the write data is written to the local cache of the core processor on the primary bus expected to next use the write data.” Examiner’s Note: one of the processor cores represents the second processor).
Claim(s) 9, 14, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Agarwal and further in view of Hux et al. Pat. No. US 20190102859 A1 (hereafter Hux).
With regard to claim 9, Teng and Agarwal teach the graph streaming processing system as claimed in claim 8. Agarwal additionally teaches wherein the thread scheduler is further configured to: dispatch at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node (¶ [0024] states “Each of the nodes 101-113 may be processed in parallel with multiple threads, wherein each thread may or may not be dependent on the processing of one or more other threads” and “A thread may be dependent on data generated by other threads of the same node, and/or data generated by threads of other nodes.” ¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).” ¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.” Examiner’s Note: multiple threads of a same node are scheduled. In one embodiment, threads of a parent node (uncle/cousin/sibling) are scheduled before a child node. When the hardware scheduler checks if the producer nodes are done processing and outputting data for the child node, that is the determination of sufficient data).
Teng and Agarwal do not explicitly teach dispatching subsequent threads in response to a determination.
However, in an analogous art, Hux teaches wherein the thread scheduler is further configured to: dispatch at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node (¶ [0197] states “Once the preempt signal is set to low, one or more threads or thread groups are dispatched for the selected command at 2834. At 2836, it can be determined if more threads are available or needed for the current commend. If so, processing returns to 2830, if not processing returns to 2802 and the loop continues. Returning to 2830, if the watermark has not been met, then the preempt signal is set to high at 2832 and one or more threads or thread groups are dispatched for the selected command.”)
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the dispatching of additional threads after a determination if more threads are needed of Hux with the dispatching of subsequent threads of a same node the checking of completion of producer nodes of Agarwal and the system of processor elements and neural network accelerator of Teng. As a result, threads of a first node of Agarwal are dispatched. During the determination if a child node can begin execution, a determination is made as to whether additional threads should be dispatched. Accordingly, additional threads are dispatched before the child node thread is dispatched.
A person having ordinary skill in the art would have been motivated to make this combination to realize the benefits of the global thread dispatcher logic such as passing commands to a high priority streamer and a normal priority command streamer. This embodiment relates to reducing latency in a GPU which can provide an improved experience to users of VR, AR, and MR (¶ [0180] states “latency may be reduced by providing a direct path to the GPU, bypassing the traditional command submission path. Such a direct path may still provide an interactive user interface with minimal power and die area requirements.” ¶ [0178] states “Various applications, such as virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications depend on low-latency responses to head motion. Any disparity between head motion and sensory input (e.g., a head mounted display or other visual system) has been shown to trigger a biological response that induces nausea in the viewer.”)
With regard to claim 14, Teng and Agarwal teach the method as claimed in claim 11. Agarwal additionally teaches further comprising: dispatching the at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node (¶ [0024] states “Each of the nodes 101-113 may be processed in parallel with multiple threads, wherein each thread may or may not be dependent on the processing of one or more other threads” and “A thread may be dependent on data generated by other threads of the same node, and/or data generated by threads of other nodes.” ¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).” ¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.” Examiner’s Note: multiple threads of a same node are scheduled. In one embodiment, threads of a parent node (uncle/cousin/sibling) are scheduled before a child node. When the hardware scheduler checks if the producer nodes are done processing and outputting data for the child node, that is the determination of sufficient data).
Teng and Agarwal do not explicitly teach dispatching subsequent threads in response to a determination.
However, in an analogous art, Hux teaches further comprising: dispatching the at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node (¶ [0197] states “Once the preempt signal is set to low, one or more threads or thread groups are dispatched for the selected command at 2834. At 2836, it can be determined if more threads are available or needed for the current commend. If so, processing returns to 2830, if not processing returns to 2802 and the loop continues. Returning to 2830, if the watermark has not been met, then the preempt signal is set to high at 2832 and one or more threads or thread groups are dispatched for the selected command.”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the dispatching of additional threads after a determination if more threads are needed of Hux with the dispatching of subsequent threads of a same node the checking of completion of producer nodes of Agarwal and the system of processor elements and neural network accelerator of Teng. As a result, threads of a first node of Agarwal are dispatched. During the determination if a child node can begin execution, a determination is made as to whether additional threads should be dispatched. Accordingly, additional threads are dispatched before the child node thread is dispatched.
A person having ordinary skill in the art would have been motivated to make this combination to realize the benefits of the global thread dispatcher logic such as passing commands to a high priority streamer and a normal priority command streamer. This embodiment relates to reducing latency in a GPU which can provide an improved experience to users of VR, AR, and MR (¶ [0180] states “latency may be reduced by providing a direct path to the GPU, bypassing the traditional command submission path. Such a direct path may still provide an interactive user interface with minimal power and die area requirements.” ¶ [0178] states “Various applications, such as virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications depend on low-latency responses to head motion. Any disparity between head motion and sensory input (e.g., a head mounted display or other visual system) has been shown to trigger a biological response that induces nausea in the viewer.”).
With regard to claim 19, Teng and Agarwal teach the non-transitory computer-readable medium as claimed in claim 16. Agarwal additionally teaches wherein the program instructions further facilitate: dispatching the at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node (¶ [0024] states “Each of the nodes 101-113 may be processed in parallel with multiple threads, wherein each thread may or may not be dependent on the processing of one or more other threads” and “A thread may be dependent on data generated by other threads of the same node, and/or data generated by threads of other nodes.” ¶ [0039] states “Before the dispatch of a child thread, the hardware scheduler checks the status of the producer threads (uncle/cousin/sibling) in the scorecard. Once the producer thread(s) finish, the child thread is launched for execution (dispatched).” ¶ [0058] states “scheduling of a thread on the thread processors 1030 is based on availability of resources including a thread slot on a thread processor of the plurality of thread processors 1030, adequate space in the register file, space in the output command buffer for writing the commands produced by the spawn instructions.” Examiner’s Note: multiple threads of a same node are scheduled. In one embodiment, threads of a parent node (uncle/cousin/sibling) are scheduled before a child node. When the hardware scheduler checks if the producer nodes are done processing and outputting data for the child node, that is the determination of sufficient data).
Teng and Agarwal do not explicitly teach dispatching subsequent threads in response to a determination.
However, in an analogous art, Hux teaches wherein the program instructions further facilitate: dispatching the at least one subsequent thread of the first node to generate at least one subsequent data unit, before dispatching the at least one thread of the second node for execution, upon determining that the at least one data unit comprises insufficient data for execution of the at least one thread of the second node (¶ [0197] states “Once the preempt signal is set to low, one or more threads or thread groups are dispatched for the selected command at 2834. At 2836, it can be determined if more threads are available or needed for the current commend. If so, processing returns to 2830, if not processing returns to 2802 and the loop continues. Returning to 2830, if the watermark has not been met, then the preempt signal is set to high at 2832 and one or more threads or thread groups are dispatched for the selected command.”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the dispatching of additional threads after a determination if more threads are needed of Hux with the dispatching of subsequent threads of a same node the checking of completion of producer nodes of Agarwal and the system of processor elements and neural network accelerator of Teng. As a result, threads of a first node of Agarwal are dispatched. During the determination if a child node can begin execution, a determination is made as to whether additional threads should be dispatched. Accordingly, additional threads are dispatched before the child node thread is dispatched.
A person having ordinary skill in the art would have been motivated to make this combination to realize the benefits of the global thread dispatcher logic such as passing commands to a high priority streamer and a normal priority command streamer. This embodiment relates to reducing latency in a GPU which can provide an improved experience to users of VR, AR, and MR (¶ [0180] states “latency may be reduced by providing a direct path to the GPU, bypassing the traditional command submission path. Such a direct path may still provide an interactive user interface with minimal power and die area requirements.” ¶ [0178] states “Various applications, such as virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications depend on low-latency responses to head motion. Any disparity between head motion and sensory input (e.g., a head mounted display or other visual system) has been shown to trigger a biological response that induces nausea in the viewer.”).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20250094221 A1
teaches
THREAD SCHEDULING FOR MULTITHREADED DATA PROCESSING ENVIRONMENTS
US 20190235917 A1
teaches
CONFIGURABLE SCHEDULER IN A GRAPH STREAMING PROCESSING SYSTEM
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PETER L YUAN whose telephone number is (571)272-5737. The examiner can normally be reached Mon-Fri 7:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at 571-272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PETER LI YUAN/Examiner, Art Unit 2197
/BRADLEY A TEETS/Supervisory Patent Examiner, Art Unit 2197