Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mathematical concept/mental process) without significantly more.
Claim 1:
Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A system for updating transformer machine learning models to account for relative timing of events,”, and a system or machine is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process and mathematical concept but for recitation of generic computer components:
“generating a plurality of event embeddings comprising a query event embedding corresponding to the query event and a plurality of key event embeddings corresponding to the plurality of key events;” (this is a mental process, a person could mentally evaluate generating a plurality of event vector representations that comprise a query event vector representation to a query event and a plurality of key event vector representations corresponding to a plurality of key events using pen and paper, see MPEP § 2106.04(a)(2)(III)),
“determining a plurality of respective time differences between the first time of the query event and each corresponding second time of the plurality of key events;” (this is a mental process, a person could mentally determine a plurality of respective time differences between a first time of a query event and each corresponding second time of the plurality of key events, see MPEP § 2106.04(a)(2)(III)),
“generating a plurality of attention values…” (this is a mental process, a person could mentally evaluate generating a plurality of attention values, see MPEP § 2106.04(a)(2)(III)),
“…by aggregating, with the plurality of dot products, a function of the plurality of respective time differences, each attention value indicating a weight of a corresponding key event of the plurality of key events relative to the query event, and each attention value accounting for a respective time difference between the first time and a corresponding second time, wherein the function comprises an exponential decay function;” (this is a mathematical concept, generating a plurality of attention values by aggregation with a plurality of dot products is a mathematical calculation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process and mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
“A system for updating transformer machine learning models to account for relative timing of events, the system comprising: at least one processor,” (A system comprising a processor is considered generic computer component being used as tool to perform functions of the judicial exception – see MPEP § 2106.05(f)),
“at least one memory,” (Using a memory is considered generic computer component being used as tool to perform functions of the judicial exception – see MPEP § 2106.05(f)),
“and computer-readable media having computer-executable instructions stored thereon, the computer-executable instructions, when executed by the at least one processor, causing the system to perform operations comprising:” (Using microprocessor readable and executable instructions is considered mere instructions to apply an exception using generic computer – see MPEP § 2106.05(f)),
“retrieving a transformer model trained based on a plurality of events, the plurality of events comprising a query event associated with a first time and a plurality of key events associated with a plurality of second times;” (Retrivieving a model trained based on a plurality events is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)),
“inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings, the transformation comprising a plurality of dot products of the query event embedding with each of the plurality of key event embeddings;” (Inputting a plurality of event embeddings into a transformer model is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)),
“and updating the transformer model with the plurality of attention values.” (Updating a transformer model with a plurality of attention values is considered mere instructions to apply an exception using generic computer – see MPEP § 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements v and vi recites generic computer component being used as tool to perform functions of the judicial exception, additional elements vii and x recites mere instructions to apply an exception using generic computer, and additional elements viii and ix recites insignificant extra-solution activity of mere data gathering, which are well understood routines and conventional activity, see receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362, which is not indicative of significantly more.
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 2:
Regarding claim 2, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A method comprising:”, and a method or process is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process and mathematical concept but for recitation of generic computer components:
“generating a plurality of event embeddings for the plurality of events;” (this is a mental process, a person could mentally evaluate generating a plurality of event vector representations corresponding to a plurality of key events using pen and paper, see MPEP § 2106.04(a)(2)(III)),
“determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of key events;” (this is a mental process, a person could mentally determine a plurality of respective time differences between a first time of an event and each corresponding second time of the plurality of key events, see MPEP § 2106.04(a)(2)(III)),
“generating a plurality of attention values by aggregating the transformation and the plurality of respective time differences;” (this is a mental process, a person could mentally evaluate generating a plurality of attention values, see MPEP § 2106.04(a)(2)(III)),
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process and mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
“A method comprising: retrieving a transformer model trained based on a plurality of events, the plurality of events comprising a first event associated with a first time and a plurality of second events associated with a plurality of second times;” (Retrivieving a model trained based on a plurality events is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)),
“inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings;” (Inputting a plurality of event embeddings into a transformer model is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)),
“and updating the transformer model with the plurality of attention values.” (Updating a transformer model with a plurality of attention values is considered mere instructions to apply an exception using generic computer – see MPEP § 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional element vi recites mere instructions to apply an exception using generic computer, and additional elements iv and v recites insignificant extra-solution activity of mere data gathering, which are well understood routines and conventional activity, see receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362, which is not indicative of significantly more.
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 3:
Regarding claim 3, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis to claim 2. Further, claim 3 recites the following additional elements:
“The method of claim 2, wherein aggregating the transformation and the plurality of respective time differences comprises adjusting the transformation based on the plurality of respective time differences such that each attention value accounts for a respective time difference between the first time and a corresponding second time.” (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer, see MPEP § 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer - see MPEP § 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 4:
Regarding claim 4, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis to claim 2. Further, claim 4 recites the following additional elements:
“The method of claim 2, wherein aggregating the transformation and the plurality of respective time differences comprises aggregating the transformation and a function of the plurality of respective time differences.” (this is a mathematical concept, aggregating a transformation and the plurality of respective time differences comprising of aggregating the transformation and a function of the plurality of respective time differences is a mathematical function or formula, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 5:
Regarding claim 5, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis to claim 4. Further, claim 5 recites the following additional elements:
“The method of claim 4, wherein the function is an exponential decay function and wherein aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation.” (this is a mathematical concept, a function being an exponential decay function and the aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation is a mathematical function or formula and calculation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 6:
Regarding claim 6, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis to claim 4. Further, claim 6 recites the following additional elements:
“The method of claim 4, wherein the function is an exponential growth function and wherein aggregating the transformation and the plurality of respective time differences comprises subtracting the function of the plurality of respective time differences from the transformation.” (this is a mathematical concept, a function being an exponential growth function and the aggregating the transformation and the plurality of respective time differences comprises subtracting the function of the plurality of respective time differences to the transformation is a mathematical function or formula and calculation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 7:
Regarding claim 7, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis to claim 2. Further, claim 7 recites the following additional elements:
“The method of claim 2, wherein the transformation comprises a plurality of dot products of a first event embedding corresponding to the first event with each of a plurality of second event embeddings corresponding to the plurality of second events.” (this is a mathematical concept, a transformation comprising of dot products is a mathematical calculation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 8:
Regarding claim 8, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis to claim 2. Further, claim 8 recites the following additional elements:
“The method of claim 2, wherein generating the plurality of attention values further comprises normalizing, using a softmax function, values generated by aggregating the transformation and the plurality of respective time differences.” (this is a mathematical concept, normalizing using a softmax function is a mathematical operation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 9:
Regarding claim 9, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis to claim 2. Further, claim 9 recites the following additional elements:
“The method of claim 2, wherein updating the transformer model with the plurality of attention values comprises aggregating the plurality of respective time differences with a plurality of transformations at a plurality of layers of the transformer model.” (this is a mathematical concept, aggregating a plurality of respective time differences with a plurality of transformations is a mathematical operation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 10:
Regarding claim 10, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis to claim 2. Further, claim 10 recites the following additional elements:
“The method of claim 2, wherein the first event is a query event and wherein the plurality of second events is a plurality of key events, wherein each key event is weighted relative to the query event.” (under step 2A prong II and step 2B this amounts to merely indicating a field of use or technological environment in which to apply a judicial exception, see MPEP § 2106.05(h)),
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 11:
Regarding claim 11, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis to claim 2. Further, claim 11 recites the following additional elements:
“determining a corresponding plurality of time differences between a time of each event and each corresponding time of each other event of the plurality of events;” (this is a mental process, a person could mentally determine a plurality of respective time differences between a first time of an event and each corresponding second time of the plurality of key events, see MPEP § 2106.04(a)(2)(III)),
“and generating a corresponding plurality of attention values by aggregating the transformation and the corresponding plurality of time differences.” (this is a mental process, a person can mentally generate attention values by aggregating a transformation and corresponding time differences, see MPEP § 2106.04(a)(2)(III)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
“The method of claim 2, further comprising, for each event of the plurality of events: inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on each event embedding corresponding to each event relative to each other event embedding of the plurality of event embeddings.” (In step 2A, prong 2, this is considered insignificant extra-solution activity of mere data gathering, see MPEP § 2106.05(g)). (In step 2B, this is also considered insignificant extra-solution activity of mere data gathering, which is a well understood routine and conventional activity, see receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362, see MPEP § 2106.05(g)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 12:
Regarding claim 12, it is dependent upon claim 11, and thereby incorporates the limitations of, and corresponding analysis to claim 11. Further, claim 12 recites the following additional elements:
“The method of claim 11, wherein updating the transformer model further comprises aggregating, for each event of the plurality of events, the corresponding plurality of time differences with a plurality of transformations at a plurality of layers of the transformer model.” (this is a mathematical concept, aggregating a plurality of time differences with a plurality of transformations is a mathematical operation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 13:
Regarding claim 13, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “One or more non-transitory, computer-readable media storing instructions that,”, there is no definition of non-transitory computer readable media in the applicant’s specification. Therefore, under BRI non-transitory computer readable media could include signals making the claim signals per se. As such the claim is rejected under failing to fall into the one of the statutory categories of the patent eligible subject matter.
In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components:
“generating a plurality of event embeddings comprising a query event embedding corresponding to the query event and a plurality of key event embeddings corresponding to the plurality of key events;” (this is a mental process, a person could mentally evaluate generating a plurality of event vector representations that comprise a query event vector representation to a query event and a plurality of key event vector representations corresponding to a plurality of key events using pen and paper, see MPEP § 2106.04(a)(2)(III)),
“determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of second events;” (this is a mental process, a person could mentally determine a plurality of respective time differences between a first time of a first event and each corresponding second time of the plurality of key events, see MPEP § 2106.04(a)(2)(III)),
“generating a plurality of attention values by aggregating the transformation and the plurality of respective time differences;” (this is a mental process, a person could mentally evaluate generating a plurality of attention values, see MPEP § 2106.04(a)(2)(III)),
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
“One or more non-transitory, computer-readable media storing instructions that,” (A non-transitory computer readable media is considered generic computer component being used as tool to perform functions of the judicial exception – see MPEP § 2106.05(f)),
“when executed by one or more processors cause the one or more processors to perform operations comprising:” (Using a processor is considered generic computer component being used as tool to perform functions of the judicial exception – see MPEP § 2106.05(f)),
“retrieving a transformer model trained based on a plurality of events, the plurality of events comprising a first event associated with a first time and a plurality of second events associated with a plurality of second times;” (Retrivieving a model trained based on a plurality events is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)),
“inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings;” (Inputting a plurality of event embeddings into a transformer model is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)),
“and updating the transformer model with the plurality of attention values.” (Updating a transformer model with a plurality of attention values is considered mere instructions to apply an exception using generic computer – see MPEP § 2106.05(f)),
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional elements iv and v recites generic computer component being used as tool to perform functions of the judicial exception, additional element viii recites mere instructions to apply an exception using generic computer, and additional elements vi and vii recites insignificant extra-solution activity of mere data gathering, which is a well understood routine and conventional activity, see receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362, which is not indicative of significantly more.
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 14:
Regarding claim 14, it is dependent upon claim 13, and thereby incorporates the limitations of, and corresponding analysis to claim 13. Further, claim 14 recites the following additional elements:
“The one or more non-transitory, computer-readable media of claim 13, wherein aggregating the transformation and the plurality of respective time differences comprises adjusting the transformation based on the plurality of respective time differences such that each attention value accounts for a respective time difference between the first time and a corresponding second time.” (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer, see MPEP § 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using generic computer - see MPEP § 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 15:
Regarding claim 15, it is dependent upon claim 13, and thereby incorporates the limitations of, and corresponding analysis to claim 13. Further, claim 15 recites the following additional elements:
“The one or more non-transitory, computer-readable media of claim 13, wherein aggregating the transformation and the plurality of respective time differences comprises aggregating the transformation and a function of the plurality of respective time differences.” (this is a mathematical concept, aggregating a transformation and a function of a plurality of respective time differences is a mathematical operation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 16:
Regarding claim 16, it is dependent upon claim 15, and thereby incorporates the limitations of, and corresponding analysis to claim 15. Further, claim 16 recites the following additional elements:
“The one or more non-transitory, computer-readable media of claim 15, wherein the function is an exponential decay function and wherein aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation.” (this is a mathematical concept, a function being an exponential decay function and the aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation is a mathematical function or formula and calculation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 17:
Regarding claim 17, it is dependent upon claim 15, and thereby incorporates the limitations of, and corresponding analysis to claim 15. Further, claim 17 recites the following additional elements:
“The one or more non-transitory, computer-readable media of claim 15, wherein the function is an exponential growth function and wherein aggregating the transformation and the plurality of respective time differences comprises subtracting the function of the plurality of respective time differences from the transformation.” (this is a mathematical concept, a function being an exponential growth function and the aggregating the transformation and the plurality of respective time differences comprises subtracting the function of the plurality of respective time differences to the transformation is a mathematical function or formula and calculation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 18:
Regarding claim 18, it is dependent upon claim 13, and thereby incorporates the limitations of, and corresponding analysis to claim 13. Further, claim 18 recites the following additional elements:
“The one or more non-transitory, computer-readable media of claim 13, wherein the transformation comprises a plurality of dot products of a first event embedding corresponding to the first event with each of a plurality of second event embeddings corresponding to the plurality of second events.” (this is a mathematical concept, a transformation comprising of dot products is a mathematical calculation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 19:
Regarding claim 19, it is dependent upon claim 13, and thereby incorporates the limitations of, and corresponding analysis to claim 13. Further, claim 19 recites the following additional elements:
“The one or more non-transitory, computer-readable media of claim 13, wherein generating the plurality of attention values further comprises normalizing, using a softmax function, values generated by aggregating the transformation and the plurality of respective time differences.” (this is a mathematical concept, normalizing using a softmax function is a mathematical operation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 20:
Regarding claim 20, it is dependent upon claim 13, and thereby incorporates the limitations of, and corresponding analysis to claim 13. Further, claim 20 recites the following additional elements:
“The one or more non-transitory, computer-readable media of claim 13, wherein updating the transformer model with the plurality of attention values comprises aggregating the plurality of respective time differences with a plurality of transformations at a plurality of layers of the transformer model.” (this is a mathematical concept, aggregating a plurality of respective time differences with a plurality of transformations is a mathematical operation, see MPEP § 2106.04(a)(2)(I)),
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-5, 7-16, and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Son H. et al, (US. Patent Application Publication 20240153130 A1) effectively filed on June 28, 2023, (hereafter Son), in view of Xu D. et al, "Self-attention with Functional Time Representation Learning", available at https://proceedings.neurips.cc/paper_files/paper/2019/hash/cf34645d98a7630e2bcca98b3e29c8f2-Abstract.html, effectively published in 2019, (hereafter Xu), and further in view of Conway A. et al, (US. Patent Application Publication 20250190459 A1) effectively filed on August 10, 2023, (hereafter Conway).
Claim 1:
Regarding claim 1, Son teaches “A system for updating transformer machine learning models to account for relative timing of events, the system comprising: at least one processor, at least one memory, and computer-readable media having computer-executable instructions stored thereon, the computer-executable instructions, when executed by the at least one processor, causing the system to perform operations comprising: retrieving a transformer model trained based on a plurality of events,…”
See Son in paragraph [0005] describing, “In one general aspect, an apparatus may include a memory configured to store a transformer network comprising a plurality of transformer layers each performing self-attention and cross-attention; one or more processors configured to generate a plurality of feature maps having respective different resolutions based on an input image;” Here, Son establishes an apparatus or system with a memory and processors which when configured/executed store/retrieve a transformer network or transformer model. Further, see Son in paragraph [0023] describing, “The method may include generating a plurality of temporary feature maps for a training input image based on an input of the training input image to a feature map extraction model; calculating, for each of present object queries, temporary position estimation information comprising temporary first position information of a respective bounding box and temporary second position information of respective key points, based on an input of the plurality of generated temporary feature maps to an in-training transformer network;” Here, Son establishes in-training for the transformer network based on an input of the plurality of temporary feature maps which is seen as the events. There are queries associated with the events or objects, of a first temporal position which is being interpreted as a first time, and a second position from a plurality of key points/events.
Further, Son teaches “and updating the transformer model with the plurality of attention values.”
See Son in paragraph [0005] describing, “In one general aspect, an apparatus may include a memory configured to store a transformer network comprising a plurality of transformer layers each performing self-attention and cross-attention; one or more processors configured to generate a plurality of feature maps having respective different resolutions based on an input image; and update, for each of the plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of layers for the updating includes a provision of the plurality of generated feature maps to the transformer network, and wherein each of the plurality of transformer layers comprises a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image;” Here, Son establishes updating of the transformer layers of the transformer model/network and each layer does self-attention implying an update being done for a plurality of values, as known in the art self-attention is a process of calculating attention values.
However, Son did not explicitly teach “…the plurality of events comprising a query event associated with a first time and a plurality of key events associated with a plurality of second times; generating a plurality of event embeddings comprising a query event embedding corresponding to the query event and a plurality of key event embeddings corresponding to the plurality of key events; inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings, the transformation comprising a plurality of dot products of the query event embedding with each of the plurality of key event embeddings; determining a plurality of respective time differences between the first time of the query event and each corresponding second time of the plurality of key events; generating a plurality of attention values by aggregating, with the plurality of dot products, a function of the plurality of respective time differences, each attention value indicating a weight of a corresponding key event of the plurality of key events relative to the query event, and each attention value accounting for a respective time difference between the first time and a corresponding second time, wherein the function comprises an exponential decay function;”
In the same field of art, Xu teaches, ”…the plurality of events comprising a query event associated with a first time and a plurality of key events associated with a plurality of second times;”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, has queries and keys and uses a query-key pair approach using positional encoding, a first query is seen in the formula and a key is also seen, V is for the number of events which can be seen as a plurality of events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding, here t1 can be seen as the first time and the query event, wherein t2 is the second time and can be seen as the second time and a key event.
Further, Xu teaches, “generating a plurality of attention values by aggregating, with the plurality of dot products, a function of the plurality of respective time differences, each attention value indicating a weight of a corresponding key event of the plurality of key events relative to the query event, and each attention value accounting for a respective time difference between the first time and a corresponding second time, wherein the function comprises an exponential decay function;”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding using dot products. Further, see Xu on pages 5-6 section 5 Mercer Time Embedding describing, “
PNG
media_image8.png
734
540
media_image8.png
Greyscale
” Here, Xu establishes an exponential decay function with the Mercer Theorem function as the Fourier coefficients in the function decay exponentially and it is used for temporal patterns which is used to identify time differences for the time embeddings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Neither Son or Xu appear to explicitly teach “generating a plurality of event embeddings comprising a query event embedding corresponding to the query event and a plurality of key event embeddings corresponding to the plurality of key events; inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings, the transformation comprising a plurality of dot products of the query event embedding with each of the plurality of key event embeddings; determining a plurality of respective time differences between the first time of the query event and each corresponding second time of the plurality of key events;”
However Conway in the same field of art teaches, “generating a plurality of event embeddings comprising a query event embedding corresponding to the query event and a plurality of key event embeddings corresponding to the plurality of key events;”
See Conway in paragraph [0010] describing, “ In some aspects, the techniques described herein relate to a method, wherein the one or more attributes of the knowledge base include a type of encoder used to create a plurality of embeddings of the knowledge base, the plurality of embeddings representing a plurality of portions of source data.” Here, Conway establishes the creation or generating of a plurality of embeddings of a knowledge base. Further, see Conway in paragraph [0011] describing, “In some aspects, the techniques described herein relate to a method, wherein the one or more attributes of the prompt construction facility include (i) a process by which the prompt construction facility identifies one or more embeddings in the knowledge base matching an embedding representing a query, (ii) a process by which source data corresponding to the identified one or more embeddings is added to a constructed prompt, and/or (iii) a configuration of a prompt template used to construct the constructed prompt.” Here, Conway establishes the embeddings of this knowledge base comprising of a query embedding and a plurality of embeddings corresponding to source data. Further, see Conway in paragraph [0115] describing, “The knowledge base 130 contains a representation of information contained in a corpus of source data. The corpus of source data may be relevant to a user, an organization, a knowledge domain, one or more topics, etc. In some embodiments, the KB 130 also includes the source data from which the KB's information representation is derived. The source data may include any suitable type of data (e.g., text, audio, image, video, time-series, etc.).” Here, Conway establishes examples of source data which is being interpreted as events.
Further, Conway in the same field of art teaches, “inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings, the transformation comprising a plurality of dot products of the query event embedding with each of the plurality of key event embeddings;”
See Conway in paragraph [0031] describing, “In some aspects, the techniques described herein relate to a method including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a prompt into a generative model. The prompt consist of source data that consists of a plurality of embeddings as stated. Further, see Conway in paragraph [0144] describing, “ In some examples, a matching criterion may be satisfied if the value of the vector similarity metric is within a range (e.g., the Euclidean, Manhattan, Minkowski, Chebychev, or L2-squared distance between the vectors is less than a threshold distance; the difference between ‘1’ and the cosine similarity value for the vectors is less than a threshold value, the dot product of the vectors is greater than a threshold value, etc.). As another example, a nearest neighbor algorithm (e.g., k-nearest neighbors, approximate nearest neighbors (ANN), etc.) can be applied to an embedding to identify one or more embeddings that are neighbors of the molecule embedding in the vector database's embedding space. In some examples, with respect to a query embedding, a matching criterion is satisfied for any embeddings that are identified by the nearest neighbor algorithm as being neighbors of the query embedding.” Here, Conway establishes dot products of vectors in relation to embeddings to satisfy a matching criterion. This matching criterion can be satisfied as well for any of the plurality of embeddings, which are seen to be key embeddings here, that are identified as being a neighbor of the query embedding. Further, see Conway in paragraph [0037] describing, “In some aspects, the techniques described herein relate to a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a plurality of embeddings into a generative model. Further, see Conway in paragraph [0115] describing, “The knowledge base 130 contains a representation of information contained in a corpus of source data. The corpus of source data may be relevant to a user, an organization, a knowledge domain, one or more topics, etc. In some embodiments, the KB 130 also includes the source data from which the KB's information representation is derived. The source data may include any suitable type of data (e.g., text, audio, image, video, time-series, etc.). In some embodiments, the information contained in the knowledge database 130 includes information that was not present in the training dataset of the generative model 140. Thus, the knowledge database 130 can address the generative model's “knowledge gap” by making such information accessible to the generative model 140. In some embodiments, the knowledge base 130 includes a structured representation (e.g., a vector representation, for example, vector embeddings) of information extracted from the source data. The generative model 140 may be any suitable generative model (e.g., a deep neural network (DNN), an LLM, a Generative Pre-trained Transformer (GPT), GPT-3, GPT-4, etc.). Any suitable technique (e.g., encoding technique) can be used to generate the knowledge base's structured representations of information.” Here, Conway establishes the knowledge base comprising of the embeddings comprising data of events such as images, audio and time-series, and establishes the generative model being able to be a transformer model. Further, see Conway in paragraph [0144] describing, “Some non-limiting examples of generative models include generative adversarial networks (GANs), variational autoencoders (VAEs), autoregressive models (e.g., large language models (LLMs)), recurrent neural networks (RNNs), transformer-based models, reinforcement learning models for generative tasks, etc. Transformer-based models generally have an encoder-decoder architecture, use an attention mechanism (e.g., scaled dot-product attention, multi-head attention, masked attention, etc.) to model the relationships between different elements in a sequence of content, and perform well when processing long sequences of content. Some non-limiting examples of transformer-based models include Generalized Pre-trained Transformer 4 (GPT-4), DALL-E3, etc.” Here, Conway further establishes the generative model, which can be a transformer model as well, which is able to perform scaled-dot product attention mechanism. As known in the art, attention-mechanism in transformers perform transformations and scaled-dot product attention deals with plurality of dot products of a plurality of query and key event pairs, as known in the art. This mechanism is used as stated to model relationships between different elements in a sequence of content which as previously established is done with the similarity matching of a criterion of the embeddings.
Further, Conway in the same field of art teaches, “determining a plurality of respective time differences between the first time of the query event and each corresponding second time of the plurality of key events;”
See Conway in paragraph [0193] describing, “ In some embodiments, this time may be measured from the time when query (e.g., user input) is provided to (or received by) the monitored AI system to the time when the monitored AI system provides generated content in response to that query. In some embodiments, this time may be measured from the time when a constructed prompt is provided to (or received by) the generative model 140 to the time when the generative model 140 provides generated content in response to that prompt. Additionally or alternatively, the quantitative assessment facility 333 can measure other components of latency, for example, a latency of the prompt construction operation, or latency of any operation or set of operations performed by the monitored AI system during the process of processing a query and/or generating a completion. Latency metrics may be useful for developing and/or monitoring generative AI systems used for time-sensitive applications such as applications that interact with users in real-time. In some embodiments, the values of a latency metric can be monitored individually or in aggregate. In some examples, aggregate latency can be determined and/or reported as a numeric average of the latency across all responses generated by the monitored AI system when processing an evaluation dataset.” Here, Conway establishes a measurement of time between a query and a corresponding second time of generated content which is being interpreted as the key event and also establishes a plurality of responses being generated which can be the plurality of key events.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base references of Xu and Son with the teachings of Conway
by using Son’s teachings of a method and apparatus with attention based object analysis, and Xu’s teachings of self-attention with functional time representation learning and incorporate with Conway’s teachings of a generative AI system using transformer-based models to perform various methods.
One of ordinary skill in the art would be motivated to do so because by integrating Conway’s frameworks into the methods of Xu and Son, which are all in relation to transformer models or networks, one of ordinary skill in the art would bring “The techniques described herein [that] may be used to improve any knowledge base, and may be particularly well-suited to improving retrieval-augmented Gen AI systems. In some cases, these techniques may be applied to Gen AI systems configured to function as retrieval bots that retrieve relevant information in response to a query.” (Conway, paragraph [0121]).
Claim 2:
Regarding claim 2, Son teaches “A method comprising: retrieving a transformer model trained based on a plurality of events,…”
See Son in paragraph [0005] describing, “In one general aspect, an apparatus may include a memory configured to store a transformer network comprising a plurality of transformer layers each performing self-attention and cross-attention; one or more processors configured to generate a plurality of feature maps having respective different resolutions based on an input image;” Here, Son establishes an apparatus or system with a memory and processors which when configured/executed store/retrieve a transformer network or transformer model. Further, see Son in paragraph [0023] describing, “The method may include generating a plurality of temporary feature maps for a training input image based on an input of the training input image to a feature map extraction model; calculating, for each of present object queries, temporary position estimation information comprising temporary first position information of a respective bounding box and temporary second position information of respective key points, based on an input of the plurality of generated temporary feature maps to an in-training transformer network;” Here, Son establishes in-training for the transformer network based on an input of the plurality of temporary feature maps which is seen as the events. There are queries associated with the events or objects, of a first temporal position which is being interpreted as a first time, and a second position from a plurality of key points/events.
Further, Son teaches “and updating the transformer model with the plurality of attention values.”
See Son in paragraph [0005] describing, “In one general aspect, an apparatus may include a memory configured to store a transformer network comprising a plurality of transformer layers each performing self-attention and cross-attention; one or more processors configured to generate a plurality of feature maps having respective different resolutions based on an input image; and update, for each of the plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of layers for the updating includes a provision of the plurality of generated feature maps to the transformer network, and wherein each of the plurality of transformer layers comprises a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image;” Here, Son establishes updating of the transformer layers of the transformer model/network and each layer does self-attention implying an update being done for a plurality of values, as known in the art self-attention is a process of calculating attention values.
However, Son did not explicitly teach “…the plurality of events comprising a first event associated with a first time and a plurality of second events associated with a plurality of second times; generating a plurality of event embeddings for the plurality of events; inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings; determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of second events; generating a plurality of attention values by aggregating the transformation and the plurality of respective time differences;”
In the same field of art, Xu teaches, ”…the plurality of events comprising a query event associated with a first time and a plurality of key events associated with a plurality of second times;”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, has queries and keys and uses a query-key pair approach using positional encoding, a first query is seen in the formula and a key is also seen, V is for the number of events which can be seen as a plurality of events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding, here t1 can be seen as the first time and the query event, wherein t2 is the second time and can be seen as the second time and a key event.
Further, Xu teaches, “generating a plurality of attention values by aggregating the transformation and the plurality of respective time differences;”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu further establishes calculation of a plurality of time differences with the temporal difference.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Neither Son or Xu appear to explicitly teach “ generating a plurality of event embeddings for the plurality of events; inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings; determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of second events;”
However Conway in the same field of art teaches, “generating a plurality of event embeddings for the plurality of events;”
See Conway in paragraph [0010] describing, “ In some aspects, the techniques described herein relate to a method, wherein the one or more attributes of the knowledge base include a type of encoder used to create a plurality of embeddings of the knowledge base, the plurality of embeddings representing a plurality of portions of source data.” Here, Conway establishes the creation or generating of a plurality of embeddings of a knowledge base. Further, see Conway in paragraph [0011] describing, “In some aspects, the techniques described herein relate to a method, wherein the one or more attributes of the prompt construction facility include (i) a process by which the prompt construction facility identifies one or more embeddings in the knowledge base matching an embedding representing a query, (ii) a process by which source data corresponding to the identified one or more embeddings is added to a constructed prompt, and/or (iii) a configuration of a prompt template used to construct the constructed prompt.” Here, Conway establishes the embeddings of this knowledge base comprising of a query embedding and a plurality of embeddings corresponding to source data. Further, see Conway in paragraph [0115] describing, “The knowledge base 130 contains a representation of information contained in a corpus of source data. The corpus of source data may be relevant to a user, an organization, a knowledge domain, one or more topics, etc. In some embodiments, the KB 130 also includes the source data from which the KB's information representation is derived. The source data may include any suitable type of data (e.g., text, audio, image, video, time-series, etc.).” Here, Conway establishes examples of source data which is being interpreted as events.
Further, Conway in the same field of art teaches, “inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings;”
See Conway in paragraph [0031] describing, “In some aspects, the techniques described herein relate to a method including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a prompt into a generative model. The prompt consist of source data that consists of a plurality of embeddings as stated. Further, see Conway in paragraph [0144] describing, “ In some examples, a matching criterion may be satisfied if the value of the vector similarity metric is within a range (e.g., the Euclidean, Manhattan, Minkowski, Chebychev, or L2-squared distance between the vectors is less than a threshold distance; the difference between ‘1’ and the cosine similarity value for the vectors is less than a threshold value, the dot product of the vectors is greater than a threshold value, etc.). As another example, a nearest neighbor algorithm (e.g., k-nearest neighbors, approximate nearest neighbors (ANN), etc.) can be applied to an embedding to identify one or more embeddings that are neighbors of the molecule embedding in the vector database's embedding space. In some examples, with respect to a query embedding, a matching criterion is satisfied for any embeddings that are identified by the nearest neighbor algorithm as being neighbors of the query embedding.” Here, Conway establishes dot products of vectors in relation to embeddings to satisfy a matching criterion. This matching criterion can be satisfied as well for any of the plurality of embeddings, which are seen to be key embeddings here, that are identified as being a neighbor of the query embedding. Further, see Conway in paragraph [0037] describing, “In some aspects, the techniques described herein relate to a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a plurality of embeddings into a generative model. Further, see Conway in paragraph [0115] describing, “The knowledge base 130 contains a representation of information contained in a corpus of source data. The corpus of source data may be relevant to a user, an organization, a knowledge domain, one or more topics, etc. In some embodiments, the KB 130 also includes the source data from which the KB's information representation is derived. The source data may include any suitable type of data (e.g., text, audio, image, video, time-series, etc.). In some embodiments, the information contained in the knowledge database 130 includes information that was not present in the training dataset of the generative model 140. Thus, the knowledge database 130 can address the generative model's “knowledge gap” by making such information accessible to the generative model 140. In some embodiments, the knowledge base 130 includes a structured representation (e.g., a vector representation, for example, vector embeddings) of information extracted from the source data. The generative model 140 may be any suitable generative model (e.g., a deep neural network (DNN), an LLM, a Generative Pre-trained Transformer (GPT), GPT-3, GPT-4, etc.). Any suitable technique (e.g., encoding technique) can be used to generate the knowledge base's structured representations of information.” Here, Conway establishes the knowledge base comprising of the embeddings comprising data of events such as images, audio and time-series, and establishes the generative model being able to be a transformer model. Further, see Conway in paragraph [0144] describing, “Some non-limiting examples of generative models include generative adversarial networks (GANs), variational autoencoders (VAEs), autoregressive models (e.g., large language models (LLMs)), recurrent neural networks (RNNs), transformer-based models, reinforcement learning models for generative tasks, etc. Transformer-based models generally have an encoder-decoder architecture, use an attention mechanism (e.g., scaled dot-product attention, multi-head attention, masked attention, etc.) to model the relationships between different elements in a sequence of content, and perform well when processing long sequences of content. Some non-limiting examples of transformer-based models include Generalized Pre-trained Transformer 4 (GPT-4), DALL-E3, etc.” Here, Conway further establishes the generative model, which can be a transformer model as well, which is able to perform scaled-dot product attention mechanism. As known in the art, attention-mechanism in transformers perform transformations and scaled-dot product attention deals with plurality of dot products of a plurality of query and key event pairs, as known in the art. This mechanism is used as stated to model relationships between different elements in a sequence of content which as previously established is done with the similarity matching of a criterion of the embeddings.
Further, Conway in the same field of art teaches, “determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of second events;”
See Conway in paragraph [0193] describing, “ In some embodiments, this time may be measured from the time when query (e.g., user input) is provided to (or received by) the monitored AI system to the time when the monitored AI system provides generated content in response to that query. In some embodiments, this time may be measured from the time when a constructed prompt is provided to (or received by) the generative model 140 to the time when the generative model 140 provides generated content in response to that prompt. Additionally or alternatively, the quantitative assessment facility 333 can measure other components of latency, for example, a latency of the prompt construction operation, or latency of any operation or set of operations performed by the monitored AI system during the process of processing a query and/or generating a completion. Latency metrics may be useful for developing and/or monitoring generative AI systems used for time-sensitive applications such as applications that interact with users in real-time. In some embodiments, the values of a latency metric can be monitored individually or in aggregate. In some examples, aggregate latency can be determined and/or reported as a numeric average of the latency across all responses generated by the monitored AI system when processing an evaluation dataset.” Here, Conway establishes a measurement of time between a first query and a corresponding second time of generated content which is being interpreted as the second event and also establishes a plurality of responses being generated which can be the plurality of second events.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base references of Xu and Son with the teachings of Conway
by using Son’s teachings of a method and apparatus with attention based object analysis, and Xu’s teachings of self-attention with functional time representation learning and incorporate with Conway’s teachings of a generative AI system using transformer-based models to perform various methods.
One of ordinary skill in the art would be motivated to do so because by integrating Conway’s frameworks into the methods of Xu and Son, which are all in relation to transformer models or networks, one of ordinary skill in the art would bring “The techniques described herein [that] may be used to improve any knowledge base, and may be particularly well-suited to improving retrieval-augmented Gen AI systems. In some cases, these techniques may be applied to Gen AI systems configured to function as retrieval bots that retrieve relevant information in response to a query.” (Conway, paragraph [0121]).
Claim 3:
Regarding claim 3, Son in view of Xu, and further in view of Conway teaches the limitations of claim 2.
Son does not appear to explicitly teach “The method of claim 2, wherein aggregating the transformation and the plurality of respective time differences comprises adjusting the transformation based on the plurality of respective time differences such that each attention value accounts for a respective time difference between the first time and a corresponding second time.”
Further, Xu teaches “The method of claim 2, wherein aggregating the transformation and the plurality of respective time differences comprises adjusting the transformation based on the plurality of respective time differences such that each attention value accounts for a respective time difference between the first time and a corresponding second time.”
See Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on page 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the concatenation or aggregation of time differences for time embeddings and an adjustment to a transformation is done here, with the mapping of the time embeddings. The time embeddings gets a difference between two times with the temporal difference.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 4:
Regarding claim 4, Son in view of Xu, and further in view of Conway teaches the limitations of claim 2.
Son does not appear to explicitly teach “The method of claim 2, wherein aggregating the transformation and the plurality of respective time differences comprises aggregating the transformation and a function of the plurality of respective time differences.”
Further, Xu teaches “The method of claim 2, wherein aggregating the transformation and the plurality of respective time differences comprises aggregating the transformation and a function of the plurality of respective time differences.”
See Xu on page 2 section 2 Related Work describing, “Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings.” Here, Xu establishes the self-attention mechanism of the previous limitation having a concatenation or aggregation of a position from positional encoding with event embeddings. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on page 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the concatenation or aggregation of time differences as a function for time embeddings and an adjustment to a transformation is done here, with the mapping of the time embeddings. The time embeddings gets a difference between two times with the temporal difference. Further, see Xu on page 3 section 3 Preliminaries describing, “So the task of learning temporal patterns is converted to a kernel learning problem with Φ as feature map. Also, the interactions between event embedding and time can now be recognized with some other mappings as
PNG
media_image9.png
20
134
media_image9.png
Greyscale
, which we will discuss in Section 6. By relating time embedding to kernel function learning, we hope to identify Φ with some functional forms which are compatible with current deep learning frameworks.” Here, Xu establishes another function for aggregating a transformation which is the mappings for time differences.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 5:
Regarding claim 5, Son in view of Xu, and further in view of Conway teaches the limitations of claim 4.
Son does not appear to explicitly teach “The method of claim 4, wherein the function is an exponential decay function and wherein aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation.”
Further, Xu teaches “The method of claim 4, wherein the function is an exponential decay function and wherein aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation.”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding using dot products. You can also see that the function here is being added to the transformation as the time differences are being added to the transformation in the formula given. Further, see Xu on pages 5-6 section 5 Mercer Time Embedding describing, “
PNG
media_image8.png
734
540
media_image8.png
Greyscale
” Here, Xu establishes an exponential decay function with the Mercer Theorem function as the Fourier coefficients in the function decay exponentially and it is used for temporal patterns which is used to identify time differences for the time embeddings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 7:
Regarding claim 7, Son in view of Xu, and further in view of Conway teaches the limitations of claim 2.
Further, Son teaches “The method of claim 2, wherein the transformation comprises a plurality of dot products of a first event embedding corresponding to the first event with each of a plurality of second event embeddings corresponding to the plurality of second events.”
See Son in paragraph [0068] describing, “For example, the self-attention model 320 may calculate a score matrix representing a relation between a query and a key. The score matrix may be derived by a scaled dot-product attention operation as shown in Equation 1 below.
PNG
media_image10.png
48
64
media_image10.png
Greyscale
” Here, Son establishes self-attention being done using a dot product between a query which can be seen as a first event and a key which can be a second event. Further, see Son in paragraph [0069] describing, “In Equation 1, S represents a score matrix, Q represents a query vector, K represents a key vector, T represents a transpose matrix, and dk represents a dimension of a key vector.” Here, Son establishes the query and keys as vectors which are embeddings. Further, see Son in paragraph [0015] describing, “In another general aspect, a processor-implemented method may include generating a plurality of feature maps having respective different resolutions based on an input image; and updating, for each of a plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of transformer layers for the updating includes a provision of the plurality of generated feature maps to a transformer network comprising the plurality of layers each performing self-attention and cross-attention, wherein each of the plurality of transformer layers comprises: a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image; and a cross-attention model configured to generate respective output data by performing cross-attention on respective one or more feature maps among the plurality of feature maps and the respective generated intermediate data.” Here, Son establishes the feature maps which is a type of transformation, comprising of a plurality of transformer layers who each do self-attention, since each does self-attention it can be seen as a plurality of dot products being done.
Claim 8:
Regarding claim 8, Son in view of Xu, and further in view of Conway teaches the limitations of claim 2.
Further, Son teaches “The method of claim 2, wherein generating the plurality of attention values further comprises normalizing, using a softmax function,...”
See Son in paragraph [0070] describing, “The self-attention model 320 may apply softmax to a result (e.g., the score matrix S) of an operation between the query vector and the key vector. Softmax may be understood as normalizing the result of the operation between the query vector and the key vector, and may be performed by a corresponding Softmax layer.” Here, Son establishes normalizing of the attention value using softmax.
Son does not appear to explicitly teach “values generated by aggregating the transformation and the plurality of respective time differences.”
Further, Xu teaches “values generated by aggregating the transformation and the plurality of respective time differences.”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embeddings that calculates time differences.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 9:
Regarding claim 9, Son in view of Xu, and further in view of Conway teaches the limitations of claim 2.
Further, Son teaches “The method of claim 2, wherein updating the transformer model with…at a plurality of layers of the transformer model.”
See Son in paragraph [0015] describing, “In another general aspect, a processor-implemented method may include generating a plurality of feature maps having respective different resolutions based on an input image; and updating, for each of a plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of transformer layers for the updating includes a provision of the plurality of generated feature maps to a transformer network comprising the plurality of layers each performing self-attention and cross-attention, wherein each of the plurality of transformer layers comprises: a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image; and a cross-attention model configured to generate respective output data by performing cross-attention on respective one or more feature maps among the plurality of feature maps and the respective generated intermediate data.” Here, Son establishes updating the model using a plurality of layers of the transformer network or model, which each do self-attention known to get an attention value, the respective position information is being interpreted here as the time differences and in an analogous system can be. The respective position information is for the transformation of feature maps.
Son does not appear to explicitly teach “…the plurality of attention values comprises aggregating the plurality of respective time differences with a plurality of transformations...”
Further, Xu teaches “…the plurality of attention values comprises aggregating the plurality of respective time differences with a plurality of transformations...”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embeddings that calculates time differences. Xu also shows a plurality time differences being added or aggregated to a plurality of transformations as the formula establishes at least two mappings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Son with the teaching of Xu by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 10:
Regarding claim 10, Son in view of Xu, and further in view of Conway teaches the limitations of claim 2.
Further, Son teaches “The method of claim 2, wherein the first event is a query event and wherein the plurality of second events is a plurality of key events, wherein each key event is weighted relative to the query event.”
See Son in paragraph [0005] describing, “In one general aspect, an apparatus may include a memory configured to store a transformer network comprising a plurality of transformer layers each performing self-attention and cross-attention; one or more processors configured to generate a plurality of feature maps having respective different resolutions based on an input image; and update, for each of the plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of layers for the updating includes a provision of the plurality of generated feature maps to the transformer network, and wherein each of the plurality of transformer layers comprises a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image;” Here, Son establishes a first position which can be seen as a first event corresponding to a query, and a second position which can be a second event from a plurality of key points or key events which corresponds to the query which can be seen as it being relative to the query event. Further, see Son in paragraph [0065] describing, “he self-attention model 320 of the first layer 311 may perform self-attention by applying weights to portions to be inferred important in specific data and reflecting this again. The self-attention model 320 of the first layer 311 may include an attention layer, which may be configured to calculate attention weightings. Thus, the computing apparatus may improve an object query through the self-attention. In other words, a key, query, and value input to the self-attention model 320 may be information included in the object query. The value of the self-attention model 320 may be set based on the content information corresponding to the object query, and the key and the query of the self-attention model 320 may be set based on the content information and the position estimation information corresponding to the object query.” Here, Son further establishes weighting for the self-attention and how weighting applies to the query and the key.
Claim 11:
Regarding claim 11, Son in view of Xu, and further in view of Conway teaches the limitations of claim 2.
Son does not appear to explicitly teach “The method of claim 2, further comprising, for each event of the plurality of events: inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on each event embedding corresponding to each event relative to each other event embedding of the plurality of event embeddings; determining a corresponding plurality of time differences between a time of each event and each corresponding time of each other event of the plurality of events; and generating a corresponding plurality of attention values by aggregating the transformation and the corresponding plurality of time differences.”
Further, Xu teaches “and generating a corresponding plurality of attention values by aggregating the transformation and the corresponding plurality of time differences.”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding using dot products. Further, see Xu on pages 5-6 section 5 Mercer Time Embedding describing, “
PNG
media_image8.png
734
540
media_image8.png
Greyscale
” Here, Xu establishes an exponential decay function with the Mercer Theorem function as the Fourier coefficients in the function decay exponentially and it is used for temporal patterns which is used to identify time differences for the time embeddings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Neither Son or Xu appear to explicitly teach “The method of claim 2, further comprising, for each event of the plurality of events: inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on each event embedding corresponding to each event relative to each other event embedding of the plurality of event embeddings; determining a corresponding plurality of time differences between a time of each event and each corresponding time of each other event of the plurality of events;”
However Conway in the same field of art teaches, “The method of claim 2, further comprising, for each event of the plurality of events: inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on each event embedding corresponding to each event relative to each other event embedding of the plurality of event embeddings;”
Conway in paragraph [0031] describing, “In some aspects, the techniques described herein relate to a method including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a prompt into a generative model. The prompt consist of source data that consists of a plurality of embeddings as stated, by inputting the prompt to the generative model a plurality of embeddings is also input. Further, see Conway in paragraph [0144] describing, “ In some examples, a matching criterion may be satisfied if the value of the vector similarity metric is within a range (e.g., the Euclidean, Manhattan, Minkowski, Chebychev, or L2-squared distance between the vectors is less than a threshold distance; the difference between ‘1’ and the cosine similarity value for the vectors is less than a threshold value, the dot product of the vectors is greater than a threshold value, etc.). As another example, a nearest neighbor algorithm (e.g., k-nearest neighbors, approximate nearest neighbors (ANN), etc.) can be applied to an embedding to identify one or more embeddings that are neighbors of the molecule embedding in the vector database's embedding space. In some examples, with respect to a query embedding, a matching criterion is satisfied for any embeddings that are identified by the nearest neighbor algorithm as being neighbors of the query embedding.” Here, Conway establishes dot products of vectors in relation to embeddings to satisfy a matching criterion. This matching criterion can be satisfied as well for any of the plurality of embeddings, which are seen to be key embeddings here, that are identified as being a neighbor of the query embedding. Further, see Conway in paragraph [0037] describing, “In some aspects, the techniques described herein relate to a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a plurality of embeddings into a generative model. Further, see Conway in paragraph [0115] describing, “The knowledge base 130 contains a representation of information contained in a corpus of source data. The corpus of source data may be relevant to a user, an organization, a knowledge domain, one or more topics, etc. In some embodiments, the KB 130 also includes the source data from which the KB's information representation is derived. The source data may include any suitable type of data (e.g., text, audio, image, video, time-series, etc.). In some embodiments, the information contained in the knowledge database 130 includes information that was not present in the training dataset of the generative model 140. Thus, the knowledge database 130 can address the generative model's “knowledge gap” by making such information accessible to the generative model 140. In some embodiments, the knowledge base 130 includes a structured representation (e.g., a vector representation, for example, vector embeddings) of information extracted from the source data. The generative model 140 may be any suitable generative model (e.g., a deep neural network (DNN), an LLM, a Generative Pre-trained Transformer (GPT), GPT-3, GPT-4, etc.). Any suitable technique (e.g., encoding technique) can be used to generate the knowledge base's structured representations of information.” Here, Conway establishes the knowledge base comprising of the embeddings comprising data of events such as images, audio and time-series, and establishes the generative model being able to be a transformer model. Further, see Conway in paragraph [0144] describing, “Some non-limiting examples of generative models include generative adversarial networks (GANs), variational autoencoders (VAEs), autoregressive models (e.g., large language models (LLMs)), recurrent neural networks (RNNs), transformer-based models, reinforcement learning models for generative tasks, etc. Transformer-based models generally have an encoder-decoder architecture, use an attention mechanism (e.g., scaled dot-product attention, multi-head attention, masked attention, etc.) to model the relationships between different elements in a sequence of content, and perform well when processing long sequences of content. Some non-limiting examples of transformer-based models include Generalized Pre-trained Transformer 4 (GPT-4), DALL-E3, etc.” Here, Conway further establishes the generative model, which can be a transformer model as well, which is able to perform scaled-dot product attention mechanism. As known in the art, attention-mechanism in transformers perform transformations. This mechanism is used as stated to model relationships between different elements in a sequence of content which as previously established is done with the similarity matching of a criterion of the embeddings, modeling of relationships can be further seen as embeddings corresponding to each event relative to other event embeddings.
Further, Conway teaches, “determining a corresponding plurality of time differences between a time of each event and each corresponding time of each other event of the plurality of events;”
See Conway in paragraph [0193] describing, “ In some embodiments, this time may be measured from the time when query (e.g., user input) is provided to (or received by) the monitored AI system to the time when the monitored AI system provides generated content in response to that query. In some embodiments, this time may be measured from the time when a constructed prompt is provided to (or received by) the generative model 140 to the time when the generative model 140 provides generated content in response to that prompt. Additionally or alternatively, the quantitative assessment facility 333 can measure other components of latency, for example, a latency of the prompt construction operation, or latency of any operation or set of operations performed by the monitored AI system during the process of processing a query and/or generating a completion. Latency metrics may be useful for developing and/or monitoring generative AI systems used for time-sensitive applications such as applications that interact with users in real-time. In some embodiments, the values of a latency metric can be monitored individually or in aggregate. In some examples, aggregate latency can be determined and/or reported as a numeric average of the latency across all responses generated by the monitored AI system when processing an evaluation dataset.” Here, Conway establishes a measurement of time between a query and a corresponding second time of generated content which is being interpreted as the key event and also establishes a plurality of responses being generated which can be the plurality of key events.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base references of Xu and Son with the teachings of Conway by using Son’s teachings of a method and apparatus with attention based object analysis, and Xu’s teachings of self-attention with functional time representation learning and incorporate with Conway’s teachings of a generative AI system using transformer-based models to perform various methods.
One of ordinary skill in the art would be motivated to do so because by integrating Conway’s frameworks into the methods of Xu and Son, which are all in relation to transformer models or networks, one of ordinary skill in the art would bring “The techniques described herein [that] may be used to improve any knowledge base, and may be particularly well-suited to improving retrieval-augmented Gen AI systems. In some cases, these techniques may be applied to Gen AI systems configured to function as retrieval bots that retrieve relevant information in response to a query.” (Conway, paragraph [0121]).
Claim 12:
Regarding claim 12, Son in view of Xu, and further in view of Conway teaches the limitations of claim 11.
Further, Son teaches “The method of claim 11, wherein updating the transformer model further comprises…at a plurality of layers of the transformer model.”
See Son in paragraph [0015] describing, “In another general aspect, a processor-implemented method may include generating a plurality of feature maps having respective different resolutions based on an input image; and updating, for each of a plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of transformer layers for the updating includes a provision of the plurality of generated feature maps to a transformer network comprising the plurality of layers each performing self-attention and cross-attention, wherein each of the plurality of transformer layers comprises: a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image; and a cross-attention model configured to generate respective output data by performing cross-attention on respective one or more feature maps among the plurality of feature maps and the respective generated intermediate data.” Here, Son establishes updating the model using a plurality of layers of the transformer network or model, which each do self-attention known to get an attention value, the respective position information is being interpreted here as the time differences and in an analogous system can be. The respective position information is for the transformation of feature maps.
Son does not appear to explicitly teach “…aggregating, for each event of the plurality of events, the corresponding plurality of time differences with a plurality of transformations...”
Further, Xu teaches “…aggregating, for each event of the plurality of events, the corresponding plurality of time differences with a plurality of transformations...”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embeddings that calculates time differences. Xu also shows a plurality time differences being added or aggregated to a plurality of transformations as the formula establishes at least two mappings, and the time differences are each corresponding to its mapping in the formula.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Son with the teaching of Xu by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 13:
Regarding claim 13, Son teaches “One or more non-transitory, computer-readable media storing instructions that, when executed by one or more processors cause the one or more processors to perform operations comprising: retrieving a transformer model trained based on a plurality of events,…”
See Son in paragraph [0005] describing, “In one general aspect, an apparatus may include a memory configured to store a transformer network comprising a plurality of transformer layers each performing self-attention and cross-attention; one or more processors configured to generate a plurality of feature maps having respective different resolutions based on an input image;” Here, Son establishes an apparatus or system with a memory and processors which when configured/executed store/retrieve a transformer network or transformer model. Further, see Son in paragraph [0023] describing, “The method may include generating a plurality of temporary feature maps for a training input image based on an input of the training input image to a feature map extraction model; calculating, for each of present object queries, temporary position estimation information comprising temporary first position information of a respective bounding box and temporary second position information of respective key points, based on an input of the plurality of generated temporary feature maps to an in-training transformer network;” Here, Son establishes in-training for the transformer network based on an input of the plurality of temporary feature maps which is seen as the events. There are queries associated with the events or objects, of a first temporal position which is being interpreted as a first time, and a second position from a plurality of key points/events.
Further, Son teaches “and updating the transformer model with the plurality of attention values.”
See Son in paragraph [0005] describing, “In one general aspect, an apparatus may include a memory configured to store a transformer network comprising a plurality of transformer layers each performing self-attention and cross-attention; one or more processors configured to generate a plurality of feature maps having respective different resolutions based on an input image; and update, for each of the plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of layers for the updating includes a provision of the plurality of generated feature maps to the transformer network, and wherein each of the plurality of transformer layers comprises a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image;” Here, Son establishes updating of the transformer layers of the transformer model/network and each layer does self-attention implying an update being done for a plurality of values, as known in the art self-attention is a process of calculating attention values.
However, Son did not explicitly teach “…the plurality of events comprising a first event associated with a first time and a plurality of second events associated with a plurality of second times; generating a plurality of event embeddings for the plurality of events; inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings; determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of second events; generating a plurality of attention values by aggregating the transformation and the plurality of respective time differences;”
In the same field of art, Xu teaches, ”… the plurality of events comprising a first event associated with a first time and a plurality of second events associated with a plurality of second times;”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, has queries and keys and uses a query-key pair approach using positional encoding, a first query is seen in the formula and a key is also seen, V is for the number of events which can be seen as a plurality of events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding, here t1 can be seen as the first time and the query event, wherein t2 is the second time and can be seen as the second time and a key event.
Further, Xu teaches, “generating a plurality of attention values by aggregating the transformation and the plurality of respective time differences;”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu further establishes calculation of a plurality of time differences with the temporal difference.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Neither Son or Xu appear to explicitly teach “ generating a plurality of event embeddings for the plurality of events; inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings; determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of second events;”
However Conway in the same field of art teaches, “generating a plurality of event embeddings for the plurality of events;”
See Conway in paragraph [0010] describing, “ In some aspects, the techniques described herein relate to a method, wherein the one or more attributes of the knowledge base include a type of encoder used to create a plurality of embeddings of the knowledge base, the plurality of embeddings representing a plurality of portions of source data.” Here, Conway establishes the creation or generating of a plurality of embeddings of a knowledge base. Further, see Conway in paragraph [0011] describing, “In some aspects, the techniques described herein relate to a method, wherein the one or more attributes of the prompt construction facility include (i) a process by which the prompt construction facility identifies one or more embeddings in the knowledge base matching an embedding representing a query, (ii) a process by which source data corresponding to the identified one or more embeddings is added to a constructed prompt, and/or (iii) a configuration of a prompt template used to construct the constructed prompt.” Here, Conway establishes the embeddings of this knowledge base comprising of a query embedding and a plurality of embeddings corresponding to source data. Further, see Conway in paragraph [0115] describing, “The knowledge base 130 contains a representation of information contained in a corpus of source data. The corpus of source data may be relevant to a user, an organization, a knowledge domain, one or more topics, etc. In some embodiments, the KB 130 also includes the source data from which the KB's information representation is derived. The source data may include any suitable type of data (e.g., text, audio, image, video, time-series, etc.).” Here, Conway establishes examples of source data which is being interpreted as events.
Further, Conway in the same field of art teaches, “inputting, into the transformer model, the plurality of event embeddings to cause the transformer model to perform a transformation on the plurality of event embeddings;”
See Conway in paragraph [0031] describing, “In some aspects, the techniques described herein relate to a method including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a prompt into a generative model. The prompt consist of source data that consists of a plurality of embeddings as stated. Further, see Conway in paragraph [0144] describing, “ In some examples, a matching criterion may be satisfied if the value of the vector similarity metric is within a range (e.g., the Euclidean, Manhattan, Minkowski, Chebychev, or L2-squared distance between the vectors is less than a threshold distance; the difference between ‘1’ and the cosine similarity value for the vectors is less than a threshold value, the dot product of the vectors is greater than a threshold value, etc.). As another example, a nearest neighbor algorithm (e.g., k-nearest neighbors, approximate nearest neighbors (ANN), etc.) can be applied to an embedding to identify one or more embeddings that are neighbors of the molecule embedding in the vector database's embedding space. In some examples, with respect to a query embedding, a matching criterion is satisfied for any embeddings that are identified by the nearest neighbor algorithm as being neighbors of the query embedding.” Here, Conway establishes dot products of vectors in relation to embeddings to satisfy a matching criterion. This matching criterion can be satisfied as well for any of the plurality of embeddings, which are seen to be key embeddings here, that are identified as being a neighbor of the query embedding. Further, see Conway in paragraph [0037] describing, “In some aspects, the techniques described herein relate to a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: selecting, by one or more processors, a plurality of clusters of embeddings from an embedding space of a knowledge base; for each cluster of embeddings in the plurality of clusters of embeddings, selecting, by the one or more processors, one or more embeddings from the respective cluster of embeddings; obtaining a respective portion of source data represented by the selected one or more embeddings; constructing, by the one or more processors, a prompt based on the obtained portion of source data, wherein the prompt relates to generating a question about the portion of source data and an answer to the question; providing the prompt as input to a generative model;” Here, Conway establishes inputting a plurality of embeddings into a generative model. Further, see Conway in paragraph [0115] describing, “The knowledge base 130 contains a representation of information contained in a corpus of source data. The corpus of source data may be relevant to a user, an organization, a knowledge domain, one or more topics, etc. In some embodiments, the KB 130 also includes the source data from which the KB's information representation is derived. The source data may include any suitable type of data (e.g., text, audio, image, video, time-series, etc.). In some embodiments, the information contained in the knowledge database 130 includes information that was not present in the training dataset of the generative model 140. Thus, the knowledge database 130 can address the generative model's “knowledge gap” by making such information accessible to the generative model 140. In some embodiments, the knowledge base 130 includes a structured representation (e.g., a vector representation, for example, vector embeddings) of information extracted from the source data. The generative model 140 may be any suitable generative model (e.g., a deep neural network (DNN), an LLM, a Generative Pre-trained Transformer (GPT), GPT-3, GPT-4, etc.). Any suitable technique (e.g., encoding technique) can be used to generate the knowledge base's structured representations of information.” Here, Conway establishes the knowledge base comprising of the embeddings comprising data of events such as images, audio and time-series, and establishes the generative model being able to be a transformer model. Further, see Conway in paragraph [0087] describing, “Some non-limiting examples of generative models include generative adversarial networks (GANs), variational autoencoders (VAEs), autoregressive models (e.g., large language models (LLMs)), recurrent neural networks (RNNs), transformer-based models, reinforcement learning models for generative tasks, etc. Transformer-based models generally have an encoder-decoder architecture, use an attention mechanism (e.g., scaled dot-product attention, multi-head attention, masked attention, etc.) to model the relationships between different elements in a sequence of content, and perform well when processing long sequences of content. Some non-limiting examples of transformer-based models include Generalized Pre-trained Transformer 4 (GPT-4), DALL-E3, etc.” Here, Conway further establishes the generative model, which can be a transformer model as well, which is able to perform scaled-dot product attention mechanism. As known in the art, attention-mechanism in transformers perform transformations and scaled-dot product attention deals with plurality of dot products. Further, see Conway in paragraph [0144] describing, “Some non-limiting examples of generative models include generative adversarial networks (GANs), variational autoencoders (VAEs), autoregressive models (e.g., large language models (LLMs)), recurrent neural networks (RNNs), transformer-based models, reinforcement learning models for generative tasks, etc. Transformer-based models generally have an encoder-decoder architecture, use an attention mechanism (e.g., scaled dot-product attention, multi-head attention, masked attention, etc.) to model the relationships between different elements in a sequence of content, and perform well when processing long sequences of content. Some non-limiting examples of transformer-based models include Generalized Pre-trained Transformer 4 (GPT-4), DALL-E3, etc.” Here, Conway further establishes the generative model, which can be a transformer model as well, which is able to perform scaled-dot product attention mechanism. As known in the art, attention-mechanism in transformers perform transformations and scaled-dot product attention deals with plurality of dot products of a plurality of query and key event pairs, as known in the art. This mechanism is used as stated to model relationships between different elements in a sequence of content which as previously established is done with the similarity matching of a criterion of the embeddings.
Further, Conway in the same field of art teaches, “determining a plurality of respective time differences between the first time of the first event and each corresponding second time of the plurality of second events;”
See Conway in paragraph [0193] describing, “ In some embodiments, this time may be measured from the time when query (e.g., user input) is provided to (or received by) the monitored AI system to the time when the monitored AI system provides generated content in response to that query. In some embodiments, this time may be measured from the time when a constructed prompt is provided to (or received by) the generative model 140 to the time when the generative model 140 provides generated content in response to that prompt. Additionally or alternatively, the quantitative assessment facility 333 can measure other components of latency, for example, a latency of the prompt construction operation, or latency of any operation or set of operations performed by the monitored AI system during the process of processing a query and/or generating a completion. Latency metrics may be useful for developing and/or monitoring generative AI systems used for time-sensitive applications such as applications that interact with users in real-time. In some embodiments, the values of a latency metric can be monitored individually or in aggregate. In some examples, aggregate latency can be determined and/or reported as a numeric average of the latency across all responses generated by the monitored AI system when processing an evaluation dataset.” Here, Conway establishes a measurement of time between a first query and a corresponding second time of generated content which is being interpreted as the second event and also establishes a plurality of responses being generated which can be the plurality of second events.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base references of Xu and Son with the teachings of Conway
by using Son’s teachings of a method and apparatus with attention based object analysis, and Xu’s teachings of self-attention with functional time representation learning and incorporate with Conway’s teachings of a generative AI system using transformer-based models to perform various methods.
One of ordinary skill in the art would be motivated to do so because by integrating Conway’s frameworks into the methods of Xu and Son, which are all in relation to transformer models or networks, one of ordinary skill in the art would bring “The techniques described herein [that] may be used to improve any knowledge base, and may be particularly well-suited to improving retrieval-augmented Gen AI systems. In some cases, these techniques may be applied to Gen AI systems configured to function as retrieval bots that retrieve relevant information in response to a query.” (Conway, paragraph [0121]).
Claim 14:
Regarding claim 14, Son in view of Xu, and further in view of Conway teaches the limitations of claim 13.
Son does not appear to explicitly teach “The one or more non-transitory, computer-readable media of claim 13, wherein aggregating the transformation and the plurality of respective time differences comprises adjusting the transformation based on the plurality of respective time differences such that each attention value accounts for a respective time difference between the first time and a corresponding second time.”
Further, Xu teaches “The one or more non-transitory, computer-readable media of claim 13, wherein aggregating the transformation and the plurality of respective time differences comprises adjusting the transformation based on the plurality of respective time differences such that each attention value accounts for a respective time difference between the first time and a corresponding second time.”
See Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on page 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the concatenation or aggregation of time differences for time embeddings and an adjustment to a transformation is done here, with the mapping of the time embeddings. The time embeddings gets a difference between two times with the temporal difference.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 15:
Regarding claim 15, Son in view of Xu, and further in view of Conway teaches the limitations of claim 13.
Son does not appear to explicitly teach “The one or more non-transitory, computer-readable media of claim 13, wherein aggregating the transformation and the plurality of respective time differences comprises aggregating the transformation and a function of the plurality of respective time differences.”
Further, Xu teaches “The one or more non-transitory, computer-readable media of claim 13, wherein aggregating the transformation and the plurality of respective time differences comprises aggregating the transformation and a function of the plurality of respective time differences.”
See Xu on page 2 section 2 Related Work describing, “Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings.” Here, Xu establishes the self-attention mechanism of the previous limitation having a concatenation or aggregation of a position from positional encoding with event embeddings. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on page 3 section 3 Preliminaries describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the concatenation or aggregation of time differences as a function for time embeddings and an adjustment to a transformation is done here, with the mapping of the time embeddings. The time embeddings gets a difference between two times with the temporal difference. Further, see Xu on page 3 section 3 Preliminaries describing, “So the task of learning temporal patterns is converted to a kernel learning problem with Φ as feature map. Also, the interactions between event embedding and time can now be recognized with some other mappings as
PNG
media_image9.png
20
134
media_image9.png
Greyscale
, which we will discuss in Section 6. By relating time embedding to kernel function learning, we hope to identify Φ with some functional forms which are compatible with current deep learning frameworks.” Here, Xu establishes another function for aggregating a transformation which is the mappings for time differences.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 16:
Regarding claim 16, Son in view of Xu, and further in view of Conway teaches the limitations of claim 15.
Son does not appear to explicitly teach “The one or more non-transitory, computer-readable media of claim 15, wherein the function is an exponential decay function and wherein aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation.”
Further, Xu teaches “The one or more non-transitory, computer-readable media of claim 15, wherein the function is an exponential decay function and wherein aggregating the transformation and the plurality of respective time differences comprises adding the function of the plurality of respective time differences to the transformation.”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding using dot products. You can also see that the function here is being added to the transformation as the time differences are being added to the transformation in the formula given. Further, see Xu on pages 5-6 section 5 Mercer Time Embedding describing, “
PNG
media_image8.png
734
540
media_image8.png
Greyscale
” Here, Xu establishes an exponential decay function with the Mercer Theorem function as the Fourier coefficients in the function decay exponentially and it is used for temporal patterns which is used to identify time differences for the time embeddings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 18:
Regarding claim 18, Son in view of Xu, and further in view of Conway teaches the limitations of claim 13.
Further, Son teaches “The one or more non-transitory, computer-readable media of claim 13, wherein the transformation comprises a plurality of dot products of a first event embedding corresponding to the first event with each of a plurality of second event embeddings corresponding to the plurality of second events.”
See Son in paragraph [0068] describing, “For example, the self-attention model 320 may calculate a score matrix representing a relation between a query and a key. The score matrix may be derived by a scaled dot-product attention operation as shown in Equation 1 below.
PNG
media_image10.png
48
64
media_image10.png
Greyscale
” Here, Son establishes self-attention being done using a dot product between a query which can be seen as a first event and a key which can be a second event. Further, see Son in paragraph [0069] describing, “In Equation 1, S represents a score matrix, Q represents a query vector, K represents a key vector, T represents a transpose matrix, and dk represents a dimension of a key vector.” Here, Son establishes the query and keys as vectors which are embeddings. Further, see Son in paragraph [0015] describing, “In another general aspect, a processor-implemented method may include generating a plurality of feature maps having respective different resolutions based on an input image; and updating, for each of a plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of transformer layers for the updating includes a provision of the plurality of generated feature maps to a transformer network comprising the plurality of layers each performing self-attention and cross-attention, wherein each of the plurality of transformer layers comprises: a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image; and a cross-attention model configured to generate respective output data by performing cross-attention on respective one or more feature maps among the plurality of feature maps and the respective generated intermediate data.” Here, Son establishes the feature maps which is a type of transformation, comprising of a plurality of transformer layers who each do self-attention, since each does self-attention it can be seen as a plurality of dot products being done.
Claim 19:
Regarding claim 19, Son in view of Xu, and further in view of Conway teaches the limitations of claim 13.
Further, Son teaches “The one or more non-transitory, computer-readable media of claim 13, wherein generating the plurality of attention values further comprises normalizing, using a softmax function,...”
See Son in paragraph [0070] describing, “The self-attention model 320 may apply softmax to a result (e.g., the score matrix S) of an operation between the query vector and the key vector. Softmax may be understood as normalizing the result of the operation between the query vector and the key vector, and may be performed by a corresponding Softmax layer.” Here, Son establishes normalizing of the attention value using softmax.
Son does not appear to explicitly teach “values generated by aggregating the transformation and the plurality of respective time differences.”
Further, Xu teaches “values generated by aggregating the transformation and the plurality of respective time differences.”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embeddings that calculates time differences.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim 20:
Regarding claim 20, Son in view of Xu, and further in view of Conway teaches the limitations of claim 13.
Further, Son teaches “The one or more non-transitory, computer-readable media of claim 13, wherein updating the transformer model with…at a plurality of layers of the transformer model.”
See Son in paragraph [0015] describing, “In another general aspect, a processor-implemented method may include generating a plurality of feature maps having respective different resolutions based on an input image; and updating, for each of a plurality of transformer layers, respective position estimation information comprising first position information of a respective bounding box corresponding to one object query and second position information of respective key points corresponding to the one object query, wherein an implementation of the plurality of transformer layers for the updating includes a provision of the plurality of generated feature maps to a transformer network comprising the plurality of layers each performing self-attention and cross-attention, wherein each of the plurality of transformer layers comprises: a self-attention model configured to generate respective intermediate data by performing self-attention on respective content information on a feature of the input image; and a cross-attention model configured to generate respective output data by performing cross-attention on respective one or more feature maps among the plurality of feature maps and the respective generated intermediate data.” Here, Son establishes updating the model using a plurality of layers of the transformer network or model, which each do self-attention known to get an attention value, the respective position information is being interpreted here as the time differences and in an analogous system can be. The respective position information is for the transformation of feature maps.
Son does not appear to explicitly teach “…the plurality of attention values comprises aggregating the plurality of respective time differences with a plurality of transformations...”
Further, Xu teaches “…the plurality of attention values comprises aggregating the plurality of respective time differences with a plurality of transformations...”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embeddings that calculates time differences. Xu also shows a plurality time differences being added or aggregated to a plurality of transformations as the formula establishes at least two mappings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Son with the teaching of Xu by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Claim(s) 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Son H. et al, in view of Xu D. et al, further in view of Conway A. et al, and further Wang Y. et al, " Characterization of MPC-based Private Inference for Transformer-based Models", available at https://ieeexplore.ieee.org/abstract/document/9804616, effectively published on June 27, 2022, (hereafter Wang).
Claim 6:
Regarding claim 6, Son in view of Xu, and further in view of Conway teaches the limitations of claim 4.
Son does not appear to explicitly teach “The method of claim 4, wherein the function is an exponential growth function and wherein aggregating the transformation and the plurality of respective time differences comprises subtracting the function of the plurality of respective time differences from the transformation.”
Further, Xu teaches “…and wherein aggregating the transformation and the plurality of respective time differences comprises...”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding using dot products. You can also see that the function here is being added to the transformation as the time differences are being added to the transformation in the formula given.
Further, Xu teaches “…the plurality of respective time differences...”
See Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes a function here of a plurality of time differences are being added to the transformation in the formula given to its respective mapping.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Neither Son, Xu, or Conway appear to explicitly teach “The method of claim 4, wherein the function is an exponential growth function…comprises subtracting the function of…from the transformation.”
However Wang in an analogous system teaches, “The method of claim 4, wherein the function is an exponential growth function…comprises subtracting the function of…from the transformation.”
See Wang on Page 8 Section III subsection D. Analysis of Softmax in MPC describing, “The exponential function is approximated using the limit approximation (Section II-A5). The exponential function can explode quickly even in plaintext, causing numerical overflow when some input values are large. To achieve numerical stability, softmax is practically implemented as
PNG
media_image11.png
74
274
media_image11.png
Greyscale
where xmax is the maximum value in the given vector. The subtraction of the maximum value does not change the final value of the softmax function, but it greatly improves the numerical stability since the largest input to the exponential function becomes 0.” Here, Wang establishes an exponential growth function within the softmax function given known in the art to be a part of attention mechanism calculation. The exponential growth function is being subtracted from another value within the transformation to modify the transformation here of the softmax function, which is being interpreted here as the subtracting of a function from a transformation.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base references of Xu, Son and Conway with the teachings of Wang by using Son’s teachings of a method and apparatus with attention based object analysis, Xu’s teachings of self-attention with functional time representation learning, and Conway’s teachings of a generative AI system using transformer-based models to perform various methods and incorporate with Wang’s teachings of Transformer models with secure multi-party computation using softmax computation.
One of ordinary skill in the art would be motivated to do so because by integrating Wang’s frameworks into the methods of Conway, Xu, and Son, which are all in relation to transformer models or networks, one of ordinary skill in the art would bring “the first in-depth study on MPC-based private inference of Transformer-based models.” (Wang, page 1 section I. Introduction) and “trade-off between performance improvements and the corresponding impact on model accuracy using detailed experiments.” (Wang, page 1 Abstract).
Claim 17:
Regarding claim 17, Son in view of Xu, and further in view of Conway teaches the limitations of claim 15.
Son does not appear to explicitly teach “The one or more non-transitory, computer-readable media of claim 15, wherein the function is an exponential growth function and wherein aggregating the transformation and the plurality of respective time differences comprises subtracting the function of the plurality of respective time differences from the transformation.”
Further, Xu teaches “…and wherein aggregating the transformation and the plurality of respective time differences comprises...”
See Xu on page 2 section 2 Related Work describing, “The original self-attention uses dot-product attention [20], defined via:
PNG
media_image1.png
40
220
media_image1.png
Greyscale
, where Q denotes the queries, K denotes the keys and V denotes the values (representations) of events in the sequence. Self-attention mechanism relies on the positional encoding to recognize and capture sequential information, where the vector representation for each position, which is shared across all sequences, is added or concatenated to the corresponding event embeddings. The above Q, K and V matrices are often linear (or identity) projections of the combined event-position representations. Attention patterns are detected through the inner products of query-key pairs, and propagate to the output as the weights for combining event values. Several variants of self-attention have been developed under different use cases including online recommendation [10], where sequence representations are often given by the attention-weighted sum of event embeddings.” Here, Xu teaches a self-attention mechanism which, as known in the art, generates a plurality of attention values by concatenating or aggregating, and establishes that each attention value here is given a weight as Q denotes queries, K denotes key events and V is the value of the events the function given demonstrates an attention value indicating a weight of a corresponding key event relative to a query event of a plurality of key events. Further, see Xu on pages 2-3 section 2 Related Work describing, “A recent work proposes a time-aware RNN with time encoding [12]. The functional time embeddings proposed in our work have sound theoretical justifications and interpretations. Also, by replacing positional encoding with time embedding we inherit the advantages of self-attention such as computation efficiency and model interpretability.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding. Further, see Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes the self-attention mechanism using positional encoding and replacing it with time embedding using dot products. You can also see that the function here is being added to the transformation as the time differences are being added to the transformation in the formula given.
Further, Xu teaches “…the plurality of respective time differences...”
See Xu on pages 2-3 section 2 Related Work describing, “Embedding time from an interval (suppose starting from origin) T = [0,tmax] to
PNG
media_image2.png
20
16
media_image2.png
Greyscale
is equivalent to finding a mapping
PNG
media_image3.png
20
78
media_image3.png
Greyscale
. Time embeddings can be added or concatenated to event embedding
PNG
media_image4.png
17
48
media_image4.png
Greyscale
, where Zi gives the vector representation of event ei, i =1,...,V for a total of V events. The intuition is that upon concatenation of the event and time representations, the dot product between two time-dependent events (e1,t1) and (e2,t2) becomes
PNG
media_image5.png
26
130
media_image5.png
Greyscale
=
PNG
media_image6.png
21
234
media_image6.png
Greyscale
represents relationship between events, we expect that
PNG
media_image7.png
21
80
media_image7.png
Greyscale
captures temporal patterns, specially those related with the temporal difference t1−t2 as we discussed before.” Here, Xu establishes a function here of a plurality of time differences are being added to the transformation in the formula given to its respective mapping.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base reference of Son with the teaching of Xu
by using Son’s teachings of a method and apparatus with attention based object analysis, and incorporate with Xu’s teachings of self-attention with functional time representation learning.
One of ordinary skill in the art would be motivated to do so because by integrating Xu’s frameworks into the method Son, which are both in relation to transformer models or networks, one of ordinary skill in the art would bring “translation-invariant time kernel which motivates several functional forms of time feature mapping justified from classic functional analysis theories, namely Bochner’s Theorem [13] and Mercer’s Theorem [15]. Compared with the other heuristic-driven time to vector methods” (Xu, page 2 section 1 Introduction), and “feasible time embeddings according to the time feature mappings such that they are properly parameterized and compatible with self-attention. We further discuss the interpretations of the proposed time embeddings and how to model their interactions with event representations under self-attention.” (Xu, page 2 section 1 Introduction).
Neither Son, Xu, or Conway appear to explicitly teach “The one or more non-transitory, computer-readable media of claim 15, wherein the function is an exponential growth function…comprises subtracting the function of…from the transformation.”
However Wang in an analogous system teaches, “The one or more non-transitory, computer-readable media of claim 15, wherein the function is an exponential growth function…comprises subtracting the function of…from the transformation.”
See Wang on Page 8 Section III subsection D. Analysis of Softmax in MPC describing, “The exponential function is approximated using the limit approximation (Section II-A5). The exponential function can explode quickly even in plaintext, causing numerical overflow when some input values are large. To achieve numerical stability, softmax is practically implemented as
PNG
media_image11.png
74
274
media_image11.png
Greyscale
where xmax is the maximum value in the given vector. The subtraction of the maximum value does not change the final value of the softmax function, but it greatly improves the numerical stability since the largest input to the exponential function becomes 0.” Here, Wang establishes an exponential growth function within the softmax function given known in the art to be a part of attention mechanism calculation. The exponential growth function is being subtracted from another value within the transformation to modify the transformation here of the softmax function, which is being interpreted here as the subtracting of a function from a transformation.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the base references of Xu, Son and Conway with the teachings of Wang by using Son’s teachings of a method and apparatus with attention based object analysis, Xu’s teachings of self-attention with functional time representation learning, and Conway’s teachings of a generative AI system using transformer-based models to perform various methods and incorporate with Wang’s teachings of Transformer models with secure multi-party computation using softmax computation.
One of ordinary skill in the art would be motivated to do so because by integrating Wang’s frameworks into the methods of Conway, Xu, and Son, which are all in relation to transformer models or networks, one of ordinary skill in the art would bring “the first in-depth study on MPC-based private inference of Transformer-based models.” (Wang, page 1 section I. Introduction) and “trade-off between performance improvements and the corresponding impact on model accuracy using detailed experiments.” (Wang, page 1 Abstract).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HASSAN R SESAY whose telephone number is (571)272-8493. The examiner can normally be reached Monday-Friday 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HASSAN RAMADAN SESAY/Examiner, Art Unit 2146
/USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146