DETAILED ACTION
1. Claims 1-15 are pending in this application.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. §102 and §103 (or as subject to pre-AIA 35 U.S.C. §102 and §103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Information Disclosure Statement
3. The information disclosure statement filed 08/14/2025 is in compliance with the provisions of 37 CFR 1.97, 1.98 and MPEP § 609. It has been placed in the application file and the information referred to therein has been considered as to the merits.
Claim Rejections - 35 USC § 101
4. 35 U.S.C. §101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 15 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter as follows.
Claim 15 define “… a recording medium …”. The specification in paragraphs [0090] and [0266] does not clearly define the recording medium or give a meaning for it. In view of the specification, the recording medium may include transitory and non-transitory propagation signals, whereas a transitory propagating signal is not a process, machine, manufacture, or composition of matter. The claim and/or specification fails to refrain the claimed tangible medium from propagating a signal. Therefore, the claim is “a signal per se” claim. A “signal per se” claim refer to a claim that covers a signal itself, rather than a device or process that uses the signal. A signal per se claim is non-statutory subject matter. See MPEP 2106.03(II) “In re Nuijten, 500 F.3d 1346, 84 USPQ2d 1495 (Fed. Cir. 2007). When the BRI encompasses transitory forms of signal transmission, a rejection under 35 U.S.C. 101 as failing to claim statutory subject matter would be appropriate.”
Any amendment to the claim should be commensurate with its corresponding disclosure.
Claims 1-15 are rejected under 35 U.S.C. §101 because the claimed invention is directed to an abstract idea (Mental Process) without significantly more. The claims similarly describe steps for providing similar content in a content streaming system.
The following is an analysis based on 2019 Revised Patent Subject Matter Eligibility Guidance (2019 PEG).
Step 1, Statutory Category?
Claims 1-13 are directed to a method.
Claim 14 is directed to a system.
Claim 20 is directed to a program stored in a recording medium.
Therefore, claims 1-14 fall into at least one of the four statutory categories.
Claim 15 is a signal per se claim and is not statutory or eligible for patenting, see the signal per se rejection above.
Step 2A, Prong I: Judicial Exception Recited?
The examiner submits that the foregoing claim limitations constitute a “Mental Process”, as the claims cover performance of the limitations in the human mind, given the broadest reasonable interpretation.
As per independent claims 1, 14 and 15, the claims similarly recite the limitations of:
“determining a first vector corresponding to the first sequence-type text data and a second vector corresponding to the second sequence-type text data using a language model learned based on synopsis information included in metadata of content items;” A Human can mentally observe data and use that information to create a vector. A human can also evaluate vector data using specific criteria. The machine learning model is a simple element used herein to implement the abstract idea. There is nothing so complex in the limitation that could not be doing in the human mind.
“determining a similarity between the first content item and the second content item using the first vector and the second vector;” A human can mentally observe vectors and mentally judge them to identify similarity between their information. There is nothing so complex in the limitation that could not be doing in the human mind.
As per dependent claim 7, the claim recites the limitation of:
“converting text metadata describing contents of the content items into the sequence-type text data;” A human can mentally observe an item's metadata and select words to build a sequence of the words observed. There is nothing so complex in the limitation that could not be doing in the human mind.
“masking a synopsis token located between tokens indicating the synopsis area among a plurality of tokens included in the sequence-type text data;” A human can observe information and mentally visualize parts of the observed information that are masked or hidden. There is nothing so complex in the limitation that could not be doing in the human mind.
As per dependent claim 8, the claim recites the limitation of:
“dividing the text metadata into a plurality of tokens;” A human can mentally observe metadata and visualize it divided into two or more parts. There is nothing so complex in the limitation that could not be doing in the human mind.
“generating the sequence-type text data by inserting at least one separator between the tokens” a human can mentally create a sequence and insert information as they visualize it. There is nothing so complex in the limitation that could not be doing in the human mind.
As per dependent claim 9, the claim recites the limitation of:
“selecting an independent token from among synopsis tokens located between tokens indicating the synopsis area;” A human can mentally select information and identify data about the selected information. There is nothing so complex in the limitation that could not be doing in the human mind.
“masking the selected independent token” A human can observe information and mentally visualize parts of the observed information that are masked or hidden. There is nothing so complex in the limitation that could not be doing in the human mind.
As per dependent claim 13, the claim recites the limitation of:
“determining a third vector corresponding to the third sequence-type text data using the learned language model;” A Human can mentally observe data and use that information to create a vector. A human can also evaluate vector data using specific criteria. The machine learning model is a simple element used herein to implement the abstract idea. There is nothing so complex in the limitation that could not be doing in the human mind.
“determining a similarity between the first content item and the third content item using the first vector and the third vector, wherein the providing the content list comprises:” A human can mentally observe vectors and mentally judge them to identify similarity between their information. There is nothing so complex in the limitation that could not be doing in the human mind.
“selecting the second content item from among the second content item and the third content item based on the similarity between the first content item and the second content item and the similarity between the first content item and the third content item.” A human can mentally select information and identify data about the selected information. There is nothing so complex in the limitation that could not be doing in the human mind.
Accordingly, claims 1-15 recite at least one abstract idea.
Step 2A, Prong II: Integrated into a Practical Application?
The claims recite the following additional limitations/elements:
As per independent claims 1, 14 and 15, the claims similarly recite the limitations/elements of:
The additional element of “a server” is a mere element (generic computer) used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere element to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
The additional limitation of “obtaining first sequence-type text data including information included in first metadata of a first content item;” is an example of an insignificant extra-solution activity of data gathering and can be understood as activities incidental to the primary process or product that are merely a nominal or tangential addition to the claim (see MPEP 2106.05(g)).
The additional limitation of “obtaining second sequence-type text data including information included in second metadata of a second content item;” is an example of an insignificant extra-solution activity of data gathering and can be understood as activities incidental to the primary process or product that are merely a nominal or tangential addition to the claim (see MPEP 2106.05(g)).
The additional limitation of “providing a content list including at least one content item including the second content item selected based on the similarity.” is an example of an insignificant extra-solution activity of data transmitting (In computer science, providing data is commonly known as transmitting data) and can be understood as activities incidental to the primary process or product that are merely a nominal or tangential addition to the claim (see MPEP 2106.05(g)).
As per dependent claim 2, the claim recites the limitation of:
The additional limitation/element “wherein the language model is learned through training to predict synopsis information of the content items based on a masked language model (MLM).” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 3, the claim recites the limitation of:
The additional limitation/element “wherein the language model is primarily learned through training to predict hashtag information of the content items based on the MLM and is secondarily learned through training to predict synopsis information of the content items based on the MLM.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 4, the claim recites the limitation of:
The additional limitation/element “wherein the language model is primarily learned through training to predict synopsis information of the content items based on the MLM and is secondarily learned through training to predict hashtag information of the content items based on the MLM.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 5, the claim recites the limitation of:
The additional limitation/element “wherein the language model is learned through training to predict a masked token located between tokens indicating a synopsis area among a plurality of tokens included in input sequence-type text data.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 6, the claim recites the limitation of:
The additional limitation/element “wherein tokens indicating the synopsis area includes at least one of a separator token for separating different types of features or a special token for different types of features other than the synopsis.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 7, the claim recites the limitation of:
The additional limitation/element “performing learning on the language model through training to predict the masked synopsis token, wherein the text metadata includes at least one of title, synopsis, genre, director, actor or hashtag information.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 8, the claim recites the limitation of:
The additional limitation/element “wherein the at least one separator further includes at least one of tokens indicating the synopsis area, a separator token for separating different types of features, or special tokens indicating an area of a specific type of feature.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 9, the claim recites the limitation of:
The additional limitation/element “wherein the independent token is a token that does not start with a specified symbol.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 10, the claim recites the limitation of:
The additional limitation/element “wherein the training is performed using a prediction model, and wherein the prediction model includes the language model that receives, as input, sequence-type text data including the masked synopsis token and outputs vector values corresponding to the sequence-type text data, and a masked language model (MLM) head layer configured to predict at least one input token corresponding to at least one vector value output from the language model.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 11, the claim recites the limitation of:
The additional limitation/element “wherein the determining the similarity between the first content item and the second content item comprises calculating a similarity between the first vector and the second vector using a cosine similarity algorithm, wherein each of the first vector and the second vector is obtained by performing average pooling for output vector values of a last hidden layer of the learned language model.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 12, the claim recites the limitation of:
The additional limitation/element “wherein each of the first vector and the second vector is determined by assigning a weight to a vector value corresponding to a position of a specified feature among the output vector values of the last hidden layer of the learned language model.” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per dependent claim 13, the claim recites the limitation of:
The additional limitation/element “obtaining third sequence-type text data including information included in third metadata of a third content item;” is a mere instruction/element used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere elements to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
As per independent claim 14, the claim recites the limitations/elements of:
The additional limitations/elements of “a communication unit configured to transmit and receive signals to and from at least one client device; and a processor electrically connected to the communication unit, wherein the processor is configured to:” is a mere element (generic computer) used to apply an exception. A recitation of the words "apply it" (or an equivalent) are mere element to implement an abstract idea or other exception on a computer. (See MPEP 2106.05(f)).
Therefore, claims 1-15 do not integrate the recited abstract ideas into a practical application.
Step 2B: Claim provides an Inventive Concept?
With respect to the limitations identified as insignificant extra-solution activity above the conclusions are carried over, and both the “obtaining …; and providing …;” is well-understood, routine, and conventional operations.
For support as being well-understood, routine, and conventional for “obtaining …; and providing …;” as noted by the courts is well understood routine and conventional, see MPEP 2106.05(d)(ii) “i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); … buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network);” and/or MPEP 2106.05(d)(ii) “iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93;”, and/or MPEP 2106.05(d)(II) “iii. Ultramercial, 772 F.3d at 716, 112 USPQ2d at 1755 (updating an activity log);”.
Looking at the limitations in combination and the claim as a whole does not change this conclusion and the claim is ineligible.
Therefore, the claims 1-15 are not patent eligible.
Claim Rejections - 35 USC § 103
5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. § 102 and § 103 (or as subject to pre-AIA 35 U.S.C. § 102 and § 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section § 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under pre-AIA 35 U.S.C. § 103(a) are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
6. Claims 1 and 14-15 are rejected under 35 U.S.C. § 103 as being unpatentable over Hropak et al. (US 20220222470 A1) in view of Chao (US 20200159773 A1).
As per claim 1, Hropak teaches a method (i.e. “a method 400 “; fig.4, para. [0077]) of operating a server (i.e. “server(s) 116 may be operated”; fig. 1, para. [0051])
in a content streaming system (i.e. “a content streaming system 800”; fig.8, para. [0106]), the method comprising:
obtaining first sequence-type text data including information included in first metadata of a first content item (i.e. “identify a content item(s) in the image data.”; fig.4, para. [0078]. Further, i.e. “the metadata includes an indication of one or more frames of the video 156 and content item information 152”; para. [0037]. Further, i.e. “there may be any of a number of content item information 152 related to the content items 150, which may take any of a variety of different forms. Examples include overlays, annotations, arrows, color changes, highlighted appearance, text, and/or any other indication or information related to the content items 150.”; para. [0042]; Examiner note: using a BRI the first sequence-type text data is interpreted as the number of content item information);
obtaining second sequence-type text data including information included in second metadata of a second content item (i.e. “A second MLM and/or algorithm may then be used to identify a particular content item(s) from the detected objects.”; para. [0021]; Examiner note: using a BRI the second sequence-type text data including information included in second metadata of a second content item is interpreted as the particular content item(s));
However, it is noted that the prior art of Hropak do not explicitly teach “determining a first vector corresponding to the first sequence-type text data and a second vector corresponding to the second sequence-type text data using a language model learned based on synopsis information included in metadata of content items; determining a similarity between the first content item and the second content item using the first vector and the second vector; and providing a content list including at least one content item including the second content item selected based on the similarity.”
On the other hand, in the same field of endeavor, Chao teaches determining a first vector corresponding to the first sequence-type text data and a second vector corresponding to the second sequence-type text data (i.e. “generating at least one vector, each vector corresponding to one text component”; para. [0012])
using a language model learned (i.e. “using the machine learning process”; para. [0108])
based on synopsis information included in metadata of content items (i.e. “the natural language description may be one of: a long synopsis of the content item; a short overview of the content item; an abbreviated summary of the content item;”; para. [0041]. Further, i.e. “The content metadata store 4 stores metadata for a collection of content.”; fig.1, para. [0057]);
determining a similarity between the first content item and the second content item using the first vector and the second vector (i.e. “computing a score for each possible match based on vector similarities;”; para. [0032]. Further, i.e. “determining associations within vectors by identifying common text components which are text components which are represented in the same vector, wherein an indicated association strength depends on the frequency with which the text components occur in the same natural language description;”; para. [0034]); and
providing a content list including at least one content item including the second content item selected based on the similarity (i.e. “comparing the search vector to each metadata vector derived from the training inputs to generate a list of possible matches and identifying at least one match to be presented to a user”; para. [0043]; Examiner note: using a BRI the providing a content list is interpreted as the generate a list of possible matches. Using a BRI the second content item selected is interpreted as the identifying at least one match).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Chao that teaches methods and systems for content access and storage into the prior art of Hropak that teaches automatic content recognition and information in live streaming suitable for video games. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to operate a process that successfully identifies this new metadata as belonging to the existing record because it can improve the breadth of information held about the episode in this existing metadata (Chao, para. [0075]).
As per claim 14, Hropak teaches a server in a content streaming system (i.e. “a content streaming system 800, in accordance with some embodiments of the present disclosure. FIG. 8 includes application server(s)”; fig.8, para. [0106]), the server comprising:
a communication unit configured to transmit and receive signals to and from at least one client device (i.e. “a communication interface 110C to transmit streaming data to the identification server(s) 116 and/or client device(s) 104”; para. [0034]); and
a processor electrically connected to the communication unit (i.e. “processor(s), cause the processor(s) to,…, in response receive a video stream from the video server(s) 130 and/or the identification server(s) 116 using the communication interface 110A,”; para. [0037]),
wherein the processor is configured to:
obtain first sequence-type text data including information included in first metadata of a first content item (i.e. “identify a content item(s) in the image data.”; fig.4, para. [0078]. Further, i.e. “the metadata includes an indication of one or more frames of the video 156 and content item information 152”; para. [0037]. Further, i.e. “there may be any of a number of content item information 152 related to the content items 150, which may take any of a variety of different forms. Examples include overlays, annotations, arrows, color changes, highlighted appearance, text, and/or any other indication or information related to the content items 150.”; para. [0042]; Examiner note: using a BRI the first sequence-type text data is interpreted as the number of content item information);
obtain second sequence-type text data including information included in second metadata of a second content item (i.e. “A second MLM and/or algorithm may then be used to identify a particular content item(s) from the detected objects.”; para. [0021]; Examiner note: using a BRI the second sequence-type text data including information included in second metadata of a second content item is interpreted as the particular content item(s));
However, it is noted that the prior art of Hropak do not explicitly teach “determine a first vector corresponding to the first sequence-type text data and a second vector corresponding to the second sequence-type text data using a language model learned based on synopsis information included in metadata of content items; determine a similarity between the first content item and the second content item using the first vector and the second vector; and provide a content list including at least one content item including the second content item selected based on the similarity.”
On the other hand, in the same field of endeavor, Chao teaches determine a first vector corresponding to the first sequence-type text data and a second vector corresponding to the second sequence-type text data (i.e. “generating at least one vector, each vector corresponding to one text component”; para. [0012])
using a language model learned (i.e. “using the machine learning process”; para. [0108])
based on synopsis information included in metadata of content items (i.e. “the natural language description may be one of: a long synopsis of the content item; a short overview of the content item; an abbreviated summary of the content item;”; para. [0041]. Further, i.e. “The content metadata store 4 stores metadata for a collection of content.”; fig.1, para. [0057]);
determine a similarity between the first content item and the second content item using the first vector and the second vector (i.e. “computing a score for each possible match based on vector similarities;”; para. [0032]. Further, i.e. “determining associations within vectors by identifying common text components which are text components which are represented in the same vector, wherein an indicated association strength depends on the frequency with which the text components occur in the same natural language description;”; para. [0034]); and
provide a content list including at least one content item including the second content item selected based on the similarity (i.e. “comparing the search vector to each metadata vector derived from the training inputs to generate a list of possible matches and identifying at least one match to be presented to a user”; para. [0043]; Examiner note: using a BRI the providing a content list is interpreted as the generate a list of possible matches. Using a BRI the second content item selected is interpreted as the identifying at least one match).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Chao that teaches methods and systems for content access and storage into the prior art of Hropak that teaches automatic content recognition and information in live streaming suitable for video games. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to operate a process that successfully identifies this new metadata as belonging to the existing record because it can improve the breadth of information held about the episode in this existing metadata (Chao, para. [0075]).
As per claim 15, Hropak teaches a program stored in a recording medium to execute the method according to claim 1 when operated by a processor (i.e. “method 500, … various functions may be carried out by a processor executing instructions stored in memory. The method may also be embodied as computer-usable instructions stored on computer storage media.”; para. [0080]; Examiner note: using a BRI the program stored in a recording medium is interpreted as the instructions stored in memory).
7. Claims 2 and 5-10 are rejected under 35 U.S.C. § 103 as being unpatentable over Hropak et al. (US 20220222470 A1) in view of Chao (US 20200159773 A1) in further view of Malon (US 20220327586 A1).
As per claim 2, Hropak and Chao teach all the limitations as discussed in claim 1 above.
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “wherein the language model is learned through training to predict synopsis information of the content items based on a masked language model (MLM).”
On the other hand, in the same field of endeavor, Malon teaches wherein the language model is learned through training to predict synopsis information of the content items based on a masked language model (MLM) (i.e. “by using a claim generation module, one can pinpoint opinions about key phrases beyond just “positive” and “negative” sentiment. A pretrained transformer model for masked language modeling, such as Bidirectional Encoder Representations from Transformers (BERT) or XLNet, can be fine-tuned to become the claim generation model, where the fine-tuned claim generation model predicts the sequence of words for the claim.”; para. [0021]. Further, i.e. “an opinion summarization tool 100 can count opinions about a product and provide summaries of the opinions.”; para. [0024]; Examiner note: using a BRI the synopsis information of the content items is interpreted as the summaries of the opinions).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews into the combination of prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to perform a frequency analysis on an input list of single-item product reviews and a broader category corpus. This analysis identifies frequent phrases to facilitate fine-tuning a pretrained transformer model, ultimately producing a trained claim-generator model (Malon, para. [0007]).
As per claim 5, Hropak and Chao teach all the limitations as discussed in claim 1 above.
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “wherein the language model is learned through training to predict a masked token located between tokens indicating a synopsis area among a plurality of tokens included in input sequence-type text data.”
On the other hand, in the same field of endeavor, Malon teaches wherein the language model is learned through training to predict a masked token located between tokens indicating a synopsis area among a plurality of tokens included in input sequence-type text data (i.e. “A pretrained transformer model for masked language modeling, such as Bidirectional Encoder Representations from Transformers (BERT) or XLNet, can be fine-tuned to become the claim generation model, where the fine-tuned claim generation model predicts the sequence of words for the claim. The vocabulary for such a model includes tokens for common words and pieces of words, and several special tokens used in training.”; para. [0021]. Further, i.e. “At block 250, one masked word in the template can be predicted using a refined claim generator, T1, applied to the input reviews”; para. [0045]. Further, i.e. “At block 240, for each key word or phrase, k, a summary template, f, is taken, with a slot for the key word or phrase, and zero or one masked tokens to be filled in.”; para. [0043]; Examiner not: using a BRI the synopsis rea is interpreted as the summary template ,f, is taken, with a slot. The input sequence-type text data is the input review).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews into the combination of prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to perform a frequency analysis on an input list of single-item product reviews and a broader category corpus. This analysis identifies frequent phrases to facilitate fine-tuning a pretrained transformer model, ultimately producing a trained claim-generator model (Malon, para. [0007]).
As per claim 6, Hropak and Chao teach all the limitations as discussed in claim 5 above.
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “wherein tokens indicating the synopsis area includes at least one of a separator token for separating different types of features or a special token for different types of features other than the synopsis.”
On the other hand, in the same field of endeavor, Malon teaches wherein tokens indicating the synopsis area includes at least one of a separator token for separating different types of features or a special token for different types of features other than the synopsis (i.e. “A sequence of tokens, x, can be constructed by concatenating the classification token, each of the sentences, r(1); . . . ; r(n), the separator token, the template output, fj, and another copy of the separator token. The separator token is a special token that tells the model that different sentences have different purposes. The separator token also keeps sentences from being concatenated or mixed up.”; para. [0044]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews into the combination of prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to perform a frequency analysis on an input list of single-item product reviews and a broader category corpus. This analysis identifies frequent phrases to facilitate fine-tuning a pretrained transformer model, ultimately producing a trained claim-generator model (Malon, para. [0007]).
As per claim 7, Hropak, Chao and Malon teach all the limitations as discussed in claim 5 above.
Additionally, Hropak teaches wherein the text metadata includes at least one of title, synopsis, genre, director, actor or hashtag information (i.e. “The MLM(s) 122 may be trained to detect objects in certain contexts (e.g., trained to detect objects within a single game title) or may be trained to detect objects over any of a number of video contexts (e.g., multiple game titles, versions, expansions, DLC, genres, game systems, etc.).”; para. [0057]).
However, it is noted that the combination of the prior arts of Hropak and Malon do not explicitly teach “converting text metadata describing contents of the content items into the sequence-type text data; masking a synopsis token located between tokens indicating the synopsis area among a plurality of tokens included in the sequence-type text data; and performing learning on the language model through training to predict the masked synopsis token;”
On the other hand, in the same field of endeavor, Chao teaches converting text metadata describing contents of the content items into the sequence-type text data (i.e. “each natural language description for each item is converted into classified text components.”; para. [0178]);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Chao that teaches methods and systems for content access and storage into the prior art of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to operate a process that successfully identifies this new metadata as belonging to the existing record because it can improve the breadth of information held about the episode in this existing metadata (Chao, para. [0075]).
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “masking a synopsis token located between tokens indicating the synopsis area among a plurality of tokens included in the sequence-type text data; and performing learning on the language model through training to predict the masked synopsis token;”
On the other hand, in the same field of endeavor, Malon teaches masking a synopsis token located between tokens indicating the synopsis area among a plurality of tokens included in the sequence-type text data (i.e. “Each summary template quotes the key word or phrase, ki, or a word or phrase derived from ki, and has zero or one positions which are masked and are to be filled in with a word.”; para. [0073]); and
performing learning on the language model through training to predict the masked synopsis token (i.e. “a pretrained transformer model, T, that has been pretrained for masked language modeling, such as BERT, can be fine-tuned. The vocabulary for such a model includes tokens for common words and pieces of words, and several special tokens used in training, including a classification token and a separator token.”; para. [0054]);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews into the combination of prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to perform a frequency analysis on an input list of single-item product reviews and a broader category corpus. This analysis identifies frequent phrases to facilitate fine-tuning a pretrained transformer model, ultimately producing a trained claim-generator model (Malon, para. [0007]).
As per claim 8, Hropak, Chao and Malon teach all the limitations as discussed in claim 7 above.
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “wherein the converting the text metadata into the sequence-type text data comprises: dividing the text metadata into a plurality of tokens; and generating the sequence-type text data by inserting at least one separator between the tokens; wherein the at least one separator further includes at least one of tokens indicating the synopsis area, a separator token for separating different types of features, or special tokens indicating an area of a specific type of feature.”
On the other hand, in the same field of endeavor, Malon teaches wherein the converting the text metadata into the sequence-type text data comprises: dividing the text metadata into a plurality of tokens (i.e. “A sequence of tokens, x, can be constructed”; para. [0044]); and
generating the sequence-type text data by inserting at least one separator between the tokens (i.e. “The separator token is a special token that tells the model that different sentences have different purposes. The separator token also keeps sentences from being concatenated or mixed up.”; para. [0062]),
wherein the at least one separator further includes at least one of tokens indicating the synopsis area, a separator token for separating different types of features, or special tokens indicating an area of a specific type of feature (i.e. “The separator token is a special token that tells the model that different sentences have different purposes.”; para. [0062]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews into the combination of prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to perform a frequency analysis on an input list of single-item product reviews and a broader category corpus. This analysis identifies frequent phrases to facilitate fine-tuning a pretrained transformer model, ultimately producing a trained claim-generator model (Malon, para. [0007]).
As per claim 9, Hropak, Chao and Malon teach all the limitations as discussed in claim 7 above.
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “wherein the masking the synopsis token comprises: selecting an independent token from among synopsis tokens located between tokens indicating the synopsis area; and masking the selected independent token, wherein the independent token is a token that does not start with a specified symbol.”
On the other hand, in the same field of endeavor, Malon teaches wherein the masking the synopsis token comprises: selecting an independent token from among synopsis tokens located between tokens indicating the synopsis area (i.e. “The key word or phrase, ki, can be selected by a user”; para. [0038]); and
masking the selected independent token (i.e. “If ki is a noun or noun phrase, the template is “The ki is [MASK].” If ki is an adjective, the templates are “It is ki.” and “The [MASK] is ki.””; para. [0041]),
wherein the independent token is a token that does not start with a specified symbol (i.e. “special tokens used in training”; para. [0021]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews into the combination of prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to perform a frequency analysis on an input list of single-item product reviews and a broader category corpus. This analysis identifies frequent phrases to facilitate fine-tuning a pretrained transformer model, ultimately producing a trained claim-generator model (Malon, para. [0007]).
As per claim 10, Hropak, Chao and Malon teach all the limitations as discussed in claim 7 above.
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “wherein the training is performed using a prediction model, and wherein the prediction model includes the language model that receives, as input, sequence-type text data including the masked synopsis token and outputs vector values corresponding to the sequence-type text data, and a masked language model (MLM) head layer configured to predict at least one input token corresponding to at least one vector value output from the language model.”
On the other hand, in the same field of endeavor, Malon teaches wherein the training is performed using a prediction model, and wherein the prediction model includes the language model that receives, as input, sequence-type text data including the masked synopsis token and outputs vector values corresponding to the sequence-type text data, and a masked language model (MLM) head layer configured to predict at least one input token corresponding to at least one vector value output from the language model (i.e. “The trained neural network claim generator model 750 can be a neural network configured to generate claims utilizing one or more templates. A claim generated by the claim generator model 750 that summarizes the input claims can be output 140.”; para. [0104]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews into the combination of prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to perform a frequency analysis on an input list of single-item product reviews and a broader category corpus. This analysis identifies frequent phrases to facilitate fine-tuning a pretrained transformer model, ultimately producing a trained claim-generator model (Malon, para. [0007]).
8. Claims 3-4 are rejected under 35 U.S.C. § 103 as being unpatentable over Hropak et al. (US 20220222470 A1) in view of Chao (US 20200159773 A1) in further view of Malon (US 20220327586 A1) still in further view of Maurer et al. (US 20240176960 A1).
As per claim 3, Hropak, Chao and Malon teach all the limitations as discussed in claim 2 above.
However, it is noted that the combination the prior arts of Hropak, Chao and Malon do not explicitly teach “wherein the language model is primarily learned through training to predict hashtag information of the content items based on the MLM and is secondarily learned through training to predict synopsis information of the content items based on the MLM.”
On the other hand, in the same field of endeavor, Maurer teaches wherein the language model is primarily learned through training to predict hashtag information of the content items based on the MLM and is secondarily learned through training to predict synopsis information of the content items based on the MLM (i.e. “a ML model(s) may be configured to analyze channel contextual data based in part on channel data associated with the channel that the teleconferencing meeting was initiated from (e.g., as shown in FIG. 6 , a meeting was initiated from the #Team-native-ai channel). In some examples, the channel contextual data may be input into a ML model configured to output a summary of the teleconferencing meeting.”; para. [0159]. Further, i.e. “The summarization engine 120 may utilize a machine-learning (ML) model(s) 142 (or MLM) that accepts inputs, and using the inputs, outputs such a summary document(s).”; para. [0037]; Examiner note: using a BRI the hashtag information is interpreted as the #Team-native-ai channel. Using a BRI the synopsis information is interpreted as the summary).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Maurer that teaches techniques for transcribing and/or summarizing multimedia collaboration sessions into the combination of the prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, Chao that teaches methods and systems for content access and storage, and Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to provide transcripts of the ad hoc discussions because they can facilitate work-related communications (Maurer, para. [0002]).
As per claim 4, Hropak, Chao and Malon teach all the limitations as discussed in claim 2 above.
However, it is noted that the combination the prior arts of Hropak, Chao and Malon do not explicitly teach “wherein the language model is primarily learned through training to predict synopsis information of the content items based on the MLM and is secondarily learned through training to predict hashtag information of the content items based on the MLM.”
On the other hand, in the same field of endeavor, Maurer teaches wherein the language model is primarily learned through training to predict synopsis information of the content items based on the MLM and is secondarily learned through training to predict hashtag information of the content items based on the MLM (i.e. “a ML model(s) may be configured to analyze channel contextual data based in part on channel data associated with the channel that the teleconferencing meeting was initiated from (e.g., as shown in FIG. 6 , a meeting was initiated from the #Team-native-ai channel). In some examples, the channel contextual data may be input into a ML model configured to output a summary of the teleconferencing meeting.”; para. [0159]. Further, i.e. “The summarization engine 120 may utilize a machine-learning (ML) model(s) 142 (or MLM) that accepts inputs, and using the inputs, outputs such a summary document(s).”; para. [0037]; Examiner note: using a BRI the hashtag information is interpreted as the #Team-native-ai channel. Using a BRI the synopsis information is interpreted as the summary).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Maurer that teaches techniques for transcribing and/or summarizing multimedia collaboration sessions into the combination of the prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, Chao that teaches methods and systems for content access and storage, and Malon that teaches extracting and counting frequent opinions within a corpus of customer reviews. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to provide transcripts of the ad hoc discussions because they can facilitate work-related communications (Maurer, para. [0002]).
9. Claims 11-13 are rejected under 35 U.S.C. § 103 as being unpatentable over Hropak et al. (US 20220222470 A1) in view of Chao (US 20200159773 A1) in further view of Ma (US 20240220556 A1).
As per claim 11, Hropak and Chao teach all the limitations as discussed in claim 1 above.
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “wherein the determining the similarity between the first content item and the second content item comprises calculating a similarity between the first vector and the second vector using a cosine similarity algorithm, wherein each of the first vector and the second vector is obtained by performing average pooling for output vector values of a last hidden layer of the learned language model.”
On the other hand, in the same field of endeavor, Ma teaches wherein the determining the similarity between the first content item and the second content item comprises calculating a similarity between the first vector and the second vector using a cosine similarity algorithm (i.e. “one or more operations (e.g., mathematical operations) may be performed using the first vector representation 620 associated with the topic “Sports” and the second vector representation 656 associated with the topic “Politics” to determine the first similarity score (e.g., the first similarity score may be based upon (and/or may be equal to) a measure of similarity between the first vector representation 620 and the second vector representation 656, such as a cosine similarity between the first vector representation 620 and the second vector representation 656).”; para. [0109]),
wherein each of the first vector and the second vector is obtained by performing average pooling for output vector values of a last hidden layer of the learned language model (i.e. “Values of the inferred activity distribution array may be automatically pruned (e.g., via TopK pooling, max pooling, average pooling, etc.) to generate a filtered subset of activity distribution values, which may be included in a set of features used to train a machine learning model configured to select content.”; para. [0047]. Further, i.e. “First vector representations of the first entities and second vector representations of the second entities may be processed (e.g., using an attention model) to generate an attention distribution array.”; para. [0045]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Ma that teaches provide platforms for viewing media into the combination of the prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to improve the accuracy of the filtered subset of activity distribution values and/or the machine learning model trained, thereby providing for more accurate selection of content by the machine learning model (Ma, para. [0048]).
As per claim 12, Hropak, Chao and Ma teach all the limitations as discussed in claim 11 above.
However, it is noted that the combination the prior arts of Hropak and Ma do not explicitly teach “wherein each of the first vector and the second vector is determined by assigning a weight to a vector value corresponding to a position of a specified feature among the output vector values of the last hidden layer of the learned language model.”
On the other hand, in the same field of endeavor, Chao teaches wherein each of the first vector and the second vector is determined by assigning a weight to a vector value corresponding to a position of a specified feature among the output vector values of the last hidden layer of the learned language model (i.e. “the outputs are further weighted (ordered) according to the weights given to each vectorisation.”; para. [0098]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Chao that teaches methods and systems for content access and storage into the combination of the prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Ma that teaches provide platforms for viewing media. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to operate a process that successfully identifies this new metadata as belonging to the existing record because it can improve the breadth of information held about the episode in this existing metadata (Chao, para. [0075]).
As per claim 13, Hropak and Chao teach all the limitations as discussed in claim 11 above.
Additionally, Hropak teaches obtaining third sequence-type text data including information included in third metadata of a third content item (i.e. “identify a content item(s) in the image data.”; fig.4, para. [0078]. Further, i.e. “the metadata includes an indication of one or more frames of the video 156 and content item information 152”; para. [0037]. Further, i.e. “there may be any of a number of content item information 152 related to the content items 150, which may take any of a variety of different forms. Examples include overlays, annotations, arrows, color changes, highlighted appearance, text, and/or any other indication or information related to the content items 150.”; para. [0042]; Examiner note: using a BRI the first sequence-type text data is interpreted as the number of content item information);
However, it is noted that the combination the prior arts of Hropak and Ma do not explicitly teach “determining a third vector corresponding to the third sequence-type text data using the learned language model; and determining a similarity between the first content item and the third content item using the first vector and the third vector, wherein the providing the content list comprises: selecting the second content item from among the second content item and the third content item based on the similarity between the first content item and the second content item and the similarity between the first content item and the third content item.”
On the other hand, in the same field of endeavor, Chao teaches determining a third vector corresponding to the third sequence-type text data using the learned language model (i.e. “generating matching metadata vectors for identifying content items in a store searchable by input vectors, …vectorising the natural language description into a vector representing a set of classified text components and assigning a weight to each text component of the vector based on the frequency of occurrence of the text component of other vectors derived from training inputs having the same content identifier;”; para. [0034]); and
determining a similarity between the first content item and the third content item using the first vector and the third vector, wherein the providing the content list comprises (i.e. “computing a score for each possible match based on vector similarities;”; para. [0032]. Further, i.e. “determining associations within vectors by identifying common text components which are text components which are represented in the same vector, wherein an indicated association strength depends on the frequency with which the text components occur in the same natural language description”; para. [0034]):
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Chao that teaches methods and systems for content access and storage into the combination of the prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Ma that teaches provide platforms for viewing media. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to operate a process that successfully identifies this new metadata as belonging to the existing record because it can improve the breadth of information held about the episode in this existing metadata (Chao, para. [0075]).
However, it is noted that the combination the prior arts of Hropak and Chao do not explicitly teach “selecting the second content item from among the second content item and the third content item based on the similarity between the first content item and the second content item and the similarity between the first content item and the third content item.”
On the other hand, in the same field of endeavor, Ma teaches selecting the second content item from among the second content item and the third content item based on the similarity between the first content item and the second content item and the similarity between the first content item and the third content item (i.e. “determine a plurality of content item scores associated with a second plurality of content items (e.g., news articles, informational articles, videos, advertisements, images, links, etc.), and use the plurality of content item scores to select the one or more content items from the plurality of content items.”; para. [0137]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Ma that teaches provide platforms for viewing media into the combination of the prior arts of Hropak that teaches automatic content recognition and information in live streaming suitable for video games, and Chao that teaches methods and systems for content access and storage. Additionally, this can operate as a facilitator for retrieving and applying video data to an appropriate MLM(s).
The motivation for doing so would be to improve the accuracy of the filtered subset of activity distribution values and/or the machine learning model trained, thereby providing for more accurate selection of content by the machine learning model (Ma, para. [0048]).
Prior Art of Record
10. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Kaul et al. (US 12596740 B2), teaches a categorization method.
Margolin et al. (US 11544460 B1), teaches anonymizing content suggestive of a particular characteristic while preserving relevant content.
Walczak et al. (US 11321580 B1), teaches learning item types of items listed in an electronic repository, and for training a machine learning model to predict the item type of a given input item.
Ishida et al. (US 20180336190 A1), teaches an information processing device acts as a neural network based on time-series data.
Conclusion
11. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANTONIO CAIA DO whose telephone number is (469)295-9251. The examiner can normally be reached on Monday - Friday / 06:30 to 16:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ng, Amy can be reached on (571) 270-1698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANTONIO J CAIA DO/
Examiner, Art Unit 2164