DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action has been issued in response to Applicant’s Communication of application S/N 19/040,453 filed on January 29, 2025. Claims 1 to 20 are currently pending with the application.
Priority
The instant application claims priority from provisional Application No. 63/626,752, filed on January 30, 2024. Applicant’s claim for the benefit of the prior-filed application under 35 U.S.C. 119(e), 120, 121, or 365(c), or 386(c) is acknowledged.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1 to 5, 7, 8, and 15 to 18 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1 to 3, 5, 7, and 10 of copending Application No. 19/040,448. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims in the instant application are anticipated by the claims in the copending Application. This is a provisional nonstatutory double patenting rejection.
Following mapping of claims 1 to 5, 7, and 8 of Instant Application to claims 1 to 3, 5, 7, and 10 of the copending application. Similar mapping applies to claims 15 to 18 of instant application, since they recite similar limitations.
Instant Application
Application No. 19/040,448
1. A method of automated generation of contextually-relevant images for a music segment, the method comprising: receiving at least one of basic metadata information and lyric information for the music segment; generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, the first prompt including a first request for context information based on the at least one of the basic metadata information and the lyric information; receiving the context information from the computer-implemented machine- learning language model in response to the first prompt; generating a second prompt for the computer-implemented machine-learning language model based on the context information, the second prompt including a second request to generate a third prompt for a computer- implemented machine-learning image-generation model including a third request to generate an image descriptive of the music segment; generating the third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model; and generating the image descriptive of the music segment by providing the third prompt as an input to the computer-implemented machine-learning image generation model.
8. The method of claim 1, and further comprising modifying data of an image database to store the image and to retrievably associate the image with an identifier for the music segment.
1. A method of automated generation of descriptive tags for a music segment, the method comprising: receiving at least one of basic metadata information and lyric information for the music segment; generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, the first prompt including a first request for first context information based on the at least one of the basic metadata information and the lyric information; generating the first context information by providing the first prompt as an input to the computer-implemented machine-learning language model; generating a second prompt for the computer-implemented machine-learning language model based on the first context information, the second prompt including a second request to generate a plurality of tags based on the first context information; generating the plurality of tags by providing the second prompt as an input to the computer-implemented machine-learning language model; and modifying electronic data of a queryable electronic database to retrievably associate the plurality of tags with the music segment.
2. The method of claim 1, wherein the generating the second prompt comprises generating the second prompt based on the context information and the at least one of the basic metadata information and the lyric information.
5. The method of claim 4, wherein the second prompt also includes the at least one of the basic metadata information and the lyric information.
3. The method of claim 1, wherein the context information comprises at least one of artist context information and historical context information.
10. The method of claim 9, wherein the first context information is historical context information and the second context information is artist context information.
4. The method of claim 1, and further comprising: generating a first database query based on the at least one of the basic metadata information and the lyric information; querying a first database with the first database query; and receiving first database data from the first database in response to the first database query; wherein generating the first prompt comprises generating the first prompt based on the first database data and the at least one of basic metadata information and lyric information.
2. The method of claim 1, and further comprising: generating a database query based on the at least one of the basic metadata information and the lyric information; querying a first database with the database query; and receiving database data from the first database in response to the database query; wherein generating the first prompt comprises generating the first prompt based on the database data and the at least one of basic metadata information and lyric information.
5. The method of claim 4, and further comprising: generating a second database query based on the context information; querying the first database with the second database query; and receiving second database data from the first database in response to the second database query; wherein generating the second prompt comprises generating the first prompt based on the second database data and context information.
3. The method of claim 1, and further comprising: generating a first database query based on the first context information; querying a first database with the database query; and receiving first database data from the first database in response to the database query; wherein generating the second prompt comprises generating the second prompt based on the first database data and the first context information.
7. The method of claim 1, wherein the context information is artist context information, and further comprising: generating a fourth prompt for the computer-implemented machine-learning language model based on the at least one of the basic metadata information 2 and the lyric information, the fourth prompt including a fourth request for historical context information based on the at least one of the basic metadata information and the lyric information; and receiving the historical context information from the computer-implemented machine-learning language model in response to the first prompt; wherein generating the second prompt comprises generating the second prompt based on the artist context information and the historical context information.
7. The method of claim 6, and further comprising: generating a third prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, the first prompt including a third request for second context information based on the at least one of the basic metadata information and the lyric information; and2 generating the second context information by providing the third prompt as an input to the computer-implemented machine-learning language model; wherein the second prompt is further based on the second context information and the second request is to generate the plurality of tags based on the first context information, the first database data, and the second context information.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 to 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 1, 15, and 16 recite generating a first and a second prompt.
The limitation of generating a first prompt, which specifically recites “generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, the first prompt including a first request for context information based on the at least one of the basic metadata information and the lyric information”, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind, but for the recitation of generic computer components. That is, other than reciting “by a processor” (in claim 16), nothing in the claim element precludes the steps from practically being performed in a human mind. For example, but for the “by a processor” language, “generating”, in the context of this claim encompasses the user mentally, with the aid of pen and paper, writing down a question or prompt requesting context information based on either metadata
or lyric information.
The limitation of generating a second prompt, which specifically recites “generating a second prompt for the computer-implemented machine-learning language model based on the context information, the second prompt including a second request to generate a third prompt for a computer-implemented machine-learning image-generation model including a third request to generate an image descriptive of the music segment”, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind, but for the recitation of generic computer components. That is, other than reciting “by a processor”, nothing in the claim element precludes the steps from practically being performed in a human mind. For example, but for the “by a processor” language, “generating”, in the context of this claim encompasses the user mentally, with the aid of pen and paper, writing down a second question or prompt that includes a request to generate a third prompt. If a claim limitation, under its broadest reasonable interpretation, covers mental processes but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claims recite the additional elements – “receiving at least one of basic metadata information and lyric information for the music segment”, “receiving the context information from the computer-implemented machine-learning language model in response to the first prompt”, “generating the third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model”, “generating the image descriptive of the music segment by providing the third prompt as an input to the computer-implemented machine-learning image generation model”, a processor, and at least one memory. The limitations “receiving at least one of basic metadata information and lyric information for the music segment”, “receiving the context information from the computer-implemented machine-learning language model in response to the first prompt”, “providing the second prompt as an input to the computer-implemented machine-learning language model”, and “providing the third prompt as an input to the computer-implemented machine-learning image generation model” amount to data-gathering steps which is considered to be insignificant extra-solution activity (See MPEP 2106.05(g)).
Continuing the analysis of the additional limitations, the limitations “generating the third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model” and “generating the image descriptive of the music segment by providing the third prompt as an input to the computer-implemented machine-learning image generation model” are recited at a high-level of generality, with no restriction on how the result is accomplished and no description of the mechanism for accomplishing the result, and is equivalent to merely saying “applying it”. Additionally, as further described above, the “providing” limitations amount to data gathering. The processor, and at least one memory in these steps are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The insignificant extra-solution activity identified above, which include the data gathering steps, is recognized by the courts as well-understood, routine, and conventional activity when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity (See MPEP 2106.05(d)(II)(i) Receiving or transmitting data over a network, e.g., using the Internet to gather data, buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)). The
claims are not patent eligible.
Claim 2 is dependent on claim 1 and includes all the limitations of claim 1. Therefore, claim 2 recites the same abstract idea of claim 1. The claim recites the additional limitations of “generating the second prompt based on the context information and the at least one of the basic metadata information and the lyric information”, which can be performed in the human mind with the aid of pen and paper, and therefore is further elaborating on the abstract idea. The claim does not amount to significantly more. Same rationale applies to claims 12, 13, 17, and 20.
Claim 3 is dependent on claim 1 and includes all the limitations of claim 1. Therefore, claim 3 recites the same abstract idea of claim 1. The claim recites the additional limitations of “the context information comprises at least one of artist context information and historical context information”, which is tying the abstract idea to a field of use by further specifying the target data, and which is simply an attempt to limit the application of the abstract idea to a particular technological environment; merely indicating a field of use or technological environment in which to apply the judicial exception does not meaningfully limit the claim (See MPEP 2106.05(h)). Same rationale applies to claim 18.
Claim 4 is dependent on claim 1 and includes all the limitations of claim 1. Therefore, claim 4 recites the same abstract idea of claim 1. The claim recites the additional limitations of “generating a first database query based on the at least one of the basic metadata information and the lyric information; querying a first database with the first database query; and receiving first database data from the first database in response to the first database query; wherein generating the first prompt comprises generating the first prompt based on the first database data and the at least one of basic metadata information and lyric information”. The generating limitations can be performed in the human mind with the aid of pen and paper, and therefore are further elaborating on the abstract idea. The querying is recited at a high-level of generality, with no restriction on how the result is accomplished and no description of the mechanism for accomplishing the result, and is equivalent to merely saying “applying it”. The receiving limitation amounts to data gathering steps, which is considered to be insignificant extra-solution activity (See MPEP 2106.05(g)), and recognized by the courts as well-understood, routine, and conventional activities when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity (See MPEP 2106.05(d) (II)(i) Receiving or transmitting data over a network, e.g., using the Internet to gather data, buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)). The claim does not amount to significantly more than the abstract idea. Same rationale applies to claims 5 to 7, 9, 14, and 19 since they recite similar limitations.
Claim 8 is dependent on claim 1 and includes all the limitations of claim 1. Therefore, claim 8 recites the same abstract idea of claim 1. The claim recites the additional limitation of “modifying data of an image database to store the image and to retrievably associate the image with an identifier for the music segment”, which amounts to data storing steps, and which is considered to be insignificant extra-solution activity, (See MPEP 2106.05(g)). The identified insignificant extra-solution activity, including the data storing step, is recognized by the courts as well-understood, routine, and conventional activities when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity (See MPEP 2106.05(d)(II)(iv) Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Mm., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015)). Therefore, does not amount to significantly more than the abstract idea.
Claim 10 is dependent on claim 1 and includes all the limitations of claim 1. Therefore, claim 10 recites the same abstract idea of claim 1. The claim recites the additional limitation of “receiving, by an application instance operating on a user device, a request for the music segment, and wherein receiving the at least one of basic metadata information and lyric information comprises receiving at least one of basic metadata information and lyric information in response to receiving the request for the music segment”, which amounts to data-gathering steps, and is considered to be insignificant extra-solution activity (See MPEP 2106.05(g)), and recognized by the courts as well-understood, routine, and conventional activities when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity (See MPEP 2106.05(d) (II)(i) Receiving or transmitting data over a network, e.g., using the Internet to gather data, buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)). Therefore, the claim does not amount to significantly more than the abstract idea. Same rationale applies to claim 11.
Additionally, the claims do not include a requirement of anything other than conventional, generic computer technology for executing the abstract idea, and therefore, do not amount to significantly more than the abstract idea.
Claims 1 to 20 are therefore not drawn to eligible subject matter as they are directed to an abstract idea without significantly more.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 to 7, and 10 to 20 are rejected under 35 U.S.C. 103 as being unpatentable over Sugden (U.S. Publication No. 2025/0095222), and further in view of Liu et al. (U.S. Publication No. 2023/0326488) hereinafter Liu.
As to claim 1:
Sugden discloses:
A method of automated generation of contextually-relevant images for a music segment, the method comprising:
receiving at least one of basic metadata information and lyric information [Paragraph 0004 teaches receiving a set of descriptors; Paragraph 0016 teaches generating artistic representations of the descriptors; Paragraph 0020 teaches descriptors may include words that describe a result for the generated image];
generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, the first prompt including a first request for context information based on the at least one of the basic metadata information and the lyric information [Paragraph 0004 teaches generating a prompt input using the set of descriptors; Paragraph 0028 teaches context information may include information related to the prompt that is not directly stated in the prompt, but information generated from a database of additional information, and based on similarity metrics between a query and the additional information, where the context may be generated based on a prompt inputted into the generative model];
receiving the context information from the computer-implemented machine- learning language model in response to the first prompt [Paragraph 0028 teaches context may be generated based on a prompt inputted into the generative AI model; Paragraph 0049 teaches collecting context information];
generating a second prompt for the computer-implemented machine-learning language model based on the context information, the second prompt including a second request to generate a third prompt for a computer- implemented machine-learning image-generation model including a third request to generate an image descriptive of the music segment [Paragraph 0017 teaches the prompt engine may generate the image prompt using context that is generated using the descriptors; Paragraph 0028 teaches context may be included in the image prompt; Paragraph 0062 teaches generating a prompt input that includes context based on the descriptors];
generating the third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model [Paragraph 0017 teaches the prompt engine may generate the image prompt using context that is generated using the descriptors; Paragraph 0036 teaches generating an image prompt by preprocessing the natural language input, and descriptors; Paragraph 0056 teaches a prompt AI model that generates the prompts; Paragraph 0062 teaches generating an image prompt based on the prompt input received]; and
generating the image descriptive of the music segment by providing the third prompt as an input to the computer-implemented machine-learning image generation model [Paragraph 0032 teaches upon receipt of the request, generating an image based on the request; Paragraph 0057 teaches generating the image by submitting the image prompt to the generative model; Paragraph 0063 teaches generative AI model may then utilize the image prompt to generate an image based on the descriptors].
Sugden does not appear to expressly disclose at least one of metadata information and lyric information for the music segment; generating an image descriptive of the music segment.
Liu discloses:
at least one of metadata information and lyric information for the music segment [Paragraph 0039 teaches input text comprises one or more song lyrics];
generating an image descriptive of the music segment [Paragraph 0039 teaches generating at least one image based at least in part on text, where the text comprises song lyrics].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, to combine the teachings of the cited references and modify the invention as taught by Sugden, by at least one of metadata information and lyric information for the music segment; generating an image descriptive of the music segment, as taught by Liu [Paragraph 0039], because both applications are directed to content generation based on text prompts, including generation of images; receiving metadata information, or lyric information, as part of the request, in addition of, or as opposed to, text descriptors, is a simple substitution of one known element for another to obtain predictable results.
As to claim 2:
The combination of Sugden and Liu discloses:
generating the second prompt based on the context information and the at least one of the basic metadata information and the lyric information [Sugden - Paragraph 0017 teaches the prompt engine may generate the image prompt using context that is generated using the descriptors, therefore, based on the context information and the descriptors; Liu - Paragraph 0039 teaches input text comprises one or more song lyrics].
As to claim 3:
Sugden discloses:
the context information comprises at least one of artist context information and historical context information [Paragraph 0038 teaches context information may identify a sub-database or a filter on a database based on the descriptors, such as a filter applied to an artistic style, color, an identified emotion, etc., therefore, artist context information].
As to claim 4:
The combination of Sugden and Liu discloses:
generating a first database query based on the at least one of the basic metadata information and the lyric information [Sugden – Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information; Liu – Paragraph 0039 teaches the song may be a song that is pre-stored, where if the song is already stored, the lyrics, i.e., text, associated with the song may already be known; Paragraph 0041 teaches the song may be entered by the user, and the associated lyrics may be known, in other words, also stored in the database];
querying a first database with the first database query [Sugden – Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information, therefore, querying a database]; and
receiving first database data from the first database in response to the first database query [Sugden – Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information; Liu – Paragraph 0039 teaches the song may be a song that is pre-stored, where if the song is already stored, the lyrics, i.e., text, associated with the song may already be known; Paragraph 0041 teaches the song may be entered by the user, and the associated lyrics may be known, in other words, also retrieved from the database];
wherein generating the first prompt comprises generating the first prompt based on the first database data and the at least one of basic metadata information and lyric information [Sugden - Paragraph 0020 teaches the descriptors used to generate the image prompt include words that describe a desired result for the image that will be generated, the words input by the user, and including artistic style, colors, ideas, etc.; Paragraph 0028 teaches context may be included in the image prompt].
As to claim 5:
Sugden discloses:
generating a second database query based on the context information [Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information];
querying the first database with the second database query [Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information, therefore, querying a database]; and
receiving second database data from the first database in response to the second database query [Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information; Paragraph 0038 teaches context information may identify a sub-database or a filter on a database based on the descriptors, such as a filter applied to a color, an artistic style, an identified emotion, and so forth];
wherein generating the second prompt comprises generating the first prompt based on the first database data and context information [Paragraph 0020 teaches the descriptors used to generate the image prompt include words that describe a desired result for the image that will be generated, the words input by the user, and including artistic style, colors, ideas, etc.; Paragraph 0028 teaches context may be included in the image prompt].
As to claim 6:
Sugden discloses:
generating a second database query based on the context information [Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information];
querying a second database with the second database query [Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information, therefore, querying a database]; and
receiving second database data from the first database in response to the second database query [Paragraph 0028 teaches context is information generated based on similarity metrics between a query and a database of additional information; Paragraph 0038 teaches context information may identify a sub-database or a filter on a database based on the descriptors, such as a filter applied to a color, an artistic style, an identified emotion, and so forth];
wherein generating the second prompt comprises generating the first prompt based on the second database data and context information [Paragraph 0020 teaches the descriptors used to generate the image prompt include words that describe a desired result for the image that will be generated, the words input by the user, and including artistic style, colors, ideas, etc.; Paragraph 0028 teaches context may be included in the image prompt].
As to claim 7:
The combination of Sugden and Liu discloses:
the context information is artist context information [Sugden - Paragraph 0038 teaches context information may identify a sub-database or a filter on a database based on the descriptors, such as a filter applied to an artistic style, color, an identified emotion, etc., therefore, artist context information], and further comprising: generating a fourth prompt for the computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, the fourth prompt including a fourth request for historical context information based on the at least one of the basic metadata information and the lyric information [Sugden – Paragraph 0004 teaches generating a prompt input using the set of descriptors; Paragraph 0051 teaches collecting additional descriptors from other sources, such as internet history, identifying trends, etc., therefore, historical context information; Liu - Paragraph 0039 teaches input text comprises one or more song lyrics]; and
receiving the historical context information from the computer-implemented machine-learning language model in response to the first prompt [Sugden – Paragraph 0028 teaches context may be generated based on a prompt inputted into the generative AI model; Paragraph 0049 teaches collecting context information]; wherein
generating the second prompt comprises generating the second prompt based on the artist context information and the historical context information [Sugden – Paragraph 0017 teaches the prompt engine may generate the image prompt using context that is generated using the descriptors; Paragraph 0028 teaches context may be included in the image prompt; Paragraph 0052 teaches generating an image prompt to submit to a generative model based on a combination of collected descriptors; Paragraph 0062 teaches generating a prompt input that includes context based on the descriptors].
As to claim 10:
The combination of Sugden and Liu discloses:
receiving, by an application instance operating on a user device, a request for the music segment [Liu – Paragraph 0041 teaches the text may comprise a single word, a phrase, a sentence, one or more song lyrics, where the text may not be manually entered by the user, for example, if the text includes song lyrics, the user may select a song that is pre-stored], and wherein receiving the at least one of basic metadata information and lyric information comprises receiving at least one of basic metadata information and lyric information in response to receiving the request for the music segment [Liu – Paragraph 0041 teaches if the text includes song lyrics, the user may select a song that is pre-stored, and the lyrics may already be known, in other words, receiving (obtaining from a database) the lyric information in response to receiving the user selection of the song].
As to claim 11:
Sugden discloses:
receiving a user preference for a user that submitted the request via the user device, the user preference describing a preferred image attribute for images conveyed by the application instance [Paragraph 0020 teaches the descriptors may include words input by the user that describe a desired result for the image, including artistic style, colors, ideas, concepts, or any other descriptors and combinations thereof].
As to claim 12:
Sugden discloses:
generating the second prompt based on the context information and the user
preference, and the second request is to generate an image descriptive of the music segment and that has the preferred image attribute [Sugden - Paragraph 0020 teaches the descriptors used to generate the image prompt include words that describe a desired result for the image that will be generated, the words input by the user, and including artistic style, colors, ideas, etc.; Paragraph 0028 teaches context may be included in the image prompt; Liu – Paragraph 0040 teaches generating the at least one image based on both the text (song lyrics) and based on a style selection, which may indicate at least one of a color, texture, or artistic style].
As to claim 13:
Sugden discloses:
generating user sentiment information by analyzing the request with a natural-language processing model, wherein the generating the second prompt further comprises generating the second prompt based on the context information, the user preference, and the sentiment information [Paragraph 0020 teaches the descriptors used to generate the image prompt include words that describe a desired result for the image that will be generated, including words describing an emotional state of the user, artistic style, colors, ideas, concepts, etc., and input by the user using natural language input, or any other input or combination; Paragraph 0028 teaches context may be included in the image prompt; Paragraph 0033 teaches identifying factors to include in the image from the natural language request submitted by the user, including abstract concepts, in other words, analyzing the request with a natural language processing model].
As to claim 14:
Sugden discloses:
receiving text data previously submitted to the application instance by the user that
submitted the request via the user device [Paragraph 0051 teaches collecting and inferring one or more descriptors from media content, including a user’s social media content, posts, viewed content, keywords, internet history, emails, text messages, etc.]; and
generating user sentiment information by analyzing the text data [Paragraph 0051 teaches collecting descriptors from the user’s content, to infer user’s emotion and user’s state of mind];
wherein generating the second prompt comprises generating the second prompt based on the context information and the sentiment information [Paragraph 0020 teaches the descriptors used to generate the image prompt include words that describe a desired result for the image that will be generated, including words describing an emotional state of the user, colors, ideas, concepts, etc., and input by the user using natural language input, or any other input or combination; Paragraph 0028 teaches context may be included in the image prompt; Paragraph 0052 teaches generating an image prompt to submit to a generative model based on a combination of collected descriptors].
As to claim 19:
The combination of Sugden and Liu discloses:
receive a request for the music segment from an application instance operating on a user device [Liu – Paragraph 0041 teaches the text may comprise a single word, a phrase, a sentence, one or more song lyrics, where the text may not be manually entered by the user, for example, if the text includes song lyrics, the user may select a song that is pre-stored],
receive the at least one of basic metadata information and lyric information after receiving the request [Liu – Paragraph 0041 teaches if the text includes song lyrics, the user may select a song that is pre-stored, and the lyrics may already be known, in other words, receiving (obtaining from a database) the lyric information in response to receiving the user selection of the song],
receive a user preference for a user that submitted the request, the user preference describing a preferred image attribute for images conveyed by the application instance [Liu – Paragraph 0040 teaches style selection may be input such as by selecting an icon corresponding to the desired style on an interface of the content application, to further generate the at least one image based on both the text (song lyrics) and based on the style selection, which may indicate at least one of a color, texture, or artistic style],
and generate the second prompt based on the context information and the user preference, wherein the second request is to generate an image descriptive of the music segment and that has the preferred image attribute [Sugden - Paragraph 0020 teaches the descriptors used to generate the image prompt include words that describe a desired result for the image that will be generated, the words input by the user, and including artistic style, colors, ideas, etc.; Liu – Paragraph 0040 teaches generating the at least one image based on both the text (song lyrics) and based on a style selection, which may indicate at least one of a color, texture, or artistic style].
Same rationale applies to claims 15 to 18, and 20, since they recite similar limitations, and are therefore, similarly rejected.
Claims 8 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Sugden (U.S. Publication No. 2025/0095222), in view of Liu et al. (U.S. Publication No. 2023/0326488) hereinafter Liu, and further in view of Murray (U.S. Publication No. 2009/0307207).
As to claim 8:
Sugden discloses:
modifying data of an image database to store the image [Paragraph 0058 teaches storing the image in the image storage; Paragraph 0022 teaches generating or obtaining images].
Sugden does not appear to expressly disclose retrievably associate the image with an identifier for the music segment.
Murray discloses:
retrievably associate the image with an identifier for the music segment [Paragraph 0068 teaches the particular image is associated with the lyric ID in the database].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, to combine the teachings of the cited references and modify the invention as taught by Sugden, by retrievably associate the image with an identifier for the music segment, as taught by Murray [Paragraph 0068], because both applications are directed to content generation; storing the image in association with the music segment enables to provide a compelling presentation of the image in combination with the musical or lyrical piece, improving thereby the user’s experience (See Murray Para [0040]).
As to claim 9:
The combination of Sugden as modified by Liu, and Murray discloses:
receiving, by an application instance operating on a user device, a request for the music segment [Liu – Paragraph 0041 teaches the user may select a song that is pre-stored; Murray – Paragraph 0078 teaches querying the database for the information needed to create the multi-media presentation];
extracting the identifier from the request [Murray – Paragraph 0078 teaches querying the database for the information needed to create the multi-media presentation, where the request contains the relevant identifiers];
querying the image database with the identifier for the music segment to retrieve the image [Murray - Paragraph 0050 teaches autoMMP database contains the associations for each word and phrase in the lyric with timing data, image data, and importance scores; Paragraph 0078 teaches importing the assets that have been identified in the MPP database as associated with the lyrics of the media];
retrieving the music segment based on the identifier [Murray - Paragraph 0078 teaches importing the music, which includes the lyrics, instrumentals, and performer’s voice]; and
after querying the database and retrieving the music segment, transmitting the music segment and the image the application instance in response to the request [Murray – Paragraph 0019 teaches simultaneously displaying the identified media while playing the corresponding audio file; Paragraph 0022 teaches after the media assets are identified and timed, they are displayed on the computer system simultaneously while playing a music audio file].
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RAQUEL PEREZ-ARROYO whose telephone number is (571)272-8969. The examiner can normally be reached Monday - Friday, 8:00am - 5:30pm, Alt Friday, EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sherief Badawi can be reached at 571-272-9782. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RAQUEL PEREZ-ARROYO/Primary Examiner, Art Unit 2169