Prosecution Insights
Last updated: October 02, 2026
Application No. 19/179,437

TEXT-BASED PICTURE GENERATION METHOD, MODEL TRAINING METHOD AND APPARATUS, DEVICE, AND STORAGE MEDIUM

Non-Final OA §103
Filed
Apr 15, 2025
Priority
Mar 03, 2023 — CN 202310240545.9 +1 more
Examiner
LAM, CHAK FUNG ANTHONY
Art Unit
2612
Tech Center
2600 — Communications
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
16 currently pending
Career history
9
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang (Patent No. CN 115631251 A), in view of Liu (Patent No. CN 108170742 A). Regarding claim 1, Yang teaches A text-based picture generation method, comprising (Yang, Pg. 5, “The invention claims a method for generating image based on text, device, electronic device, computer readable storage medium, computer program product.”) obtaining first picture description text, the first picture description text describing picture content of a picture to be generated; (Yang, Pg. 9, “FIG. 2 shows a flowchart of a method for generating image based on text according to the embodiment of the present disclosure. As shown in FIG. 2, based on the text generating image of the method 200 comprises: step S201, obtaining the first text”) performing text expansion on the first picture description text by using a picture description text expansion model, (Yang, Pg. 9, “FIG. 2 shows a flowchart of a method for generating image based on text according to the embodiment of the present disclosure. As shown in FIG. 2, based on the text generating image of the method 200 comprises: step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) to obtain a second picture description text, the picture description text expansion model being trained based on sample standard picture description texts and sample brief picture description texts of reference pictures and configured to expand a brief picture description text into a corresponding standard picture description text, (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) the standard picture description text comprising a plurality of words that describe a primary description object of a target picture (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) and at least one word that describes a secondary description object of the target picture, (Yang, Pg. 9, “step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space”) and generating a picture based on the second picture description text. (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) However, Yang is silent about and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; Liu teaches and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; (Liu, Pg. 7, “the matching unit determines that the identifying text information and the recognized text information comprises keyword for describing the monitoring entity, or the identified image information and recognizing the portrait of portrait pictures comprises a description in the monitoring entity, determining the picture matched with the monitoring entity;”) Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Yang’s art by including and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; as taught by Liu, and use that with Yang’s Method For Generating Image Based On Text, Device, Electronic Device And Medium. The motivation for the combination is to improve the text recognition model to detecting text information and context of the text. Regarding claim 2, Yang teaches The method according to claim 1, wherein the performing text expansion on the first picture description text by using the picture description text expansion model, to obtain the second picture description text comprises: determining sampling parameters of candidate words in a vocabulary by using the picture description text expansion model, a sampling parameter indicating a probability that a corresponding candidate word is sampled as a word in the second picture description text; (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) and sampling the vocabulary based on the sampling parameters of the candidate words in the vocabulary by using the picture description text expansion model, to obtain the second picture description text. (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) Regarding claim 3, The method according to claim 2, wherein the determining sampling parameters of candidate words in a vocabulary by using the picture description text expansion model comprises: determining correlation parameters of the candidate words in the vocabulary by using the picture description text expansion model, a correlation parameter indicating a semantic correlation degree between a corresponding candidate word and the first picture description text; (Yang, Pg. 9, “step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) and obtaining a description word pair, the description word pair comprising a first word in the brief picture description text and a second word in a corresponding standard picture description text pair; (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) and determining the sampling parameter of the candidate word in the vocabulary based on a co- occurrence parameter of the description word pair and the correlation parameter of the candidate word in the vocabulary by using the picture description text expansion model, (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) the co-occurrence parameter indicating a probability that the standard picture description text comprises the second word in a case that the brief picture description text comprises the first word. (Yang, Pg. 11-12, “According to some embodiments, the second expansion unit 502 comprises: a first determining sub-unit configured to obtain a pre-constructed corpus, and based on the common occurrence frequency of the word in the corpus determining the first word set associated with the semantic of the first text; and a first expansion sub-unit, configured to expand the first text based on the first word set to obtain the corresponding second text.”) Regarding claim 4, Yang teaches The method according to claim 3, further comprising: performing statistical analysis on words in the standard picture description text and words in the brief picture description text of each reference picture, to obtain a plurality of description word pairs and the co-occurrence parameter of the plurality of description word pairs. (Yang, Pg. 9/10, “In one example, can be based on a large amount of text for describing the image to construct a corpus, so as to obtain a corpus with stronger correlation with the image, so as to improve the quality of the generated image. The corpus of the corresponding field can also be constructed according to the specific field using the method. It can be understood that there is strong correlation between the word combination with higher frequency in the same text, therefore, the common occurrence frequency can screen the semantic associated word set from the corpus as the first word set, for expanding the first text. Exemplary, the first word set at least comprises one word. / It can be understood that the different word combination based on common occurrence frequency screening are respectively added to the first text, so as to obtain different second text, and so as to generate the corresponding image, to expand the diversity of the generated image. then screening the generated image based on the vector coding, so as to obtain the image with good effect and quality.”) Regarding claim 5, Yang teaches The method according to claim 4, further comprising: screening the plurality of description word pairs based on a co-occurrence parameter threshold, and reserving the description word pair whose co-occurrence parameter is not less than the co-occurrence parameter threshold. (Yang, Pg. 10, “It can be understood that the different word combination based on common occurrence frequency screening are respectively added to the first text, so as to obtain different second text, and so as to generate the corresponding image, to expand the diversity of the generated image. then screening the generated image based on the vector coding, so as to obtain the image with good effect and quality.”) Regarding claim 6, Yang teaches The method according to claim 2, wherein the sampling the vocabulary based on the sampling parameters of the candidate words in the vocabulary by using the picture description text expansion model, to obtain the second picture description text comprises: (Yang, Pg. 9, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) sampling a plurality of words in the vocabulary whose sampling parameters satisfy a sampling condition based on the sampling parameters of the candidate words in the vocabulary by using the picture description text expansion model, to obtain a plurality of pieces of second picture description text, different pieces of second picture description text comprising different words satisfying the sampling condition; (Yang, Pg. 9, “S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) and the method further comprises: performing operations of generating pictures based on the second picture description text respectively for the plurality of pieces of second picture description text. (Yang, Pg. 9, “S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) Regarding claim 7, Yang teaches The method according to claim 1, wherein the generating a picture based on the second picture description text comprises: obtaining a plurality of random factors, the random factor indicating an initial state of a to- be-generated picture; and generating pictures based on the random factors and the second picture description text respectively for the plurality of random factors. (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts;”) Regarding claim 8, Yang teaches The method according to claim 1, wherein there are a plurality of pictures; the method further comprises: sorting the plurality of pictures based on at least one of correlation parameters and quality parameters of the plurality of pictures, the correlation parameter of the picture indicating a correlation degree between the picture and the first picture description text; (Yang, Pg. 9, “According to some embodiments, the corresponding relationship between the text and the drawing type and/or artist name can be determined by the way of pre-traversing the brush selection. Specifically, it can obtain the drawing type set and/or artist name set, then respectively based on each of the set in the text to expand, to construct such as text, artist 1, text, artist 2, "text, artist 3" such expansion text, respectively generating the corresponding image based on each extended text, and screening the image with high image quality through the similarity of the text vector and the image vector, so as to determine the corresponding relation of the text and the corresponding artist and/or drawing type set and forming a template, and applying the corresponding relation between the constructed text and the drawing type and/or artist name, namely the template to the expansion of the first text.”) and displaying at least one picture (Yang, Pg. 7, “The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface.”) based on an arrangement order of the plurality of pictures. (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”, The examiner interpret the co-occurrence frequency as the priority/order of the pictures.) Regarding claim 9 The method according to claim 8, wherein the displaying at least one picture (Yang, Pg. 7, “The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface.”) based on an arrangement order of the plurality of pictures comprises: arranging and displaying the plurality of pictures according to the arrangement order of the plurality of pictures; or displaying the picture ranking the first; or displaying a plurality of pictures ranking at top target positions based on the arrangement order of the plurality of pictures (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”, The examiner interpret the co-occurrence frequency as the priority/order of the pictures.) Regarding claim 10, Yang teaches and training the picture description text expansion model based on the sample standard picture description texts and the sample brief picture description texts. (Yang, Pg 9, “As shown in FIG. 2, based on the text generating image of the method 200 comprises: step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) However, Yang is silent about The method according to claim 1, wherein the picture description text expansion model is trained by: for each of the reference pictures, obtaining description of the reference picture in a network as the corresponding sample standard picture description text; performing keyword extraction on the sample standard picture description text, and using an extracted keyword as the sample brief picture description text Liu teaches The method according to claim 1, wherein the picture description text expansion model is trained by: for each of the reference pictures, (Liu, Pg 7, “the matching unit determines that the identifying text information and the recognized text information comprises keyword for describing the monitoring entity, or the identified image information and recognizing the portrait of portrait pictures comprises a description in the monitoring entity, determining the picture matched with the monitoring entity;”) obtaining description of the reference picture in a network as the corresponding sample standard picture description text; (Liu, Pg. 10, “FIG. 5 is based on the introduction, opinion obtaining method of the present invention the picture. As shown in FIG. 5, firstly, obtaining the definition of the information source and monitoring entity; after obtaining the real-time stream data from the information source, for each picture in the real-time stream data, respectively performing the identifying text information and image information or logo information identify it; determining whether the identification in the text information comprises a description keyword of the monitoring entity and the identified image or logo is included for the picture image or logo in describing the monitoring entity, namely, determining whether the identifying text information with the monitoring entity for describing the keyword matched and whether the identification figure or logo picture of the entity matching for describing the monitoring. if so, determining the picture matched with the monitoring entity; further, determine whether it has stored picture corresponding to sentiment information, if so, combining the public opinion information and opinion information stored picture corresponding to, or according to a predetermined structured format information, generating picture corresponding to sentiment information and store it.”) performing keyword extraction on the sample standard picture description text, and using an extracted keyword as the sample brief picture description text (Liu, Pg. 7, “the matching unit determines that the identifying text information and the recognized text information comprises keyword for describing the monitoring entity, or the identified image information and recognizing the portrait of portrait pictures comprises a description in the monitoring entity, determining the picture matched with the monitoring entity;”) Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Yang’s art by including The method according to claim 1, wherein the picture description text expansion model is trained by: for each of the reference pictures, obtaining description of the reference picture in a network as the corresponding sample standard picture description text; performing keyword extraction on the sample standard picture description text, and using an extracted keyword as the sample brief picture description text as taught by Liu, and use that with Yang’s Method For Generating Image Based On Text, Device, Electronic Device And Medium. Regarding claim 11, Kang teaches A text-based picture generation apparatus, comprising: a processor and a memory, the memory having at least one computer-readable instruction stored therein, and the at least one computer-readable instruction being loaded and executed by the processor to implement: (Kang, Pg. 5-6, “According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively coupled to the at least one processor; wherein the memory is stored with instructions executable by the at least one processor, the instructions are executed by the at least one processor, so that the at least one processor is capable of executing the method.”) obtaining first picture description text, the first picture description text describing picture content of a picture to be generated; (Yang, Pg. 9, “FIG. 2 shows a flowchart of a method for generating image based on text according to the embodiment of the present disclosure. As shown in FIG. 2, based on the text generating image of the method 200 comprises: step S201, obtaining the first text”) performing text expansion on the first picture description text by using a picture description text expansion model, (Yang, Pg. 9, “FIG. 2 shows a flowchart of a method for generating image based on text according to the embodiment of the present disclosure. As shown in FIG. 2, based on the text generating image of the method 200 comprises: step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) to obtain a second picture description text, the picture description text expansion model being trained based on sample standard picture description texts and sample brief picture description texts of reference pictures and configured to expand a brief picture description text into a corresponding standard picture description text, (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) the standard picture description text comprising a plurality of words that describe a primary description object of a target picture (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) and at least one word that describes a secondary description object of the target picture, (Yang, Pg. 9, “step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space”) and generating a picture based on the second picture description text. (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) However, Yang is silent about and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; Liu teaches and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; (Liu, Pg. 7, “the matching unit determines that the identifying text information and the recognized text information comprises keyword for describing the monitoring entity, or the identified image information and recognizing the portrait of portrait pictures comprises a description in the monitoring entity, determining the picture matched with the monitoring entity;”) Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Yang’s art by including and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; as taught by Liu, and use that with Yang’s Method For Generating Image Based On Text, Device, Electronic Device And Medium. Regarding claim 12, Yang teaches The apparatus according to claim 11, wherein the performing text expansion on the first picture description text by using the picture description text expansion model, to obtain the second picture description text comprises: determining sampling parameters of candidate words in a vocabulary by using the picture description text expansion model, a sampling parameter indicating a probability that a corresponding candidate word is sampled as a word in the second picture description text; (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) and sampling the vocabulary based on the sampling parameters of the candidate words in the vocabulary by using the picture description text expansion model, to obtain the second picture description text. (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) Regarding claim 13, The apparatus according to claim 12, wherein the determining sampling parameters of candidate words in a vocabulary by using the picture description text expansion model comprises: determining correlation parameters of the candidate words in the vocabulary by using the picture description text expansion model, a correlation parameter indicating a semantic correlation degree between a corresponding candidate word and the first picture description text; (Yang, Pg. 9, “step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) and obtaining a description word pair, the description word pair comprising a first word in the brief picture description text and a second word in a corresponding standard picture description text pair; (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) and determining the sampling parameter of the candidate word in the vocabulary based on a co- occurrence parameter of the description word pair and the correlation parameter of the candidate word in the vocabulary by using the picture description text expansion model, (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) the co-occurrence parameter indicating a probability that the standard picture description text comprises the second word in a case that the brief picture description text comprises the first word. (Yang, Pg. 11-12, “According to some embodiments, the second expansion unit 502 comprises: a first determining sub-unit configured to obtain a pre-constructed corpus, and based on the common occurrence frequency of the word in the corpus determining the first word set associated with the semantic of the first text; and a first expansion sub-unit, configured to expand the first text based on the first word set to obtain the corresponding second text.”) Regarding claim 14, Yang teaches The apparatus according to claim 13, further comprising: performing statistical analysis on words in the standard picture description text and words in the brief picture description text of each reference picture, to obtain a plurality of description word pairs and the co-occurrence parameter of the plurality of description word pairs. (Yang, Pg. 9/10, “In one example, can be based on a large amount of text for describing the image to construct a corpus, so as to obtain a corpus with stronger correlation with the image, so as to improve the quality of the generated image. The corpus of the corresponding field can also be constructed according to the specific field using the method. It can be understood that there is strong correlation between the word combination with higher frequency in the same text, therefore, the common occurrence frequency can screen the semantic associated word set from the corpus as the first word set, for expanding the first text. Exemplary, the first word set at least comprises one word. / It can be understood that the different word combination based on common occurrence frequency screening are respectively added to the first text, so as to obtain different second text, and so as to generate the corresponding image, to expand the diversity of the generated image. then screening the generated image based on the vector coding, so as to obtain the image with good effect and quality.”) Regarding claim 15, Yang teaches The apparatus according to claim 14, further comprising: screening the plurality of description word pairs based on a co-occurrence parameter threshold, and reserving the description word pair whose co-occurrence parameter is not less than the co-occurrence parameter threshold. (Yang, Pg. 10, “It can be understood that the different word combination based on common occurrence frequency screening are respectively added to the first text, so as to obtain different second text, and so as to generate the corresponding image, to expand the diversity of the generated image. then screening the generated image based on the vector coding, so as to obtain the image with good effect and quality.”) Regarding claim 16, Yang teaches The apparatus according to claim 12, wherein the sampling the vocabulary based on the sampling parameters of the candidate words in the vocabulary by using the picture description text expansion model, to obtain the second picture description text comprises: (Yang, Pg. 9, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”) sampling a plurality of words in the vocabulary whose sampling parameters satisfy a sampling condition based on the sampling parameters of the candidate words in the vocabulary by using the picture description text expansion model, to obtain a plurality of pieces of second picture description text, different pieces of second picture description text comprising different words satisfying the sampling condition; (Yang, Pg. 9, “S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) and the method further comprises: performing operations of generating pictures based on the second picture description text respectively for the plurality of pieces of second picture description text. (Yang, Pg. 9, “S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) Regarding claim 17, Yang teaches The apparatus according to claim 11, wherein the generating a picture based on the second picture description text comprises: obtaining a plurality of random factors, the random factor indicating an initial state of a to- be-generated picture; and generating pictures based on the random factors and the second picture description text respectively for the plurality of random factors. (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts;”) Regarding claim 18, Yang teaches The apparatus according to claim 11, wherein there are a plurality of pictures; the method further comprises: sorting the plurality of pictures based on at least one of correlation parameters and quality parameters of the plurality of pictures, the correlation parameter of the picture indicating a correlation degree between the picture and the first picture description text; (Yang, Pg. 9, “According to some embodiments, the corresponding relationship between the text and the drawing type and/or artist name can be determined by the way of pre-traversing the brush selection. Specifically, it can obtain the drawing type set and/or artist name set, then respectively based on each of the set in the text to expand, to construct such as text, artist 1, text, artist 2, "text, artist 3" such expansion text, respectively generating the corresponding image based on each extended text, and screening the image with high image quality through the similarity of the text vector and the image vector, so as to determine the corresponding relation of the text and the corresponding artist and/or drawing type set and forming a template, and applying the corresponding relation between the constructed text and the drawing type and/or artist name, namely the template to the expansion of the first text.”) and displaying at least one picture (Yang, Pg. 7, “The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface.”) based on an arrangement order of the plurality of pictures. (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”, The examiner interpret the co-occurrence frequency as the priority/order of the pictures.) Regarding claim 19 The apparatus according to claim 18, wherein the displaying at least one picture (Yang, Pg. 7, “The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface.”) based on an arrangement order of the plurality of pictures comprises: arranging and displaying the plurality of pictures according to the arrangement order of the plurality of pictures; or displaying the picture ranking the first; or displaying a plurality of pictures ranking at top target positions based on the arrangement order of the plurality of pictures (Yang, Pg. 10, “Exemplary, hot pot and " meat " common frequency is higher, therefore, when the first text comprises " hot pot ", can be the " meat " with higher frequency is added to the expansion result. Similarly, "hot pot" and "winter" word co-occurrence frequency is high, therefore, when the first text comprises hot pot, it can be co-occurrence frequency of the " winter " is added to the expansion result, so as to obtain the hot pot, "winter" and "meat" and other words of the second text, realizing the expansion of the first text.”, The examiner interpret the co-occurrence frequency as the priority/order of the pictures.) Regarding claim 20, Yang teaches A non-transitory computer-readable storage medium, the computer-readable storage medium having at least one computer-readable instruction stored therein, and the at least one computer-readable instruction being loaded and executed by a processor to implement: (Kang, Pg. 5-6, “According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively coupled to the at least one processor; wherein the memory is stored with instructions executable by the at least one processor, the instructions are executed by the at least one processor, so that the at least one processor is capable of executing the method.”) obtaining first picture description text, the first picture description text describing picture content of a picture to be generated; (Yang, Pg. 9, “FIG. 2 shows a flowchart of a method for generating image based on text according to the embodiment of the present disclosure. As shown in FIG. 2, based on the text generating image of the method 200 comprises: step S201, obtaining the first text”) performing text expansion on the first picture description text by using a picture description text expansion model, (Yang, Pg. 9, “FIG. 2 shows a flowchart of a method for generating image based on text according to the embodiment of the present disclosure. As shown in FIG. 2, based on the text generating image of the method 200 comprises: step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) to obtain a second picture description text, the picture description text expansion model being trained based on sample standard picture description texts and sample brief picture description texts of reference pictures and configured to expand a brief picture description text into a corresponding standard picture description text, (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) the standard picture description text comprising a plurality of words that describe a primary description object of a target picture (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions;”) and at least one word that describes a secondary description object of the target picture, (Yang, Pg. 9, “step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space”) and generating a picture based on the second picture description text. (Yang, Pg. 9, “step S201, obtaining the first text, and based on a plurality of rules to expand the first text to obtain a plurality of second text, wherein the plurality of rules for expanding the first text in different dimensions; step S202, generating a plurality of corresponding images based on the plurality of second texts; step S203, encoding the first text to determine a first vector corresponding to the first text; step S204, encoding each image of the plurality of images, to determine a second vector corresponding to each image, wherein the first vector and the second vector corresponding to each image are located in the same semantic space; and step S205, based on the similarity between the first vector and the second vector corresponding to each image, screening the plurality of images.”) However, Yang is silent about and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; Liu teaches and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; (Liu, Pg. 7, “the matching unit determines that the identifying text information and the recognized text information comprises keyword for describing the monitoring entity, or the identified image information and recognizing the portrait of portrait pictures comprises a description in the monitoring entity, determining the picture matched with the monitoring entity;”) Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Yang’s art by including and the brief picture description text being a keyword that describes the primary description object in the standard picture description text; as taught by Liu, and use that with Yang’s Method For Generating Image Based On Text, Device, Electronic Device And Medium. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHAK FUNG A LAM whose telephone number is (571)272-9823. The examiner can normally be reached Monday-Friday 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said Broome can be reached at 5712722931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center andhttps://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.A.L./Examiner, Art Unit 2612 /Said Broome/Supervisory Patent Examiner, Art Unit 2612
Read full office action

Prosecution Timeline

Apr 15, 2025
Application Filed
Sep 11, 2026
Non-Final Rejection mailed — §103
Sep 18, 2026
Interview Requested

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month