Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant’s election without traverse of the restriction requirement dated 5/04/2026 in the reply filed on 7/07/2026 is acknowledged. Examiner has noted the species of claims 7-14 have been cancelled and the species of claims 1-6 and 15-20 have been elected.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 4, 15, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ramesh et al (US 11942116 B1) and Axen (US 9129642 B2), hereinafter Ramesh and Axen respectively.
Regarding claim 15, Ramesh teaches a system comprising: one or more computers and one or more storage devices on which are stored instructions that are operable,
"A computing system comprising a processor and a non-transitory computer-readable medium having stored thereon program instructions" - Ramesh Claim 1
when executed by the one or more computers, to cause the one or more computers to perform operations comprising: receiving one or more user input identifying information associated with one or more media elements and one or more characteristics of the video to be generated;
"The set of user attributes can include a name of the user, a geographic area of the user (e.g., a current country, state, county, and/or address of the user), an employer of the user, a race and ethnicity of the user, an age of the user, a gender of the user, a marital status of the user, a salary of the user, a user-preferred language (e.g., English, Spanish, Japanese), a user-preferred secondary language, a search history of the user, a description or list of user interests/hobbies, a description or list of user-preferred products, a description or list of user-preferred services, a list of products or services purchased by the user, a user-preferred travel destination, a user-preferred spokesperson (e.g., a celebrity or other individual), a physical attribute of the user (e.g., skin color, hair color, body type, etc.), a user-preferred video content genre, and/or a user-preferred music genre or artist, among other possibilities." - Col 5, Lines 9-24
NOTE: Ramesh describes providing user attributes from a user profile. Some of the user attributes include a list of products or services which can be understood as media elements and preferred language as disclosed in Pg 14, Lines 8-13. The user attributes may also include age, region, geographic area which can be understood as user defined characteristics as disclosed in Pg 8, lines 22-25. These user attributes are then used to generate a synthetic video.
and generating video content based on the received one or more user inputs,
"In some implementations, during the process of obtaining the structured data, or during the process of generating the synthetic video described in more detail below, the video-generation system 100 can use the set of user attributes to determine a spokesperson (i.e., a human, or talking animal/object, depicted in the synthetic video that speaks synthesized speech in accordance with the targeted advertisement) to be included in the synthetic video" - Col 8, Lines 33-40
the generating comprising: identifying assets to include in the video, the assets including an avatar,
"For example, if the set of user attributes specifies a set of user-preferred physical characteristics (e.g., hair, skin color, height, weight), the video-generation system 100 can use those physical characteristics to select a set of characteristics of a spokesperson according to which to render the spokesperson in the synthetic video." - Col 8, Lines 40-46 (identifying avatar asset portion)
generating a script for the video,
“Using deep learning, the video-generation system 100 can create scripts that accurately resemble the cadence, structure, and vocabulary found in advertisements and that target specific audiences and user attributes.” – Col 9, Lines 53-56
Ramesh does not teach and assembling a video layout. However, Axen teaches generating comprising:
"A synthesis engine 66--After all the actors or media engines 64 have played their part and all of the different parts have been created, avatar, background, music, animation, effect, there is a step that composites the entire scene, frame by frame to generate a standard video file, and this step is carried out by the synthesis engine 66 to produce an output of video 68." - Col 12, Lines 9-16
NOTE: Col 36-37 describes a scene-by-scene breakdown of an advertisement for a football product. This functionally corresponds to the assembling of a video layout. After the combination, the process of assembling a video layout based on the generated video created from user input as taught by Axen can be added to Ramesh's system for generating synthetic videos based on user attributes. This modification will then allow Ramesh’s system to generate videos wherein the generating comprises: identifying assets to include in the video, the assets including an avatar, generating a script for the video, and assembling a video layout.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Ramesh by incorporating the teachings of Axen to have the generating of the video content comprise the assembling a video layout. One would be motivated to use the automatic generation of a video based on a product as taught by Axen and add it to Ramesh’s system of generating a targeted advertisement video based on the user input to produce video content that is more tailored to the target audience. Applying the teachings of Axen to Ramesh for assembling a video layout would also provide an automated manner of organizing the assets, avatar, script, and scenes for generation which reduces manual video production work.
Regarding claim 1, the claim recites similar limitations to claim 15. Therefore, method claim 1 corresponds to the system disclosed in claim 15 and is rejected for the same reasons of obviousness as used above.
Regarding claim 18, Ramesh in view of Axen teaches the system of claim 15.Ramesh as modified teaches wherein assembling the video layout comprises defining a sequence of scenes, each scene having a particular avatar look and segment of the generated script.
" A synchronization engine 62--manages the different media activation tasks and tells each media activation engine what it should do. For example: for an avatar it builds the script that includes the sentences that need to be said, the gestures with time stamps and so on." – Axen Col 11, Lines 65-67 and Col 12, Lines 1-2
NOTE: Axen discloses using a synchronization engine for building the script that will be spoken and the avatar gestures that will be performed in the video. This implies that during Axen’s video layout process: "A synthesis engine 66--After all the actors or media engines 64 have played their part and all of the different parts have been created, avatar, background, music, animation, effect, there is a step that composites the entire scene, frame by frame to generate a standard video file, and this step is carried out by the synthesis engine 66 to produce an output of video 68.", see Col 12, Lines 9-16, the composition of the scenes that were assembled frame by frame takes into consideration the avatar look (gesture or pose) and the segment of the generated script that needed to be spoken during that particular scene.
Regarding claim 4, the claim recites similar limitations to claim 18. Therefore, method claim 4 corresponds to the system disclosed in claim 18 and is rejected for the same reasons of obviousness as used above.
Regarding claim 19, Ramesh in view of Axen teaches the system of claim 15.
the generating further comprising adding one or more video decorations, wherein adding video decorations comprise adding particular voice content to give voice to the script
“For example, the end user may record his/her own voice and add the audio file to the video” – Axen Col 40, Lines 36-38
and/or identifying music to include in the video.
“The storyboard may be more advanced and the video may include additional media elements such as a virtual narrator, and other media elements such as background music” – Axen Col 15, Lines 4-7
Regarding claim 5, the claim recites similar limitations to claim 19. Therefore, method claim 5 corresponds to the system disclosed in claim 19 and is rejected for the same reasons of obviousness as used above.
Claim(s) 2 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ramesh, Axen, and Saraee et al (US 20220198779 A1), hereinafter Saraee respectively.
Regarding claim 16, Ramesh in view of Axen teaches the system of claim 15. Ramesh as modified does not teach wherein identifying assets includes using a multi-modal machine learning model to select wherein identifying assets includes using a multi-modal machine learning model to select
“For instance, the data processing system may determine the difference between a performance score of an image and the benchmark. The data processing system may then compare the difference to a threshold (e.g., a predetermined threshold). If the performance score exceeds the threshold, the data processing system generates a record recommending to upload the image to a web-based property.” – Par 854, Lines 4-11
NOTE: Saraee discloses: “The content evaluation system 1005 executes a neural network (or any other machine learning model) to generate a performance score (e.g., a score indicating a likelihood that a particular user would interact with an image) for each of the retrieved images.”, see par 809. The neural network is trained using image interaction data such as amount of views, likes, ratings, comments (audience interest), see Par 834. A target audience threshold determines that if the performance score exceeds the predetermined threshold, then the image is recommended to be uploaded for use. The neural network model can also generate the performance score for a plurality of images or videos, see Abstract. This functionally corresponds to a multi-modal machine learning model selecting an image asset that satisfies a threshold probability of being an interest to a target based on the understanding of as subject of the video. After the combination, the steps for selecting an image asset that satisfied a threshold probability of target audience interest as taught by Saraee can be applied to selecting the avatar for the synthetic video for a subject, such as a product or service, as taught by Ramesh. This modification will then allow both the image asset and avatar to be selected using a multi-modal machine learning model that determines the image assets and avatars that exceed a threshold probability of audience interest. This combination would then fully teach wherein identifying assets includes using a multi-modal machine learning model to select an avatar and image assets that satisfy a threshold probability of being of interest to a target audience based on an understanding of a subject of the video.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Ramesh by incorporating the teachings of Saraee to use a multi-modal machine learning model to select an avatar and image assets that satisfies a threshold probability of target audience interest based on an understanding of a subject of the video. One would be motivated to make this combination to generate a synthetic video related to the subject that has the highest chance for engagement or interaction. The machine learning model would be able to predict the most likely combination of assets and avatars that best suit the preference of the target audience.
Regarding claim 2, the claim recites similar limitations to claim 16. Therefore, method claim 2 corresponds to the system disclosed in claim 16 and is rejected for the same reasons of obviousness as used above.
Claim(s) 3 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ramesh, Axen, and Kaplan (WO 2024182276 A1), hereinafter Kaplan respectively.
Regarding claim 17, Ramesh in view of Axen teaches the system of claim 15. Ramesh as modified teaches wherein generating the script
“Using deep learning, the video-generation system 100 can create scripts that accurately resemble the cadence, structure, and vocabulary found in advertisements and that target specific audiences and user attributes.” – Col 9, Lines 53-56
NOTE: Ramesh teaches a video-generation system using deep learning to generate scripts that match cadence, structure, and vocabulary found in advertisements. This functionally corresponds to generating a script that matches the style and tone of video content.
Ramesh as modified still does not teach wherein a large language model trained on short form video content of a social media platform wherein
“As with YouTube® and any video streaming/hosting service, TikTok® has a huge amount of information about users in the form of short-form videos. These videos can be transcribed and analyzed automatically, converting them from video to textual datasets (or other more structured formats which might include, without limitation, video and images) that can be used to train and tune LLMs or other Al agents” – Pg 34, Par 343
NOTE: After the combination, the concept of training a large language model (LLM) on videos from a social media platform such as YouTube or TikTok as taught by Kaplan can modify Ramesh’s model for generating scripts that match the style and tone of advertisements. This modification would then allow Ramesh’s script generating model to be trained on short form videos from social media platforms and generate a script based on the style and tone it learned through training. This would then teach wherein generating the script comprises using a large language model trained on short form video content of a social media platform such that the script matches a style and tone of video content on the social media platform
It would have been obvious to one of ordinary skill before the effective filing date of the present invention to modify Ramesh by incorporating the teachings of Kaplan to use a large language model trained on videos of a short form videos from a social media platform to generate a script that matches the style and tone of video content on the social media platform. One would be motivated to train a model to understand what qualities of the short form videos on the social media platform will perform well with an audience will produce scripts that are more likely to receive more attention and engagement.
Regarding claim 3, the claim recites similar limitations to claim 17. Therefore, method claim 3 corresponds to the system disclosed in claim 17 and is rejected for the same reasons of obviousness as used above.
Claim(s) 6 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ramesh, Axen, and Cohen-Martin (US 20150221000 A1), hereinafter Cohen-Martin respectively.
Regarding claim 20, Ramesh in view of Axen teaches the system of claim 15. Ramesh as modified does not teach wherein receiving one or more user inputs comprises receiving a user specification of a reference to a location containing information about a subject product. However, Cohen-Martin teaches wherein receiving one or more user inputs comprises receiving a user specification of a reference to a location containing information about a subject product.
“The web client may also be prompted to indicate a website/webpage (typically an address of the website, a URL) from which dynamic data objects are to be selected ” – Par 289, Lines 1-3
NOTE: Cohen-Martin discloses providing a address of the website (URL of the website). Then the dynamic data objects are extracted and used to create an advertisement video: “In some embodiments of the invention using a manual mode of selection, the user may select dynamic data objects, which may comprise references to dynamic data sources like repeaters or parts of a live webpage, databases tables or fields, etc. These may, for example, include, images of items offered for sale (such as, for example, photographs of products on sale, photographs of tourist locations subject to offered travel deals), related information, including, for example, the quoted price, a description of the item.”, see Par 277. The product URL link provided corresponds to the user input of a reference to a location about a product as disclosed in Pg 8, Lines 13-17. After the combination, the concept of providing a URL of a product to obtain the product information as taught by Cohen-Martin can modify Ramesh system for synthetic advertisement video generation. This modification would allow a user to enter a product URL to obtain information from the website linked to the URL to then be used as the subject media element for the generated advertisement video.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Ramesh by incorporating the teachings of Cohen-Martin to receive a user specification of a reference to a location containing information about a subject product. One would be motivated to make this combination so that the user would no longer need to manually enter the information regarding a product and instead have the video generation system obtain the information about the product from the reference location such as an URL and use it to generate a video with the product as a subject.
Regarding claim 6, the claim recites similar limitations to claim 20. Therefore, method claim 6 corresponds to the system disclosed in claim 20 and is rejected for the same reasons of obviousness as used above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID V. NGUYEN whose telephone number is (571)272-6111. The examiner can normally be reached M-F 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Y Poon can be reached at 571-270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID VAN NGUYEN/Examiner, Art Unit 2617 /KING Y POON/Supervisory Patent Examiner, Art Unit 2617