DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 7-11, 15, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wittman et al. (U.S. Patent Application Publication No. 2025/0356883), hereinafter referenced as Wittman in view of Hoffer (U.S. Patent Application Publication No. 2025/0265758), hereinafter referenced as Hoffer.
Regarding claim 1, Wittman teaches a method comprising: obtaining, from one or more client devices, input data corresponding to a narrative (paragraph 21 teaches “receiving user input including a description of a story that a user seeks to convey in a video”); this shows input data (from user therefore the client device user is using) which corresponds to story/narrative; determining, based at least on one or more first language models processing the input data, a series of scenes corresponding to respective portions of the input data (paragraph 21 teaches “transmitting a first prompt to a LLM system, the first prompt including at least a portion of the user input, receiving, from the LLM system, storyboard code representative of a set of scenes”); user input to LLM (large language model) shows first language model processing the aforementioned user input data and this leads to set/series of scenes being determined corresponding to respective portions of the input data; generating, based at least on the scene information, image data representing one or more images depicting one or more visual representations corresponding to the one or more scenes (paragraph 21 teaches “the storyboard including graphical display of one or more scenes of the set of scenes”, paragraph 103 teaches “A scene can include one or more headlines, text, images” and paragraph 186 teaches “storyboard 600 that is generated using AI in accordance with implementations of the present disclosure. The example of FIG. 6 is generated based on the example user input introduced above (e.g., As the Head of the new technologies organization at ACME, . . . ) and includes scenes 602, 604, 606, 608, 610”); scene including images and storyboard generated including scenes shows image data (images) depicting visual representation of scenes being generated and this would be based on scene information since are of the scene; generating one or more storyboard frames corresponding to the one or more scenes, the one or more storyboard frames including at least the scene information and the one or more images (paragraph 78 teaches “includes a storyboard generator 400, a layout selector 402, a data extractor/injector 404, and one or more tools 406. In some examples, the storyboard generator 400 interacts with a LLM system (e.g., the LLM system 360 of FIG. 3) to provide a storyboard based on user input. In some examples, the storyboard includes a set of scenes to be depicted in a video… determine a layout for each scene of the storyboard.”); storyboard having set of scenes means it includes scene information and for the storyboard to be created that includes scenes depicted in video means generation of storyboard frames including the images; and sending, to the one or more client devices, data representing one or more storyboards that include the one or more storyboard frames (paragraph 21 teaches “and displaying, in a user interface (UI), a storyboard for the video” and paragraph 72 teaches “stories can have a default resolution (e.g., 720×1280p×) and/or frame rate (e.g., 24 fps)”); this shows data representing storyboards including frames (of storyboard) would be sent to user/client device to be displayed.
However, Wittman fails to teach generating, based at least on one or more second language models processing the respective portions of the input data, text data representing scene information corresponding to one or more scenes of the series of scenes;
However, Hoffer teaches generating, based at least on one or more second language models processing the respective portions of the input data, text data representing scene information corresponding to one or more scenes of the series of scenes (Hoffer, paragraph 26 teaches “generating prompts from the labels, the image segments, and the text segments, the prompts representing script information, storyboard information, or scene information” and paragraph 55 teaches “Additionally or alternatively, the text in the bubbles can be modified using a large language model (LLM), For example, to abbreviate the text without substantively changing its meaning. Thus, the modifications to text or dialog can be made to be consistent with the storyline, such that the modifications do not disrupt of flow of the storyline”); modified/abbreviated text here shows text data generated, this is from another/second LLM, and the text data is included in prompt representing scene information which corresponds to scenes. Hoffer is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of rendering digital graphics using text as input. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Wittman’s invention with the additional/second LLM and text data representing scene information techniques of Hoffer to provide an improved user experience of the digital version of the graphic narrative. This would be done by the modification of text to abbreviate it without changing the meaning and without disrupting flow of the storyline.
Regarding claim 2, the combination of Wittman and Hoffer teaches further comprising: generating, based at least on one or more differences between a sequential pair of storyboard frames, one or more intermediate storyboard frames (Hoffer, paragraph 85 teaches “prompts 424 can include this storyboard, which can be augmented with additional information to interpolate/extrapolate and fill any gaps remaining in the storyboard”); this shows intermediate/interpolated frames being generated and one of ordinary skill in the art would understand that interpolation is done between sequential pair of frames by finding differences between such, therefore, this is based on the differences between sequential pair of storyboard frames; and generating one or more animatics including at least the one or more intermediate storyboard frames between the sequential pair of storyboard frames (Wittman, paragraph 27 teaches “As used herein, a video, also referred to as a story, can be described as a composition of scenes, visual elements, style instructions, animation and timing settings,” and paragraph 72 teaches “usable with story templates (e.g., settings, scenes, animations, audio)”); this shows for animatics (animations with timings, audio and story templates with animations) being generated using data which represent the storyboard (that the animations are for), therefore would include the aforementioned intermediate/interpolated storyboard frame between the sequential pair of storyboard frames. The same motivations used in claim 1 apply here in claim 2.
Regarding claim 7, the combination of Wittman and Hoffer teaches further comprising: applying, as input to one or more machine learning models, at least a portion of the text data representative of the scene information, (Hoffer, abstract teaches “Pages of a graphic narrative (e.g., comic book) are partitioned into panels, which are segmented into image segmented elements and text elements. The segmented elements are applied to a machine learning (ML)”); this shows portion/segmented element of text data representing scene information being applied to ML; wherein the generating of the image data representing the one or more images depicting the one or more visual representations corresponding to the one or more scenes is further based at least on the one or more machine learning models processing at least the portion of the text data (Wittman, paragraph 21 teaches “the storyboard including graphical display of one or more scenes of the set of scenes”, paragraph 103 teaches “A scene can include one or more headlines, text, images” and paragraph 186 teaches “storyboard 600 that is generated using AI in accordance with implementations of the present disclosure. The example of FIG. 6 is generated based on the example user input introduced above (e.g., As the Head of the new technologies organization at ACME, . . . ) and includes scenes 602, 604, 606, 608, 610”); scene including images and storyboard generated including scenes shows image data (images) depicting visual representation of scenes being generated, this would be corresponding to scene information since are of the scene, and it is based on the aforementioned ML processing of portion/segment of text data since is based on the user input which would be the portion of text data when viewed in combination. The same motivations used in claim 1 apply here in claim 7.
Regarding claim 8, Wittman teaches A system comprising: one or more processors to: (fig. 10 and paragraph 196 teaches “system 1000 includes a processor 1010”); and send, to one or more client devices, the data representing the one or more storyboard frames (paragraph 21 teaches “and displaying, in a user interface (UI), a storyboard for the video” and paragraph 72 teaches “stories can have a default resolution (e.g., 720×1280p×) and/or frame rate (e.g., 24 fps)”); this shows data representing storyboards including frames (of storyboard) would be sent to user/client device to be displayed.
However, Wittman fails to teach segment, using one or more language models, text data representing a narrative into a plurality of scenes corresponding to respective portions of the text data; generate, using one or more machine learning models and based at least on the respective portions of the text data, data representing one or more storyboard frames corresponding to one or more scenes of the plurality of scenes.
However, Hoffer teaches segment, using one or more language models, text data representing a narrative into a plurality of scenes corresponding to respective portions of the text data (Hoffer, abstract teaches “Pages of a graphic narrative (e.g., comic book) are partitioned into panels, which are segmented into image segmented elements and text elements… prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes). Thus, the comic book is effectively a movie storyboard that is automatically converted into full-motion rendered graphics by treating each combination of text and graphics as a unique prompt”, paragraph 74 teaches “segmentation processor 408 receives panels 406 and generates therefrom segmented elements 410, including image segments and text segments”, and paragraph 83 teaches “can use a large language model (LLM), such as those discussed above for the segmentation processor 408”); this shows segmentation of text data which represents narrative/story/comic, segmentation processor uses LLM to do so, and since each text and graphic is a unique prompt (which includes script or scene), this text data is segmented into a plurality of scenes corresponding to respective portions of the text data; generate, using one or more machine learning models and based at least on the respective portions of the text data, data representing one or more storyboard frames corresponding to one or more scenes of the plurality of scenes (Hoffer, abstract teaches “segmented elements are applied to a machine learning (ML) method that labels/identifies the segmented elements. Prompts based on the labels are then applied to a second ML model… prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes). Thus, the comic book is effectively a movie storyboard that is automatically converted” and paragraph 86 teaches “prompts 424 can represent script information, storyboard information, or scene information corresponding to one or more of the panels.”); this shows using ML (and based on segmented elements which are respective portions of text data) to generate data/prompt which represent storyboard (and frames thereof) corresponding to the scenes (when viewed in combination). Hoffer is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of rendering digital graphics using text as input. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Wittman’s invention with the segmentation and text data representing scene information techniques of Hoffer to provide an improved user experience of the digital version of the graphic narrative. This would be done by the segmenting of text into plurality of scenes using an LLM.
Regarding claim 9, the combination of Wittman and Hoffer teaches wherein the segmentation of the text data into the plurality of scenes using the one or more language models comprises, at least: preprocessing the text data to at least one of normalize or tokenize the text data (Hoffer, paragraph 125 teaches “the input embedding block 704 can be learned embeddings to convert the input tokens and output tokens to vectors”); due to the input tokens, text (segmented text from abstract) must have undergone tokenization somewhere before those tokens are supplied to the embedding block (thus as a preprocessing); processing, using one or more Bidirectional Encoder Representations from Transformers (BERT) models, the text data or a preprocessed version of the text data to generate a plurality of sentence embeddings (Hoffer, paragraph 123 teaches “transformers include a Bidirectional Encoder Representations from Transformer (BERT) and a Generative Pre-trained Transformer (GPT). The transformer architecture 700, which is illustrated in FIG. 7A through FIG. 7C, includes inputs 702, an input embedding block” and paragraph 125 teaches “input embedding block 704 is used to provide representations for words. For example, embedding can be used in text analysis. According to certain non-limiting examples, the representation is a real-valued vector that encodes the meaning of the word in such a way that words that are closer in the vector space are expected to be similar in meaning. Word embeddings can be obtained using language modeling and feature learning techniques, where words or phrases from the vocabulary are mapped to vectors of real numbers”); this shows processing using BERT, text data to generate phrase embeddings which one of ordinary skill in the art would understand are sentences; computing one or more scores indicative of degree of similarity between one or more first sentence embeddings and one or more second sentence embeddings of the plurality of sentence embeddings (Hoffer, paragraph 130 teaches “the linear block 716 projects the output from the last decode block 714c into word scores for the second language (e.g., a score value for each unique word in the target vocabulary) at each position in the sentence. For instance, if the output sentence has seven words and the provided vocabulary for the second language has 10,000 unique words, then 10,000 score values are generated for each of those seven words); score value for unique words would indicate degree of similarity between two of the aforementioned sentence embeddings; and associating, based at least on the one or more scores, one or more subsets of the plurality of sentences with the one or more scenes (Hoffer, paragraph 143 teaches “inputs 802 are applied as inputs to the trained ML model 804 to generate the outputs, which can include the outputs 806.”); this shows inputs and outputs (including sentences which are considered subset and in a step after scores thus based on such) for ML model which in the abstract takes prompt including scene therefore associated with such. The same motivations used in claim 8 apply here in claim 9.
Regarding claim 10, the combination of Wittman and Hoffer teaches the one or more processors further to: determine, using one or more second language models and based at least on the respective portions of the text data, one or more features corresponding to the one or more scenes, (Hoffer, paragraph 53 teaches “an AI model can learn how different types of characters and objects move in motion pictures and can statistically learn various context clues regarding how to apply those movements to the characters and objects identified in the panels of a graphic narrative. Consider, for example, a portion of a superhero graphic narrative representing a fight sequence, in the fight scene, a windup for a punch or kick proceeds certain after-effects from the punch or kick, and the movement of the punch or kick can be based on physical models (e.g., the physics of momentum) and/or learned from training data set that includes moving pictures of various fight scenes. Additionally, certain large language models (LLMs) can predict which text is likely to follow which other text”); types of characters, movements and objects identified show determining features corresponding to scene, this is using AI model/LLM (second LLM), and is based on text portions due to the objects/features identified in panels of graphic narrative like a comic which also includes text; wherein the generation of the one or more storyboard frames using the one or more machine learning models is further based at least on the one or more features (Hoffer, paragraph 53 teaches “generate a moving picture consistent with a narrative flow of the graphic narrative. Further, the combination of the relative locations of the panels, the image elements represented within the panels, and the text within the panels provide sufficient context clues for an AI model to determine fluid, continuous movements of the story within and between the stationary pictures represented in the panels of a print-version of the graphic narrative.”); this shows that the aforementioned storyboard generation (using ML as mentioned in abstract and claim 8 above) would be based on the features since happens in a step after and accounts for movements in such.
The same motivations used in claim 8 apply here in claim 10.
Regarding claim 11, the combination of Wittman and Hoffer teaches the one or more processors further to: compute, for at least the one or more scenes, one or more scores indicative of one or more rankings of the one or more scenes based at least on one or more contributions of the one or more scenes to a plot associated with the narrative, (Hoffer, paragraph 115 teaches “combining the first scores and the second scored to predict order in which the panels are to be viewed.”); combined score shows computing score and is to indicate ranking/order of scene (associated with the panel) to be viewed and this is based on contribution of scene to plot associated with narrative because that is how the initial scores are derived which are used to get this combined score; wherein the generation of the one or more storyboard frames corresponding to the one or more scenes is further based at least on the one or more scores for the one or more scenes meeting or exceeding a threshold (Hoffer, paragraph 130 teaches “For instance, if the output sentence has seven words and the provided vocabulary for the second language has 10,000 unique words, then 10,000 score values are generated for each of those seven words. The score values indicate the likelihood of occurrence for each word in the vocabulary in that position of the sentence.”); certain amount of score values shows the aforementioned generation of storyboard frames corresponding to scenes would be based on scores for scene meeting or exceeding a threshold (threshold would be the minimum score value allowed). The same motivations used in claim 8 apply here in claim 11.
Regarding claim 15, the combination of Wittman and Hoffer teaches the one or more processors further to: generate, based at least on a sequential pair of storyboard frames, one or more intermediate storyboard frames representative of a transition between the sequential pair of storyboard frames (Hoffer, paragraph 85 teaches “prompts 424 can include this storyboard, which can be augmented with additional information to interpolate/extrapolate and fill any gaps remaining in the storyboard”, paragraph 102 teaches “stitching processor 430 can use a generative AI model to provide transitions between the respective moving pictures”, and paragraph 64 teaches “transition from the second frame in FIG. 2B to the third frame in FIG. 2C”); this shows intermediate/interpolated storyboard frames being generated and one of ordinary skill in the art would understand that interpolation is done between sequential pair of frames, also since each frame here corresponds to moving picture and has transition, this would be representative of transition between the sequential pair of storyboard frames; generate one or more animatics including at least the sequential pair of storyboard frames and the one or more intermediate storyboard frames (Wittman, paragraph 27 teaches “As used herein, a video, also referred to as a story, can be described as a composition of scenes, visual elements, style instructions, animation and timing settings,” and paragraph 72 teaches “usable with story templates (e.g., settings, scenes, animations, audio)”); this shows for animatics (animations with timings, audio and story templates with animations) being generated using data which represent the storyboard (that the animations are for), therefore would include the aforementioned intermediate/interpolated storyboard frame between the sequential pair of storyboard frames along with the sequential pair of storyboard frames; and send, to the one or more client devices, the one or more animatics (Wittman, paragraph 127 teaches “Lottie animation of a corporate building with the ACME logo is displayed”); animatics/animation being displayed means first it must be sent to client device. The same motivations used in claim 8 apply here in claim 15.
Regarding claim 18, the combination of Wittman and Hoffer teaches wherein the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing deep learning operations;
a system for performing remote operations (Wittman, paragraph 202 teaches “computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a network, such as the described one”); this shows system for performing remote operations;
a system for performing real-time streaming (Wittman, paragraph 192 teaches “executed for real-time generation of videos”); real-time generation would also mean the system is capable of real-time streaming;
a system for generating or presenting one or more of augmented reality content,
virtual reality content, or mixed reality content;
a system implemented using an edge device;
a system implemented using a robot;
a system for performing conversational AI operations;
a system implementing one or more multi-model language models;
a system implementing one or more large language models (LLMs) (Wittman, paragraph 44 teaches "the conceptual architecture 200 includes a LLM system 230”); this shows system implementing LLMs;
a system implementing one or more small language models (SLMs);
a system implementing one or more vision language models (VLMs);
a system for generating synthetic data;
a system for generating synthetic data using AI;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Wittman, paragraph 39 teaches “video generation platform is provided as a cloud-based platform that can be provisioned in any appropriate cloud runtime. An example cloud runtime includes, without limitation, the SAP Kyma runtime (SKR), which is provided by SAP AG of Walldorf, Germany, and can be described as a fully managed Kubernetes-based runtime. Another example cloud runtime includes Cloud Foundry”); this shows system implemented using cloud computing resources.
Claim(s) 3-6, 12-14, 17 and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wittman in view of Hoffer and Rothe-Kushel (U.S. Patent Application Publication No. 2026/0088050), hereinafter referenced as Rothe.
Regarding claim 19, Wittman teaches One or more processors comprising: processing circuitry to cause presentation, on one or more displays of one or more client devices, (fig. 10 and paragraph 196 teaches “system 1000 includes a processor 1010… display graphical information for a user interface on the input/output device 1040”); this shows processor with processing circuitry would cause presentation on the display of client devices; of one or more animatics generated using data representing a storyboard corresponding to a narrative, (paragraph 27 teaches “As used herein, a video, also referred to as a story, can be described as a composition of scenes, visual elements, style instructions, animation and timing settings,” and paragraph 72 teaches “usable with story templates (e.g., settings, scenes, animations, audio)”); this shows for presentation/to display animatics (animations with timings, audio and story templates with animations) being generated using data which represent the storyboard (that the animations are for) corresponding to the story/narrative; wherein the data representing the storyboard is generated, at least, by: and image data representing one or more images depicting one or more visual representations associated with the one or more scenes (paragraph 21 teaches “the storyboard including graphical display of one or more scenes of the set of scenes”, paragraph 103 teaches “A scene can include one or more headlines, text, images” and paragraph 186 teaches “storyboard 600 that is generated using AI in accordance with implementations of the present disclosure. The example of FIG. 6 is generated based on the example user input introduced above (e.g., As the Head of the new technologies organization at ACME, . . . ) and includes scenes 602, 604, 606, 608, 610”); scene including images and storyboard generated including scenes shows image data (images) depicting visual representation of scenes being generated.
However, Wittman fails to teach segmenting, using one or more language models, text data representing the narrative into one or more scenes corresponding to one or more portions of the text data; and generating, using one or more machine learning models and based at least on the segmenting, one or more frames of the storyboard corresponding to the one or more scenes, the one or more frames including, at least: scene description data indicative of one or more scene descriptions associated with the one or more scenes; cinematographic data indicative of one or more cinematographic details associated with the one or more scenes.
However, Hoffer teaches segmenting, using one or more language models, text data representing the narrative into one or more scenes corresponding to one or more portions of the text data (Hoffer, abstract teaches “Pages of a graphic narrative (e.g., comic book) are partitioned into panels, which are segmented into image segmented elements and text elements… prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes). Thus, the comic book is effectively a movie storyboard that is automatically converted into full-motion rendered graphics by treating each combination of text and graphics as a unique prompt”, paragraph 74 teaches “segmentation processor 408 receives panels 406 and generates therefrom segmented elements 410, including image segments and text segments”, and paragraph 83 teaches “can use a large language model (LLM), such as those discussed above for the segmentation processor 408”); this shows segmentation of text data which represents narrative/story/comic, segmentation processor uses LLM to do so, and since each text and graphic is a unique prompt (which includes script or scene), this text data is segmented into a plurality of scenes corresponding to respective portions of the text data; and generating, using one or more machine learning models and based at least on the segmenting, one or more frames of the storyboard corresponding to the one or more scenes, (Hoffer, abstract teaches “segmented elements are applied to a machine learning (ML) method that labels/identifies the segmented elements. Prompts based on the labels are then applied to a second ML model… prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes). Thus, the comic book is effectively a movie storyboard that is automatically converted” and paragraph 86 teaches “prompts 424 can represent script information, storyboard information, or scene information corresponding to one or more of the panels.”); this shows using ML (and based on segmented elements therefore the segmenting) to generate prompt which represents storyboard (and frames thereof) corresponding to the scenes (when viewed in combination); Hoffer is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of rendering digital graphics using text as input. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Wittman’s invention with the segmentation and text data representing scene information techniques of Hoffer to provide an improved user experience of the digital version of the graphic narrative. This would be done by the segmenting of text into plurality of scenes using an LLM.
However, the combination of Wittman and Hoffer fails to teach the one or more frames including, at least: scene description data indicative of one or more scene descriptions associated with the one or more scenes; cinematographic data indicative of one or more cinematographic details associated with the one or more scenes
However, Rothe teaches
the one or more frames including, at least: scene description data indicative of one or more scene descriptions associated with the one or more scenes (Rothe, paragraph 35 teaches “the storyboard data including camera angle, lighting mood, and scene composition descriptors”); this is for the frames because is of storyboard data and the scene composition descriptors here show scene description data indicative of scene descriptions associated with the scenes; cinematographic data indicative of one or more cinematographic details associated with the one or more scenes (Rothe, paragraph 33 teaches “target platform specifications; (ii) the storyboard data, including camera angle, visual complexity, and scene-level mood attributes”, and paragraph 8 teaches “storyboarding module may generate visual storyboards, which include camera angles, scene composition, and lighting plans”); camera angle, scene composition and lighting plans shows cinematographic data indicative of cinematographic details associated with the scenes. Rothe is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of scene descriptions and cinematographic details applied to storyboard generation. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Wittman and Hoffe with the cinematographic data techniques of Rothe to improve subsequent video content generation (Rothe, claim 8). This would be done by accounting for the cinematographic details.
Regarding claim 3, the combination of Wittman, Hoffer and Rothe further comprising: generating, for the one or more scenes and based at least on processing at least one or more portions of the text data that are representative of one or more scene descriptions for the one or more scenes, second text data representative of cinematographic information for the one or more scenes, (Hoffer, abstract teaches “segmented…text elements…segmented elements are applied to a machine learning (ML) method that labels…Prompts based on the labels are then applied to a second ML model…prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes).” and Rothe, paragraph 35 teaches “the storyboard data including camera angle, lighting mood, and scene composition descriptors”); prompts (including storyboard and script) shows the prompt generated are second text data (which includes cinematographic information for the scenes since storyboard data is mentioned to have camera angle), this is for the scenes and based on processing portions/segments of text data that represents scene descriptions/descriptors for the scene (since the storyboard data of the comic which is processed is also mentioned to have scene composition descriptors); wherein the one or more storyboard frames further include the cinematographic information for the one or more scenes (Rothe, paragraph 33 teaches “target platform specifications; (ii) the storyboard data, including camera angle, visual complexity, and scene-level mood attributes”, and paragraph 8 teaches “storyboarding module may generate visual storyboards, which include camera angles, scene composition, and lighting plans”); storyboard frames/data (when viewed in combination) including camera angle, scene composition and lighting plans shows cinematographic information for the scenes. The same motivations used in claim 19 apply here in claim 3.
Regarding claim 4, the combination of Wittman, Hoffer and Rothe teaches wherein the generating of the image data representing the one or more images depicting the one or more visual representations corresponding to the one or more scenes is further based at least on the cinematographic information for the one or more scenes (Rothe, paragraph 33 teaches “target platform specifications; (ii) the storyboard data, including camera angle, visual complexity, and scene-level mood attributes”, and paragraph 8 teaches “storyboarding module may generate visual storyboards, which include camera angles, scene composition, and lighting plans”); camera angle, scene composition and lighting plans shows cinematographic data indicative of cinematographic details associated with the scenes, therefore when viewed in combination with above references, scene including images and storyboard generated including scenes shows image data (images) depicting visual representation of scenes being generated and this would be based on cinematographic information since that information is of the scene. The same motivations used in claim 19 apply here in claim 4.
Regarding claim 5, the combination of Wittman, Hoffer and Rothe teaches wherein the scene information includes one or more scene descriptions associated with the one or more scenes, (Rothe, paragraph 35 teaches “the storyboard data including camera angle, lighting mood, and scene composition descriptors”); the scene composition descriptors here show that scene information from above would include scene description data indicative of scene descriptions associated with the scenes; the one or more scene descriptions including at least one of: action information associated with the one or more scenes; setting information associated with the one or more scenes; contextual information associated with the one or more scenes; tone information associated with the one or more scenes (Rothe, paragraph 35 teaches “include metadata describing spatial orientation, emotional tone, animation affordances, or brand-consistency settings”); this shows scene description would include tone and setting information both associated with scenes; or audio information associated with the one or more scenes (Rothe, paragraph 21 teaches “structured script file may specify…audio”); this shows audio information associated with scenes. The same motivations used in claim 19 apply here in claim 5.
Regarding claim 6, the combination of Wittman, Hoffer and Rothe teaches wherein the scene information includes one or more cinematographic details associated with the one or more scenes, (Rothe, paragraph 33 teaches “target platform specifications; (ii) the storyboard data, including camera angle, visual complexity, and scene-level mood attributes”, and paragraph 8 teaches “storyboarding module may generate visual storyboards, which include camera angles, scene composition, and lighting plans”); camera angle, scene composition and lighting plans shows cinematographic details associated with the scenes; the one or more cinematographic details including at least one of: camera shot information associated with the one or more scenes (Rothe, paragraph 32 teaches “may parse each beat to infer one or more visual elements such as framing (e.g., wide shot, medium shot,”); this shows camera shot information such as wide or medium being associated with each scene ; camera angle information associated with the one or more scenes (Rothe, paragraph 8 teaches “storyboarding module may generate visual storyboards, which include camera angles); or camera movement information associated with the one or more scenes (Rothe, paragraph 32 teaches “may parse each beat to infer one or more visual elements such as… camera movement”). The same motivations used in claim 19 apply here in claim 6.
Regarding claim 12, the combination of Wittman, Hoffer and Rothe teaches the one or more processors further to: generate, using one or more second language models and based at least on the respective portions of the text data, second text data representing scene descriptions corresponding to the one or more scenes, (Rothe, paragraph 19 teaches “LLM may parse natural language user input and place desired video content specifications into a structured format. In one example, structured format containing desired video content specifications is JSON. Exemplary JSON of desired video content specifications is shown below” and paragraph 32 teaches “storyboarding module may generate a sequence of visual keyframes or scene descriptors”); this shows using second LLM (second when viewed in combination and based on segmented elements which are respective portions of text data since that would be user input from Hoffer) to generate second/JSON text data which is for storyboard therefore represents the scene descriptions/descriptors of storyboard corresponding to the scene; wherein the generation of the one or more storyboard frames using the one or more machine learning models is further based at least on the second text data (Rothe, paragraph 32 teaches “the storyboarding module may utilize a trained generative model (e.g., an image diffusion model or a 3D scene sketch generator) to output provisional visual renderings or textual scene descriptors encoded in a machine-readable format (e.g., JSON or equivalent), where each storyboard panel may specify fields”); since storyboard module uses the JSON/second text here, the aforementioned generation of storyboard frames using ML from the combination above would be based on the second/JSON text when viewed in combination. The same motivations used in claim 19 apply here in claim 12.
Regarding claim 13, the combination of Wittman, Hoffer and Rothe teaches the one or more processors further to: generate, using one or more second machine learning models and based at least on the respective portions of the text data, second text data representing cinematographic information corresponding to the one or more scenes, (Hoffer, abstract teaches “segmented…text elements…segmented elements are applied to a machine learning (ML) method that labels…Prompts based on the labels are then applied to a second ML model…prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes).” and Rothe, paragraph 35 teaches “the storyboard data including camera angle, lighting mood, and scene composition descriptors”); this shows using second ML (and based on segmented elements which are respective portions of text data) to generate prompt (including storyboard and script) which means the prompts generated are second text data (prompt which includes cinematographic information for the scenes since storyboard data is mentioned to have camera angle); wherein the generation of the one or more storyboard frames using the one or more machine learning models is further based at least on the second text data (Hoffer, paragraph 97 teaches “the training data can include storyboards and prompts generated from the storyboards, and the films generated based on these storyboards can be used as the outputs and corresponding to the storyboards and prompts generated therefrom”); since training data includes both storyboard and prompts (second text data), the aforementioned generation of storyboard frames using ML models from the combination above is based on the second text data/prompts. The same motivations used in claim 19 apply here in claim 13.
prompts
Regarding claim 14, the combination of Wittman, Hoffer and Rothe teaches the one or more processors further to: generate, using one or more second machine learning models and based at least on second text data representing at least one of scene descriptions or cinematographic information corresponding to the one or more scenes, (Hoffer, abstract teaches “segmented…text elements…segmented elements are applied to a machine learning (ML) method that labels…Prompts based on the labels are then applied to a second ML model…prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes).” and Rothe, paragraph 35 teaches “the storyboard data including camera angle, lighting mood, and scene composition descriptors”); this shows using second ML (and based on input of prompts thereof which are second text data ) to generate images after having segmented text input (which means the prompts generated after are second text data) and this is also based on cinematographic information for the scenes since storyboard data (which prompt includes) is mentioned to have camera angle;
one or more images depicting one or more visual representations corresponding to the one or more scenes, (Wittman,paragraph 21 teaches “the storyboard including graphical display of one or more scenes of the set of scenes”, paragraph 103 teaches “A scene can include one or more headlines, text, images” and paragraph 186 teaches “storyboard 600 that is generated using AI in accordance with implementations of the present disclosure. The example of FIG. 6 is generated based on the example user input introduced above (e.g., As the Head of the new technologies organization at ACME, . . . ) and includes scenes 602, 604, 606, 608, 610”); scene including images and storyboard generated including scenes shows image data (images) depicting visual representation of scenes being generated and these are corresponding/of the scene; wherein the generation of the one or more storyboard frames is further based at least on the one or more images (Wittman, paragraph 186 teaches “storyboard 600 that is generated using AI in accordance with implementations of the present disclosure. The example of FIG. 6 is generated based on the example user input introduced above (e.g., As the Head of the new technologies organization at ACME, . . . ) and includes scenes 602, 604, 606, 608, 610”); this shows storyboard (and aforementioned frames thereof) generated including scenes which show the images, therefore, the storyboard is based on the images. The same motivations used in claim 19 apply here in claim 14.
Regarding claim 17, the combination of Wittman, Hoffer and Rothe teaches the one or more processors further to: obtain, from the one or more client devices, input data indicating a request to update at least one of one or more scene descriptions, one or more cinematographic details, or one or more visual representations associated with the one or more storyboard frames (Rothe, paragraph 8 teaches “structured script may be further refined by a script refinement module. The script refinement module may make modifications to the script based on user specified tone, pacing, and genre specific elements” and paragraph 21 teaches “structured script file may specify characters, dialogue, camera angles, set details, audio, animation and more”); since script (which is associated with storyboard frames in Hoffer abstract) has camera angles which are considered cinematographic details, refining/modifying/updating script would also update the cinematographic details, and one of ordinary skill in the art would understand that input data indicating a request (which typically comes from user thus obtained from client device) would be needed to perform this refinement/modification/update; and generate, using the one or more machine learning models and based at least on one or more parameters included in the input data, one or more updated versions of the one or more storyboard frames (Rothe, claim 1 teaches “generating, by a storyboarding module, a storyboard associated with the structured script file… auditory component … generating, by a post-production module, a plurality of modified video sequences wherein each modified video sequence differs from the others and the intermediate video sequence in at least one visual or auditory characteristic”); this shows generating modified/updated version of video (storyboard frames from combination above), is based on parameters included in input such as auditory or visual characteristics and when viewed in combination, would use the ML from Hoffer abstract to process inputs. The same motivations used in claim 19 apply here in claim 17.
Regarding claim 20, the combination of Wittman, Hoffer and Rothe teaches wherein the one or more processors are comprised in at least one of:
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing deep learning operations;
a system for performing remote operations (Wittman, paragraph 202 teaches “computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a network, such as the described one”); this shows system for performing remote operations;
a system for performing real-time streaming (Wittman, paragraph 192 teaches “executed for real-time generation of videos”); real-time generation would also mean the system is capable of real-time streaming;
a system for generating or presenting one or more of augmented reality content,
virtual reality content, or mixed reality content;
a system implemented using an edge device;
a system implemented using a robot;
a system for performing conversational AI operations;
a system implementing one or more multi-model language models;
a system implementing one or more large language models (LLMs) (Wittman, paragraph 44 teaches "the conceptual architecture 200 includes a LLM system 230”); this shows system implementing LLMs;
a system implementing one or more small language models (SLMs);
a system implementing one or more vision language models (VLMs);
a system for generating synthetic data;
a system for generating synthetic data using AI;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Wittman, paragraph 39 teaches “video generation platform is provided as a cloud-based platform that can be provisioned in any appropriate cloud runtime. An example cloud runtime includes, without limitation, the SAP Kyma runtime (SKR), which is provided by SAP AG of Walldorf, Germany, and can be described as a fully managed Kubernetes-based runtime. Another example cloud runtime includes Cloud Foundry”); this shows system implemented using cloud computing resources.
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wittman in view of Hoffer as applied to claim 8 above, and further in view of Libert et al. (U.S. Patent Application Publication No. 2010/0050080), hereinafter referenced as Libert.
Regarding claim 16, the combination of Wittman, and Hoffer teaches the one or more processors further to: generate, using the one or more machine learning models and based at least on the respective portions of the text data, one or more second storyboard frames corresponding to one or more updated versions of the one or more scenes (Wittman, paragraph 78 teaches “includes a storyboard generator 400, a layout selector 402, a data extractor/injector 404, and one or more tools 406. In some examples, the storyboard generator 400 interacts with a LLM system (e.g., the LLM system 360 of FIG. 3) to provide a storyboard based on user input. In some examples, the storyboard includes a set of scenes to be depicted in a video… determine a layout for each scene of the storyboard.” and paragraph 5 teaches “revisions to an abstract displayed with a scene of the storyboard and in response, providing a modified scene using the LLM system responsive to the revisions”); for the storyboard to be created (which uses ML in Hoffer abstract and is based on respective segmented/portions of text data) that includes scenes depicted in a video means that a generation of storyboard frames would occur (inclusive of second storyboard frames), and these would be updated versions of storyboard frames due to the revision of scene of storyboard.
However, the combination of Wittman, and Hoffer fails to teach and obtain, from the one or more client devices, input data indicating a selection of the one or more storyboard frames over the one or more second storyboard frames, wherein the sending of the one or more storyboard frames to the one or more client devices is based at least on the selection
However, Libert teaches
and obtain, from the one or more client devices, input data indicating a selection of the one or more storyboard frames over the one or more second storyboard frames, (Libert, paragraph 10 teaches “storyboard user interface can be modified by selecting and dragging a video frame displayed by the clip player user interface to the storyboard user interface and dropping the selected video frame onto the displayed storyboard.”); this shows selecting and dragging (obtaining input data indicating selection from user/client device thereof) storyboard frame over another second storyboard frame; wherein the sending of the one or more storyboard frames to the one or more client devices is based at least on the selection (Libert, claim 2 teaches “a storyboard view renderer to render a storyboard user interface that displays a sequence of thumbnail images of selected frames of the media asset and allows user manipulation and editing of the storyboard” and claim 7 teaches “thumbnail images of a media asset displayed on the storyboard user interface are graphical objects that can be selected…to change the thumbnail icon off the media asset to the selected storyboard image”); this shows sending storyboard frame to client device for display as thumbnail is based on the selection. Libert is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of selection of storyboard frame over other/second storyboard frames for client device(s). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Wittman, and Hoffer with the storyboard frame selection techniques of Libert to enable users to search and organize content, and add user-definable metadata, frame-accurate location and video editing, and the ability to select frames in a video asset as party of thumbnail representation or part of a storyboard (Libert paragraph 27). This increases user engagement and experience by adding a selectivity/customization aspect to the invention.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Edson (U.S. Patent Application Publication No. 2024/0273796) fig. 2 teaches animated image file being generated by using user input alongside text instructions at text prompt module 206 leading to storyboard being generated via storyboard generator module 210 and paragraph 91 teaches “submitting a prompt to a model, such as a large language model (LLM)”.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAUMAN U AHMAD whose telephone number is (703)756-5306. The examiner can normally be reached Monday - Friday 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/N.U.A./Examiner, Art Unit 2611
/KEE M TUNG/Supervisory Patent Examiner, Art Unit 2611