DETAILED ACTION
The action is in response to the original filing on September 17, 2024 and the Remarks and Amendments filed on July 24, 2026. Claims 1-20 are pending and have been considered below. Claims 1, 2, 12, 18, and 19 are amended accordingly.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-9 are rejected under 35 U.S.C. 103 as being unpatentable over Mohamed Amer (Pat. Pub. US-10789755-B2, herein after “Amer”) in view of Kazi Habib Et. Al. (Pat. Pub. US-20230281904-A1, herein after “Habib”), Christopher Corwin (Pat. Pub. US-20210150263-A1, herein after “Corwin”) and further in view of Alex Edson (Pat. Pub. US-20240273796-A1, herein after “Edson”).
In regard to claim 1, Amer teaches [a] method comprising:
receiving, by a processing device, inputs including a description of a digital animation “Computing system 110 may process input 104” (Amer, ¶ [0021]), a digital image, and at least one object “A command might be a request to create an object or an actor to be included within the story. A story description might describe attributes of a scene or event, or describe an event or sequence of events within the story (e.g., “Jane called Sam using the new mobile phone”).” (Amer, ¶ [0021]);
selecting, by the processing device, one or more prompt portions based on the description “the prompt may relate to an ambiguity in a scene (or story) identified by computing system 110 that has not been resolved by other information received by computing system 110” (Amer, ¶ [0022]) where a portion of the prompt is selected based on ambiguity in the description, additionally, “ graph module 322 may identify an actor (901A) mapped to the “who” portion of the event frame, identify an action (901B) mapped to the “did what” portion of the event frame, identify an object (901C) mapped to the “to whom” portion …” (Amer, ¶ [0170]) where the animation is built by description, and if a portion of the description is absent, the computing system selects said portion for clarity based on the lacking description;
forming, by the processing device, a prompt based, at least in part, on the selected one or more prompt portions; “interaction module 108 may engage in an information exchange (e.g., a “conversational AI” exchange) whereby interaction module 108 collects input 104 from a user of user interface device 170, and based on the input 104, may generate output 106, prompting a user of user interface device 170 for further information … interaction module 108 may generate output 106 which may include a prompt for information about one or more aspects of a scene or story” (Amer, ¶ [0022]) where the “further information” requested is in regard to the selected missing portion of the description;
generating, by the processing device using one or more machine-learning models, animation settings from the prompt formed from the selected one or more prompt portions, the one or more machine-learning models including a large language model (LLM) “System 100A may ultimately generate animation 180 from data structure 151. For instance, in the example of FIG. 1A, machine learning module 109 outputs data structure 151 to animation generator 161” (Amer, ¶ [0023]);
outputting, by the processing device, the digital animation using the animation settings as animating the at least one object “Machine learning module 109 identifies a response to the query, and causes interaction module 108 to send output 106 to device 170” (Amer, ¶ [0025]).
Amer fails to explicitly teach the one or more machine-learning models including a large language model (LLM);
calculating, by the processing device, a path based on segmenting the digital image;
generating, by the processing device, one or more vector outlines based on the at least one object and motion of the at least one object derived from the animation settings;
outputting, by the processing device, the digital animation using the animation settings as animating the at least one object based on the path and the one or more vector outlines with respect to the digital image.
Habib teaches calculating, by the processing device, a path based on segmenting the digital image ““animation canvas” refers to a graphical user interface element for displaying and/or editing a digital animation and/or an animation frame. For instance, the term “animation canvas” includes an interface for generating and/or manipulating digital design objects, drawing animation paths, generating animation frames and animation layers, etc. In one or more embodiments, the animation management system generates animation frames, animation layers, and digital animations based on user interactions with an animation canvas.” (Habib, ¶ [0037]) where the digital image is provided by the user, including a line or geometry provided which the system uses to calculate an animation sequence to follow said line; and
outputting, by the processing device, the digital animation using the animation settings as animating the at least one object based on the path and the one or more vector outlines with respect to the digital image “the animation management system animates digital design objects along animation paths in one or more animation layers” (Habib, ¶ [0017]).
PNG
media_image1.png
825
441
media_image1.png
Greyscale
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of providing prompts regarding an animation to an artificial intelligence taught by Amer, with the method of an animation line taught by Habib to provide an artificial intelligence that can communicatively provide an animation for an object, which follows an animation line. The motivation to do so would be to easily produce an animation of a provided object via artificial intelligence.
Amer in view of Habib fail to teach the one or more machine-learning models including a large language model (LLM);
generating, by the processing device, one or more vector outlines based on the at least one object and motion of the at least one object derived from the animation settings; and
animating the at least one object based on the path and the one or more vector outlines.
Corwin teaches generating, by the processing device, one or more vector outlines based on the at least one object and motion of the at least one object derived from the animation settings “SVG allows three types of graphic objects: vector graphic shapes such as paths and outlines consisting of straight lines and curves, bitmap images, and text. Graphical objects can be grouped, styled, transformed and composited into previously rendered objects” (Corwin, ¶ [0117]) where a vector outline can be made through the use of scalable vector graphics, vector outlines can be produced from previously rendered objects; and
animating the at least one object based on the path and the one or more vector outlines “SVG drawings can be interactive and can include animation, defined in the SVG XML elements or via scripting that accesses the SVG Document Object Model (DOM)” (Corwin, ¶ [0117]) where scalable vector graphics can include animation, which are based on the provided objects.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of providing prompts regarding an animation to an artificial intelligence and animation line taught by Amer and Habib with the use of a vector outline based on the object for the object to follow taught by Corwin to provide an artificial intelligence that can communicatively provide an animation for an object, which follows an animation line. The motivation to do so would be to easily produce an animation of a provided object via artificial intelligence.
Amer in view of Habib and Corwin fail to efficiently teach the one or more machine-learning models including a large language model (LLM).
Edson teaches the one or more machine-learning models including a large language model (LLM) “the model-generated questions may be synthesized by submitting a prompt to a model, such as a large language model (LLM), comprising of two pieces of text: (1) the user's text prompt, and (2) a natural-language instruction such as “Generate one or more questions, e.g., clarifying questions, using the information by the user describing their request for an animated image.”” (Edson, ¶ [0091]) where an LLM can handle animation generation and prompt management.
Edson additionally teaches one or more vector outlines based on the at least one object and motion of the at least one object derived from the animation settings “The animated image file may include a sequence of images that depicts a visual element of the images as changing or moving over time. Examples of the animated image file include animated scalar vector graphics may include animated scalar vector graphics” (Edson, ¶ [0031]) where scalable vector graphics (SVG) use vector lines, or paths, to give an object motion. Additionally, “By comparing the segmented images of each extracted frame against one another, the model may track the path or difference, e.g., angular, translational, rotational, etc., of individual objects contained within the images” (Edson, ¶ [0086]) where the path is the motion of the object.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of providing prompts regarding an animation to an artificial intelligence and animation line taught by Amer and Habib with the use of an LLM and vector outline based on the object for the object to follow taught by Edson to provide an LLM that can communicatively provide an animation for an object, which follows an animation line. The motivation to do so would be to easily produce an animation of a provided object via artificial intelligence. The Examiner would like to additionally note that the pertinent art cited in the conclusion section also pertains to animated object movement.
In regard to claim 2, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 1, wherein the selected one or more prompt portions includes one or more prompt portions describing one or more animation elements, a duration of the digital animation, and description of an environment of the digital animation “ behavior analytics systems 314 may be, or may include, a camera, a sensor, or a set of cameras or sensors, or any appropriate type of image acquisition or other device enabling detection of movements, posture, poses, gestures, actions, positions, and/or gaze orientations, directions, durations, and persistence” (Amer, ¶ [0100]) where a sensor may capture animation elements including poses, movements, and durations.
Edson additionally teaches prompt portions describing one or more animation elements, and description of an environment of the digital animation “the text instructions may include a natural-language description of a modification to a visual element of the selected image. For example, the user may select an image of a cat typing on a computer, and input text instructions such as, “Change the background to blue-colored pixel art and put an astronaut helmet on the cat.”” (Edson, ¶ [0036]) where portions of the prompt are parsed to identify the subject, action, and background, which is read as the environment.
In regard to claim 3, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 1, wherein the generating of the animation settings is performed using the one or more machine-learning models configured as a large language model (LLM) “aspects of this disclosure relate to techniques for training a machine learning system to understand narrative and context, and employ both explicit and implicit knowledge about a scene …” (Amer, ¶ [0016]).
In regard to claim 4, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 1, wherein the generating of the animation settings is performed based on the prompt “computing device 370 detects input that corresponds to a search query entered by a user of device 370. Device 370 updates user interface 775A to echo, within text box 711 ... In the example illustrated, the text corresponds to a “Walk on uneven terrain” query, … computing system 310 detects a signal that search module 325 determines corresponds to input that includes the “Walk on uneven terrain” search query. Search module 325 uses the search query to generate one or more animations (e.g., animation 380) that corresponds to or is representative of the text query” (Amer, ¶ [0149]) and the path “the animation management system receives user input indicating animation frames and a path for those animation frames” (Habib, ¶ [0024]).
In regard to claim 5, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 1, wherein the generating of the animation settings includes generating animation semantics of the digital animation “a parser generates, based on textual scene information 112 and/or user input 114, a semantic representation structured around events and their various arguments” (Amer, ¶ [0058]).
Amer fails to teach using the one or more machine-learning models and wherein the calculating of the path is based at least in part on the animation semantics.
Habib teaches using the one or more machine-learning models and wherein the calculating of the path is based at least in part on the animation semantics “the animation path manager 706 generates, modifies, and implements animation paths based on user interactions” (Habib, ¶ [0147]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating animation semantics taught by Amer, with the method of an animation path management system taught by Habib to provide a machine learning model that can calculate the animation path based on the amination semantics. The motivation to do so would be to easily produce an animation of a provided object via artificial intelligence.
In regard to claim 6, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 1, wherein the animation settings select the digital animation from a plurality of preset animation options “Animation module 324 may include a library of assets enabling rapid assembly of a scene, and may enable a user (e.g., user 371) to select characters, customize their appearance, position cameras and lights around the set to prepare for the scene's key shots, issue directions to the actors, and record on-screen action in near or seemingly-near real time. Animation module 324 may include mechanisms for associating attributes to objects, actors, and props including lighting effects and mood-based modifiers” (Amer, ¶ [0112]) where a plurality of preset animation options could be understood as a library of assets which are both intended for rapid assembly of an animation.
In regard to claim 7, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 6, wherein the animations settings specify a subject as the at least one object “Animation agent 166 may process composition graph 152 to identify events occurring in the script or commands. In some examples, animation agent 166 may generate a data structure, such as an “event frame,” that uses the output of language parser 118 (or another module) to identify subject, predicate, object, and time and place (when available)” (Amer, ¶ [0066]), an entity corresponding to the path “the animation management system 102 determines an initial position of the digital design object corresponding to the animation frame 544b.” (Habib, ¶ [0108]) where the digital design object is an entity, and the path is each animation frame as the object changes its position, and a duration for output of the digital animation “Confidence bar 725 is also shown within video viewing area 720, and may represent a time span of the duration of the video” (Amer, ¶ [0153]).
In regard to claim 8, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 1
Amer in view of Habib fail to teach wherein the calculating includes:
forming at least one mask by segmenting the digital image using at least one machine-learning model;
converting the at least one mask into a vector outline; and
generating the path based on the vector outline.
Corwin teaches wherein the calculating includes:
forming at least one mask by segmenting the digital image using at least one machine-learning model “Graphical objects can be grouped, styled, transformed and composited into previously rendered objects. The feature set includes nested transformations, clipping paths, alpha masks, filter effects and template objects. SVG drawings can be interactive and can include animation, defined in the SVG XML elements or via scripting that accesses the SVG Document Object Model (DOM)” (Corwin, ¶ [0117]) where graphics objects are able to be segmented and masked in a number of ways, such as those available through scalable vector graphics and other similar web design techniques;
converting the at least one mask into a vector outline “The user system 112 includes various processors, ... The various processors, processing systems and processing modules implemented by the user system 112 can include: a Rapid Pictorial Demonstration (RPD) Controller/Editor 120, a browser 124 having a rendering engine 126, a dynamic Scalable Vector Graphics (SVG) system 150, and an automatic SVG animation system 160” (Corwin, ¶ [0053]) where it is known that using scalable vector graphics can allow for vectors to be created from masks; and
generating the path based on the vector outline “SVG allows three types of graphic objects: vector graphic shapes such as paths and outlines consisting of straight lines and curves, bitmap images, and text. Graphical objects can be grouped, styled, transformed and composited into previously rendered objects. The feature set includes nested transformations, clipping paths, alpha masks, filter effects and template objects” (Corwin, ¶ [0117]) where if the outline is generated, a path can be made based on the outline using SVG.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an animation for an object that follows an object path using AI taught by Amer in view of Habib with the method using masking and vector outlines to find and create a path taught by Corwin to provide a machine learning model that can produce an artificial intelligence which can use vector outlines to create an animation path for an object. The motivation to do so would be to easily produce an animation which can follow a path of a provided object via artificial intelligence.
In regard to claim 9, Amer in view of Habib, Corwin, and Edson teach [t]he method as described in claim 1, wherein the prompt includes:
an environment prompt portion establishing an environment, in which, the digital animation is to be output “Dialogue manager 122 may collect additional information about the scene from a user of user interface device 170. In some examples, one or more instances of user input 114 may be in the form of commands (or queries), or may be responses to queries posed by dialogue manager 122 (e.g., through output 116), or may otherwise be information about a scene. For instance, dialogue manager may detect input corresponding to a request to move an actor from one location to another, or to adorn an actor with a hat, or to animate an actor to walk toward another actor” (Amer, ¶ [0056]) where information about a scene or setting is interpreted as an environment;
an animation elements prompt portion specifying that a subject and a path are to be used as part of the digital animation “as shown in FIG. 5C … the animation management system 102 determines a modified number of instances for the digital design object 506 and generates the modified number of instances of the digital design object 506 for the digital animation. For example, because the user interaction shown in FIG. 5C increases the selected instances, the animation management system 102 increases the number of instances of the digital design object 506 within the global animation frames 501c. More specifically, as shown in the global animation frames 501c, the animation management system 102 generates four instances of the digital design object. Moreover, the animation management system 102 generates the digital animation so that each instance appears to enter the scene and follow along the animation path 505” (Habib, ¶ [0113]) a path and an animated subject are part of the digital animation;
a description variants prompt portion describing ways in which the description is usable to describe the digital animation “Casting agent 164 may generate actors that match the character descriptions in a script or other textual description, or in other input. For instance, in one example, language parser 118 may parse text such as “John wore a black hat and a tuxedo with a white shirt.” Language parser 118 identifies a set of “attire” tags for clothing. Composition graph module 121 interprets these descriptions as character attributes and includes such attributes as part of composition graph 152” (Amer, ¶ [0064]) where features of the digital animation and their use may be described;
a duration prompt portion specifying a duration for output of the digital animation “behavior analytics systems 314 may be, or may include, a camera, a sensor, or a set of cameras or sensors, or any appropriate type of image acquisition or other device enabling detection of movements, posture, poses, gestures, actions, positions, and/or gaze orientations, directions, durations, and persistence” (Amer, ¶ [0153]) where the duration may be tracked or specified by the behavior analytics system;
a task prompt portion specifying a task that the one or more machine-learning models is to undertake to discern the subject and the path and an output format of the digital animation “Many of the skills and practices used in interactive storytelling may benefit from assistance by artificial intelligence or artificially-intelligent agents or other systems. Often, the interaction between humans and agents or machines has focused on handling directives given by the human user” (Amer, ¶ [0015]) where tasks can be given to an AI to command an action to be performed, additionally, Habib teaches of moving a subject or object across a path in ¶ [0037];
an error handling portion “Computing system 310 may identify an ambiguity” (Amer, ¶ [0162]) specifying error message generation in response to inaccuracy of the description as including at least one corrective action “In response to identifying an ambiguity, computing system 310 may prompt a user for information (805)” (Amer, ¶ [0163]) additionally, “Computing system 310 may resolve the ambiguity based on the user's response (806)” (Amer, ¶ [0164]); and
an examples prompt portion including examples of the inputs and corresponding animation settings “aspects of this disclosure relate to AI-assisted generation of animations or videos for the purpose of facilitating interactive storytelling, such AI-assisted generation of animations can be applied in other contexts as well, and may be used in applications that involve human motion generally, and include applications involving interactions in visual and physical worlds, robotics, animation, surveillance, and video search.” (Amer, ¶ [0015]) where the artificial intelligence may communicate with the user to improve its output in different aspects.
Claims 10 & 11 are rejected under 35 U.S.C. 103 as being unpatentable over Amer in view of Habib, Corwin, Edson, and further in view of Bay Raitt (Pat. Pub. US-20120021828-A1, herein after “Raitt”).
In regard to claims10, Amer in view of Habib and Edson teach [t]he method as described in claim 1
Amer in view of Habib do not explicitly teach wherein the prompt includes a preset options prompt portion references a plurality of preset animation options that are available for digital animation generation.
Raitt teaches a plurality of preset animation options that are available for digital animation generation “The development subsystem may be used to generate one or more animation presets that are to be used in subsequent animation editing. The editing subsystem may receive one or more animation presets, provide an interface that enables an animator to specify a preset and a way of integrating the preset with the animation that is composed of recorded multi-dimensional video game world data, and automatically apply the preset to generate one or more frames of an animation that subsequently modifies the resultant visual display” (Raitt, ¶ [0041]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an object animation with artificial intelligence taught by Amer, with the method of an animation preset taught by Raitt to provide a machine learning model that can provide preset animation options to an uploaded object. The motivation to do so would be to allow for quick and simple animations to be applied to objects.
In regard to claim 11, Amer in view of Habib and Edson teach [t]he method as described in claim 1
Amer in view of Habib do not teach wherein the digital animation specifies a z-order of the at least one object in relation to an additional object such that the path of the at least one object passes before and behind the additional object.
Raitt teaches wherein the digital animation specifies a z-order of the at least one object in relation to an additional object such that the path of the at least one object passes before and behind the additional object “Besides the data used for creating the images and sounds that are captured, other data dimensions representing game state information such as motion, collision information, wireframe/skeleton data, timestamps, z-order of objects, and other such information may also be captured or extracted and stored for creating the new scene shot in a compositing cycle” (Raitt, ¶ [0100]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an object animation with artificial intelligence taught by Amer, with the method of an animation z-order taught by Raitt to provide a layering where objects closest to view are visible. The motivation to do so would be to allow for objects to pass by one another realistically.
Claims 12 & 19 are rejected under 35 U.S.C. 103 as being unpatentable over Amer in view of Edson and Ravi Gogna Et. Al. (Pat. Pub. US-12005925-B1, herein after “Gogna”).
In regard to claim 12, Amer teaches [a] computing device comprising:
a processing device “… this disclosure describes a system comprising a storage device; and processing circuitry having access to the storage device …” (Amer, ¶ [0013]); and
a computer-readable storage medium storing instructions that, responsive to execution by the processing device, causes the processing device to perform operations comprising “this disclosure describes a computer-readable storage medium comprising instructions that, when executed, configure processing circuitry of a computing system” (Amer, ¶ [0014]):
forming a prompt having text based on a description of a digital animation by combining one or more prompt portions “interaction module 108 may engage in an information exchange (e.g., a “conversational AI” exchange) whereby interaction module 108 collects input 104 from a user of user interface device 170, and based on the input 104, may generate output 106, prompting a user of user interface device 170 for further information … interaction module 108 may generate output 106 which may include a prompt for information about one or more aspects of a scene or story” (Amer, ¶ [0022]);
generating, using one or more machine-learning models including a large language model “System 100A may ultimately generate animation 180 from data structure 151. For instance, in the example of FIG. 1A, machine learning module 109 outputs data structure 151 to animation generator 161” (Amer, ¶ [0023]), animation settings of one or more candidate paths based on the prompt formed from combining the one or more prompt portions, “System 100A may ultimately generate animation 180 from data structure 151. For instance, in the example of FIG. 1A, machine learning module 109 outputs data structure 151 to animation generator 161” (Amer, ¶ [0023]) the animation settings identifying the digital animation from a plurality of preset animation options “Animation module 324 may include a library of assets enabling rapid assembly of a scene, and may enable a user (e.g., user 371) to select characters, customize their appearance, position cameras and lights around the set to prepare for the scene's key shots, issue directions to the actors, and record on-screen action in near or seemingly-near real time. Animation module 324 may include mechanisms for associating attributes to objects, actors, and props including lighting effects and mood-based modifiers” (Amer, ¶ [0112]) where a plurality of preset animation options could be understood as a library of assets which are both intended for rapid assembly of an animation; and
outputting the digital animation using animation settings and the selected candidate path “Machine learning module 109 identifies a response to the query, and causes interaction module 108 to send output 106 to device 170” (Amer, ¶ [0025]) where the output includes the animation details.
Amer fails to teach where the animation settings are of one or more candidate paths
receiving a selection of a candidate path from the one or more candidate paths as displayed in a user interface; and
outputting the digital animation using animation settings and the selected candidate path.
Gogna teaches one or more candidate paths “multiple candidate paths may exist to navigate past an obstacle (e.g., such as an object in the road or a construction zone) and the sensor data of the environment may not provide sufficient information or context for the vehicle to confidently identify a preferred path” (Gogna, ¶ [0007]) where multiple candidate paths can be generated using artificial intelligence that receives an input to a system for determining optional paths,
receiving a selection of a candidate path from the one or more candidate paths as displayed in a user interface “the teleoperator may be able to quickly and efficiently select a path while monitoring multiple different vehicles within the fleet” (Gogna, ¶ [0075]) where the user interface displays the optional paths to the user and the user may select a path; and
outputting the digital animation using animation settings and the selected candidate path “presentation of the one or more candidate paths may comprise causing display of an animation of the vehicle executing a candidate path” (Gogna, ¶ [0074]).
PNG
media_image2.png
773
359
media_image2.png
Greyscale
Gogna, Figs. 5A – 5C, depicts autonomous vehicle (item 500) traversing candidate paths (items 504 and 506) after user selection with an animation.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an object animation with artificial intelligence taught by Amer, with the use of animated candidate paths taught by Gogna to produce an object animation, generated by artificial intelligence, that can produce object animation candidate paths for a user to choose from based on a prompt. The motivation to do so would be to allow the user to choose from multiple generated paths for their object animation to follow.
Amer in view of Gogna fail to explicitly teach combining one or more prompt portions; and
one or more machine-learning models including a large language model.
Edson teaches combining one or more prompt portions “the description is “a cat typing on a keyboard wearing astronaut helmet, pink-colored pixel art,” as shown in FIG. 4N. It is to be understood that other details deemed relevant by the user may also be included. … In causing the animated image file to be generated, a multi-modal prompt inclusive of the user-selected emoji and natural-language text prompt may be communicated to a generative AI model to generate the image using the selected theme” (Edson, ¶ [0071]) where multiple subjects, or portions of the prompt are put together to create the final image, additionally “The user may also elect … modify the animated image file by selecting a “Modify” button, as depicted in FIG. 4O. In this case, the user selects the “Modify” button, enabling the user to input natural-language instructions to be sent to an AI model for modifying the animated image file” (Edson, ¶ [0071]) where further prompts can be added to include more features, which is read as combining prompt portions since no limiting method is specified in the claim; and
one or more machine-learning models including a large language model “the model-generated questions may be synthesized by submitting a prompt to a model, such as a large language model (LLM), comprising of two pieces of text: (1) the user's text prompt, and (2) a natural-language instruction such as “Generate one or more questions, e.g., clarifying questions, using the information by the user describing their request for an animated image.”” (Edson, ¶ [0091]) where an LLM is used for creating the animation and its settings based on the prompt.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of providing prompts regarding an animation to an artificial intelligence and candidate animation paths taught by Amer and Gogna with the use of an LLM and prompt portions that are generated together to create a final animation taught by Edson to provide an LLM that can be prompted to generate an animation for an object, which follows a chosen animation line. The motivation to do so would be to create an animation through multiple prompts via artificial intelligence with the ability for the user to decide the animation path.
In regard to claim 19, Amer teaches [o]ne or more computer-readable storage media storing instructions that (Amer, ¶ [0014]), responsive to execution by a processing device, causes the processing device to perform operations comprising:
receiving inputs including a description of a digital animation “Computing system 110 may process input 104” (Amer, ¶ [0021]) where the computing systems purpose is for creating an animation, a digital image, and at least one object “A command might be a request to create an object or an actor to be included within the story. A story description might describe attributes of a scene or event, or describe an event or sequence of events within the story (e.g., “Jane called Sam using the new mobile phone”).” (Amer, ¶ [0021]);
selecting one or more prompt portions based on the description “the prompt may relate to an ambiguity in a scene (or story) identified by computing system 110 that has not been resolved by other information received by computing system 110” (Amer, ¶ [0022]) where a portion of the prompt is selected based on ambiguity in the description, additionally, “ graph module 322 may identify an actor (901A) mapped to the “who” portion of the event frame, identify an action (901B) mapped to the “did what” portion of the event frame, identify an object (901C) mapped to the “to whom” portion …” (Amer, ¶ [0170]) where the animation is built by description, and if a portion of the description is absent, the computing system selects said portion for clarity based on the lacking description;
forming a prompt based, at least in part, on the selected one or more prompt portions “interaction module 108 may engage in an information exchange (e.g., a “conversational AI” exchange) whereby interaction module 108 collects input 104 from a user of user interface device 170, and based on the input 104, may generate output 106, prompting a user of user interface device 170 for further information … interaction module 108 may generate output 106 which may include a prompt for information about one or more aspects of a scene or story” (Amer, ¶ [0022]), additionally, “When composition graph 328 is missing information, graph module 322 may decide whether to prompt the user for a clarification versus sampling from a probabilistic model” (Amer, ¶ [0110]) where if information is missing, it can be supplied from a probabilistic model, the prompt including selected one or more prompt portions providing context about the digital animation “the present disclosure may enable one or more humans to communicate with AI agents to generate animated movies to tell a story, enable collaborative visualizing of a narrative, and/or enable communication with AI agents to understand and relate both parties' (human and AI agent) perceptions of the world” (Amer, ¶ [0015]) additionally, the probabilistic model in ¶ [0110] of Amer is capable of providing context based on a need for clarification for the selected prompt, input expectations “In a mixed-initiative system, one or more humans and one or more AI agents may operate as a team on such a narrative, in which all collaborators may perform parts of the task at which they respectively excel. Some aspects of this disclosure relate to an AI agent or system generating an animation or video in response to human input and/or human interactions, which is one significant, and perhaps fundamental, task that may be performed in a mixed-initiative system relating to interactive storytelling” (Amer, ¶ [0015]) where the user and AI agents collaborate to build the output;
generating, using one or more machine-learning models including a large language model (LLM), animation settings from the prompt formed from the selected one or more prompt portions, the animation settings describing one or more candidate paths “System 100A may ultimately generate animation 180 from data structure 151. For instance, in the example of FIG. 1A, machine learning module 109 outputs data structure 151 to animation generator 161” (Amer, ¶ [0023]) where animation settings are generated; and
outputting the digital animation using animation settings and based on the selected candidate path as animating the at least one object with respect to the digital image “Machine learning module 109 identifies a response to the query, and causes interaction module 108 to send output 106 to device 170” (Amer, ¶ [0025]) where the output is an animation of the object/character participating in the story.
Amer fails to teach a large language model (LLM), the animation settings describing one or more candidate paths;
receiving a selection of a candidate path from the one or more candidate paths via a user interface; and
outputting the digital animation using animation settings and based on the selected candidate path.
Gogna teaches animation settings based on the one or more candidate paths and the prompt “multiple candidate paths may exist to navigate past an obstacle (e.g., such as an object in the road or a construction zone) and the sensor data of the environment may not provide sufficient information or context for the vehicle to confidently identify a preferred path” (Gogna, ¶ [0007]) where multiple candidate paths can be generated using artificial intelligence that receives an input to a system for determining optional paths;
receiving a selection of a candidate path from the one or more candidate paths via a user interface “the teleoperator may be able to quickly and efficiently select a path while monitoring multiple different vehicles within the fleet” (Gogna, ¶ [0075]) where the user interface displays the optional paths to the user and the user may select a path; and
outputting the digital animation using animation settings and based on the selected candidate path “presentation of the one or more candidate paths may comprise causing display of an animation of the vehicle executing a candidate path” (Gogna, ¶ [0074]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an object animation with artificial intelligence taught by Amer, with the use of animated candidate paths taught by Gogna to produce an object animation, generated by artificial intelligence, that can produce object animation candidate paths for a user to choose from based on a prompt. The motivation to do so would be to allow the user to choose from multiple generated paths for their object animation to follow.
Amer in view of Gogna fail to explicitly teach a large language model (LLM).
Edson teaches a large language model (LLM) “the model-generated questions may be synthesized by submitting a prompt to a model, such as a large language model (LLM), comprising of two pieces of text: (1) the user's text prompt, and (2) a natural-language instruction such as “Generate one or more questions, e.g., clarifying questions, using the information by the user describing their request for an animated image.”” (Edson, ¶ [0091]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an object animation with artificial intelligence including candidate paths for a user to select from taught by Amer and Gogna, with the use of an LLM taught by Edson to produce an object animation, generated by an LLM, that can produce object animation candidate paths for a user to choose from based on a prompt. The motivation to do so would be to allow the user to choose from multiple generated paths for their object animation to follow.
Claims 13-14, 16, 18 & 20 are rejected under 35 U.S.C. 103 as being unpatentable over Amer in view of Edson, Gogna, and Habib.
In regard to claim 13, Amer in view of Edson and Gogna teach [t]he computing device as described in claim 12, wherein the animations settings specify a subject as at least one object “Animation agent 166 may process composition graph 152 to identify events occurring in the script or commands. In some examples, animation agent 166 may generate a data structure, such as an “event frame,” that uses the output of language parser 118 (or another module) to identify subject, predicate, object, and time and place (when available)” (Amer, ¶ [0066]), an entity corresponding to a path “the animation management system 102 determines an initial position of the digital design object corresponding to the animation frame 544b.” (Habib, ¶ [0108]) where the digital design object is an entity, and the path is each animation frame as the object changes its position, and a duration for output of the digital animation “Confidence bar 725 is also shown within video viewing area 720, and may represent a time span of the duration of the video” (Amer, ¶ [0153]).
In regard to claim 14, Amer in view of Edson and Gogna teach [t]he computing device as described in claim 12.
Amer in view of Edson and Gogna fail to teach further comprising calculating a path based on a digital image and wherein the digital animation is based on the path.
Habib teaches further comprising calculating a path based on a digital image ““animation canvas” refers to a graphical user interface element for displaying and/or editing a digital animation and/or an animation frame. For instance, the term “animation canvas” includes an interface for generating and/or manipulating digital design objects, drawing animation paths, generating animation frames and animation layers, etc. In one or more embodiments, the animation management system generates animation frames, animation layers, and digital animations based on user interactions with an animation canvas.” (Habib, ¶ [0037]) where the digital image is provided by the user, including a line or geometry provided which the system uses to calculate an animation sequence to follow said line and wherein the digital animation is based on the path “the animation management system animates digital design objects along animation paths in one or more animation layers” (Habib, ¶ [0017]).
In regard to claim 16, Amer in view of Edson and Gogna [t]he computing device as described in claim 14, wherein the generating of the animation settings includes generating animation semantics of the digital animation using the one or more machine-learning models “a parser generates, based on textual scene information 112 and/or user input 114, a semantic representation structured around events and their various arguments” (Amer, ¶ [0058]).
Amer fails to teach wherein the calculating of the path is based at least in part on the animation semantics.
Habib teaches wherein the calculating of the path is based at least in part on the animation semantics “the animation path manager 706 generates, modifies, and implements animation paths based on user interactions” (Habib, ¶ [0147]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating animation semantics taught by Amer, with the method of an animation path management system taught by Habib to provide a machine learning model that can calculate the animation path based on the amination semantics. The motivation to do so would be to easily produce an animation of a provided object via artificial intelligence.
In regard to claim 18, Amer in view of Edson and Gogna teach [t]he computing device as described in claim 12, wherein the one or more prompt portions include:
an environment prompt portion establishing an environment, in which, the digital animation is to be output “Dialogue manager 122 may collect additional information about the scene from a user of user interface device 170. In some examples, one or more instances of user input 114 may be in the form of commands (or queries), or may be responses to queries posed by dialogue manager 122 (e.g., through output 116), or may otherwise be information about a scene. For instance, dialogue manager may detect input corresponding to a request to move an actor from one location to another, or to adorn an actor with a hat, or to animate an actor to walk toward another actor” (Amer, ¶ [0056]) where information about a scene or setting is interpreted as an environment;
an animation elements prompt portion specifying that a subject and a path are to be used as part of the digital animation “as shown in FIG. 5C … the animation management system 102 determines a modified number of instances for the digital design object 506 and generates the modified number of instances of the digital design object 506 for the digital animation. For example, because the user interaction shown in FIG. 5C increases the selected instances, the animation management system 102 increases the number of instances of the digital design object 506 within the global animation frames 501c. More specifically, as shown in the global animation frames 501c, the animation management system 102 generates four instances of the digital design object. Moreover, the animation management system 102 generates the digital animation so that each instance appears to enter the scene and follow along the animation path 505” (Habib, ¶ [0113]) a path and an animated subject are part of the digital animation;
a description variants prompt portion describing ways in which the description is usable to describe the digital animation “Casting agent 164 may generate actors that match the character descriptions in a script or other textual description, or in other input. For instance, in one example, language parser 118 may parse text such as “John wore a black hat and a tuxedo with a white shirt.” Language parser 118 identifies a set of “attire” tags for clothing. Composition graph module 121 interprets these descriptions as character attributes and includes such attributes as part of composition graph 152” (Amer, ¶ [0064]) where features of the digital animation and their use may be described;
a duration prompt portion specifying a duration for output of the digital animation “behavior analytics systems 314 may be, or may include, a camera, a sensor, or a set of cameras or sensors, or any appropriate type of image acquisition or other device enabling detection of movements, posture, poses, gestures, actions, positions, and/or gaze orientations, directions, durations, and persistence” (Amer, ¶ [0153]) where the duration may be tracked or specified by the behavior analytics system;
a task prompt portion specifying a task that the one or more machine-learning models is to undertake to discern the subject and the path and an output format of the digital animation “Many of the skills and practices used in interactive storytelling may benefit from assistance by artificial intelligence or artificially-intelligent agents or other systems. Often, the interaction between humans and agents or machines has focused on handling directives given by the human user” (Amer, ¶ [0015]) where tasks can be given to an AI to command an action to be performed, additionally, Habib teaches of moving a subject or object across a path in ¶ [0037];
an error handling portion “Computing system 310 may identify an ambiguity” (Amer, ¶ [0162]) specifying error message generation in response to inaccuracy of the description as including at least one corrective action “In response to identifying an ambiguity, computing system 310 may prompt a user for information (805)” (Amer, ¶ [0163]) additionally, “Computing system 310 may resolve the ambiguity based on the user's response (806)” (Amer, ¶ [0164]); or
an examples prompt portion including examples of inputs and corresponding animation settings “aspects of this disclosure relate to AI-assisted generation of animations or videos for the purpose of facilitating interactive storytelling, such AI-assisted generation of animations can be applied in other contexts as well, and may be used in applications that involve human motion generally, and include applications involving interactions in visual and physical worlds, robotics, animation, surveillance, and video search.” (Amer, ¶ [0015]) where the artificial intelligence may communicate with the user to improve its output in different aspects.
In regard to claim 20, Amer in view of Gogna and Edson teach [t]he one or more computer-readable storage media as described in claim 19.
Amer in view of Gogna and Edson do not explicitly teach further comprising calculating a path based on a digital image.
Habib teaches further comprising calculating a path based on a digital image ““animation canvas” refers to a graphical user interface element for displaying and/or editing a digital animation and/or an animation frame. For instance, the term “animation canvas” includes an interface for generating and/or manipulating digital design objects, drawing animation paths, generating animation frames and animation layers, etc. In one or more embodiments, the animation management system generates animation frames, animation layers, and digital animations based on user interactions with an animation canvas.” (Habib, ¶ [0037]) where the digital image is provided by the user, including a line or geometry provided which the system uses to calculate an animation sequence to follow said line and wherein the digital animation is based on the path “the animation management system animates digital design objects along animation paths in one or more animation layers” (Habib, ¶ [0017]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an object animation with artificial intelligence including candidate paths for a user to select from taught by Amer and Gogna, with the use of path calculation taught by Habib to produce an object animation, that can calculate object animation candidate paths for a user to choose from based on a prompt. The motivation to do so would be to allow the user to choose from multiple calculated paths for their object animation to follow.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Amer in view of Edson, Gogna, Habib, and Corwin.
In regard to claim 15, Amer in view of Edson, Gogna, and Habib teach [t]he computing device as described in claim 14.
Amer in view of Edson, Gogna, and Habib fail to explicitly teach segmenting the digital image using at least one machine-learning model to form at least one mask;
converting the at least one mask into a vector outline; and
generating the path based on the vector outline.
Corwin teaches segmenting the digital image using at least one machine-learning model to form at least one mask “Graphical objects can be grouped, styled, transformed and composited into previously rendered objects. The feature set includes nested transformations, clipping paths, alpha masks, filter effects and template objects. SVG drawings can be interactive and can include animation, defined in the SVG XML elements or via scripting that accesses the SVG Document Object Model (DOM)” (Corwin, ¶ [0117]) where graphics objects are able to be segmented and masked in a number of ways, such as those available through scalable vector graphics and other similar web design techniques;
converting the at least one mask into a vector outline “The user system 112 includes various processors, ... The various processors, processing systems and processing modules implemented by the user system 112 can include: a Rapid Pictorial Demonstration (RPD) Controller/Editor 120, a browser 124 having a rendering engine 126, a dynamic Scalable Vector Graphics (SVG) system 150, and an automatic SVG animation system 160” (Corwin, ¶ [0053]) where it is known that using scalable vector graphics can allow for vectors to be created from masks; and
generating the path based on the vector outline “SVG allows three types of graphic objects: vector graphic shapes such as paths and outlines consisting of straight lines and curves, bitmap images, and text. Graphical objects can be grouped, styled, transformed and composited into previously rendered objects. The feature set includes nested transformations, clipping paths, alpha masks, filter effects and template objects” (Corwin, ¶ [0117]) where if the outline is generated, a path can be made based on the outline using SVG.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an animation for an object that follows an object path using AI taught by Amer in view of Habib with the method using masking and vector outlines to find and create a path taught by Corwin to provide a machine learning model that can produce an artificial intelligence which can use vector outlines to create an animation path for an object. The motivation to do so would be to easily produce an animation which can follow a path of a provided object via artificial intelligence.
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Amer in view of Edson, Gogna, and Raitt.
In regard to claim 17, Amer in view of Edson and Gogna teach [t]he computing device as described in claim 12.
Amer in view of Edson and Gogna fail to teach wherein the prompt includes a preset options prompt portion references a plurality of preset animation options that are available for digital animation generation.
Raitt teaches wherein the prompt includes a preset options prompt portion references a plurality of preset animation options that are available for digital animation generation “The development subsystem may be used to generate one or more animation presets that are to be used in subsequent animation editing. The editing subsystem may receive one or more animation presets, provide an interface that enables an animator to specify a preset and a way of integrating the preset with the animation that is composed of recorded multi-dimensional video game world data, and automatically apply the preset to generate one or more frames of an animation that subsequently modifies the resultant visual display” (Raitt, ¶ [0041]).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of generating an object animation with artificial intelligence taught by Amer, with the method of an animation preset taught by Raitt to provide a machine learning model that can provide preset animation options to an uploaded object. The motivation to do so would be to allow for quick and simple animations to be applied to objects.
Response to Arguments
The Examiner wishes to thank the Applicant for their efficiency and cooperation towards prosecution. Applicant’s arguments with respect to claims 1, 12, and 19 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Beginning on page 11 of the Remarks, Applicant argues that Amer fails to disclose the aspects of amended independent claim 1. Applicant’s amendments clarify that the machine learning model is a large language model (LLM).
Applicant’s arguments, see Remarks, Pages 11-12, filed July 24th, 2026, with respect to the rejection of claim 1 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground of rejection is made in view of Edson. Additionally, all amendments are found within the specification, and no new matter has been added.
Applicant argues that Amer fails to teach a large language model (LLM). The updated rejection as it applies to the amended claim 1 is applied to the similar amendments and arguments of claims 12 and 19. The new rejection under 35 U.S.C. § 103 with Amer in view of Habib, Corwin and Edson teach the amended claims and make a point of obviousness for combination, Edson discloses the use of a LLM can be utilized for object animation, both based on text or image input, including the use of several prompt capabilities. Edson recites “the model-generated questions may be synthesized by submitting a prompt to a model, such as a large language model (LLM), comprising of two pieces of text: (1) the user's text prompt, and (2) a natural-language instruction such as “Generate one or more questions, e.g., clarifying questions, using the information by the user describing their request for an animated image.”” (Edson, ¶ [0091]) where an LLM is used for an animated image.
Applicant further argues Amer in view of the cited references fail to teach where the vector outlines are based on the at least one object and motion of the at least one object derived from the animation settings. The new rejection under 35 U.S.C. § 103 with Amer in view of Habib, Corwin and Edson teach the amended claims and make a point of obviousness for combination, Edson discloses scalar vector graphics where a position and motion, or path of the object is used to define its animation. The path of the animation, including the speed and angles of animation are read as animation settings under broadest reasonable interpretation. Edson recites “The animated image file may include a sequence of images that depicts a visual element of the images as changing or moving over time. Examples of the animated image file include animated scalar vector graphics may include animated scalar vector graphics” (Edson, ¶ [0031]). Additionally, “By comparing the segmented images of each extracted frame against one another, the model may track the path or difference, e.g., angular, translational, rotational, etc., of individual objects contained within the images” (Edson, ¶ [0086]).
Beginning on page 13 of the Remarks, Applicant argues that Amer fails to disclose the aspects of amended independent claim 12. Applicant’s amendments clarify that the prompt is formed by combining one or more prompt portions. Applicant adds a limitation to specify that the candidate paths are based on the prompt formed from combining one or more prompt portions. All amendments are found within the specification, and no new matter has been added.
The new grounds of rejection under 35 U.S.C. § 103 with Amer in view of Gogna and Edson teach the amended claims and make a point of obviousness to combine. Edson discloses portions of a prompt that are used in combination to create a final animated product, as no specific formatting of the prompt portions, or combination method is specified. Edson recites “the description is “a cat typing on a keyboard wearing astronaut helmet, pink-colored pixel art,” … a multi-modal prompt inclusive of the user-selected emoji and natural-language text prompt may be communicated to a generative AI model to generate the image using the selected theme” (Edson, ¶ [0071]) where multiple subjects, or portions of the prompt are put together to create the final image, additionally “The user may also elect … modify the animated image file by selecting a “Modify” button, … enabling the user to input natural-language instructions to be sent to an AI model for modifying the animated image file” (Edson, ¶ [0071]) where further prompts can be added to include more features.
Beginning on page 15 of the Remarks, Applicant points out that a rejection of claim 19 is cited as being unpatentable over Amer in view of Habib and further in view of Corwin, but no formal rejection was provided. The Examiner would like to apologize for this mistake and any unnecessary time it may have consumed. The improper rejection was withdrawn and the correctly cited rejection of Amer in view of Gogna will be addressed, additionally, the Examiner has updated the claim rejection headings to appropriately list each claim more clearly.
Applicant argues that Amer in view of Corwin fails to disclose the aspects of amended independent claim 19. Applicant’s amendments clarify that the invention makes use of a large language model (LLM).
Applicant’s arguments, see Remarks, Pages 17-18, filed July 24th, 2026, with respect to the rejection of claim 19 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground of rejection is made in view of Edson. The updated rejection is applied in the same manner previously stated for the substantially similar amendment in claim 1.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Heng Deng (Pat. Pub. CN-117725180-A). Deng discloses an LLM text prediction model which may be interacted with by a user to generate a digital human animation (abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAIDEN ALEXANDER USSERY whose telephone number is (571)272-1192. The examiner can normally be reached Monday - Friday* 7:30AM - 5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at (571) 272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CAIDEN ALEXANDER USSERY/Examiner, Art Unit 2611
/TAMMY GODDARD/Supervisory Patent Examiner, Art Unit 2611