CTNF 18/951,203 CTNF 99224 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-fti AIA The present application is being examined under the pre-AIA first to invent provisions. Information Disclosure Statement The information disclosure statement (IDS) submitted on 12/02/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Allowable Subject Matter 12-151-08 AIA 07-43 12-51-08 Claim s 4-8,11-13,17-19 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim (s) 1,2,3,14,15,16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kreis(US-20240171788-A1) in view of Ramesh(US-11922550-B1) and Mashrabov(US-20200234483-A1) . As per claim 1, Kreis discloses receiving a text prompt that specifies an object because Kreis teaches text-to-video generation based on user-designated text prompts and object-related scene controls (Kreis teaches at (para. [0028]) “The AI-based video generation and synthesis can be based on, for example, text prompts that define the content of the videos, bounding boxes that define a location or area where objects are to be generated.” This discloses text prompts defining video content and expressly contemplates objects to be generated.) As per claim 1, Kreis discloses generating a video that comprises a respective video frame at each of a plurality of time steps in the video (Kreis teaches at (para. [0030]) “the at least one temporal attention neural network layer is updated to align multiple images generated by the image diffusion model into consecutive frames of the video.” This discloses generating a video having multiple consecutive frames.) Kreis further teaches time-step/frame generation by interpolation (Kreis teaches at (para. [0034]) “In some embodiments, a video can be synthesized in a stage-wise approach. The video diffusion model can include a first video diffusion model and a second video diffusion model. The first video diffusion model can be updated/trained to generate the multiple frames of a video (referred to as the third video) as described, where the multiple frames can be sparse frames of a video. In other words, the first video has a low frame rate, low FPS, or a low temporal resolution. The second video diffusion model can be updated to generate the first video by temporally up-sampling (e.g., frame interpolation) the third video to fill in at least one frame between two consecutive frames of the third video. The stage-wise approach can generate long-term and consistent videos, by generating a low frame-rate video (the third video) of any length first, and adding the intermediate/missing frames subsequently. The resulting video (the first video) can be up-sampled (relative to the third video) in terms of spatial and/or temporally resolution as described.” This supports frames generated across temporal positions.) Kreis discloses generating a video over plural frames/time steps, but Kreis does not expressly disclose that the generated video depicts the particular instance of the object from a received image. Mashrabov supplies that feature (Mashrabov teaches at (para. [0023]) “As used herein, a facesync actor is a person whose facial landmark parameters are used, an actor is another person whose body is used in a video template and whose skin may be recolored, and a user is a person who takes an image of his face to generate a personalized video. Thus, in some embodiments, the personalized video includes the face of the user modified to have facial expressions of the facesync actor and includes a body of the actor taken from the video template and recolored to match with the color of the face of the user. ” This discloses that the generated video depicts the particular user/face instance from the input image.) Mashrabov further teaches the video is generated based on the user’s captured or selected image (Mashrabov teaches at (para. [0024]) “The user of the computing device may capture, by the computing device, an image of a face or select an image of the face from a camera roll. The computing device may further generate, based on the image of the face and one of the pre-generated video templates, a personalized video.” This supplies the particular-instance image-to-video teaching.) As per claim 1, Kreis discloses generating the respective video frame at the time step using a video generation neural network (Kreis teaches at (para. [0029]) “an image diffusion model … can be fine-tuned … into a video generator by incorporating/adding/inserting at least one temporal attention neural network layer into at least one denoising neural network.” This discloses a neural-network video generator.) As per claim 1, Kreis at least partially discloses that the video generation neural network is conditioned on prompt/control information (Kreis teaches at (para. [0037]) “In some examples, different conditioning signals can be applied. The video diffusion model can accept text prompts that describe the desired video content. The video diffusion model can be updated (e.g., video fine-tuned) to text-to-image diffusion models. To generate videos for a particular task (e.g., driving scenarios), generating the videos can be controlled by allowing/specifying/adding bounding boxes for objects in the scene. The video diffusion model can be updated (e.g., trained) to be conditional on global class-like information, e.g., day, night, and so on. .” Kreis therefore discloses conditioning video generation on text and other control signals, but not expressly on a reference-image embedding.) Kreis also teaches conditioning signals generally (Kreis teaches at (para. [0071]) “In some examples, the training system 100 can update or train the video diffusion model 110 to generate the first video according to at least one of text prompts, bounding boxes, or conditioning signals. ” This is analogous or partial support for conditioning on additional control input.) However, Kreis does not disclose receiving a control input that comprises an image that depicts a particular instance of the object; obtaining, at each of the plurality of time steps, a text prompt embedding that is generated by a joint encoder neural network based on the text prompt; obtaining, at each of the plurality of time steps, a control input embedding, the control input embedding comprising an image embedding that is generated by the joint encoder neural network based on the image; or generating the respective video frame while the video generation neural network is conditioned specifically on the text prompt embedding and on the control input embedding. The combination of Kreis and Ramesh disclose obtaining, at each of the plurality of time steps, a text prompt embedding that is generated by a joint encoder neural network based on the text prompt is disclosed by Ramesh at least as to the text-embedding generation by a jointly trained encoder structure (Ramesh teaches at (col.8 lines 29-43) “Disclosed embodiments may involve accessing a text description and inputting the text description into a text encoder. Accessing a text description may include at least one of retrieving, requesting, receiving, acquiring, or obtaining a text description. For example, a processor may be configured to access a text description that has been inputted into a machine (e.g., by a user) or access a text description corresponding to a request to generate an image based on the description. Inputting the text description into a text encoder may involve feeding, inserting, entering, submitting, or transferring the text description into a text encoder. A text encoder may convert text into an alternative representation (e.g., digital representation and/or mathematical representation) that preserves patterns, relationships, structure, or context between components of the text. ” and at col.10 lines 27-28 “At step 206 , a method may involve receiving, from the text encoder, a text embedding. At step 208 , a method may involve inputting at least one of the text description or the text embedding into a first sub-model. ” This supplies the claimed text prompt embedding generated from the text prompt. Ramesh further teaches that the text encoder is part of a jointly trained image/text encoder arrangement at (col.5 line 60- col.6 line 2) “Disclosed embodiments may involve jointly training an image encoder on (e.g., using, based on) the first data set and the text encoder on (e.g., using, based on) the second data set. Training (e.g., an image encoder or a text encoder) may include one or more of adjusting parameters (e.g., parameters of the image encoder or text encoder), removing parameters, adding parameters, generating functions, generating connections (e.g., neural network connecting), or any other machine training operation (e.g., as discussed regarding system 400 ). ” This functionally corresponds to the claimed joint encoder neural network because the encoder arrangement processes text to generate a text embedding and image data to generate an image embedding.) Obtaining, at each of the plurality of time steps, a control input embedding, the control input embedding comprising an image embedding that is generated by the joint encoder neural network based on the image is disclosed by Ramesh (Ramesh teaches at (col.6 line 42-48) “Images may be inputted into an image encoder, and the captions or text descriptions may be inputted into the text encoder. The image encoder may output an image embedding, and the text encoder may output a text embedding.” This discloses processing image input to generate an image embedding corresponding to the claimed control input embedding.) Ramesh also defines the embeddings as encoder-output representations (Ramesh teaches at (col.6 lines 16-19) “An embedding, including a text embedding or image embedding, may include an output of an encoder, such as a numeric, vector, or spatial representation of the input to the text encoder.” This supports treating the text and image encoder outputs as the claimed embeddings.); generating the respective video frame while the video generation neural network is conditioned specifically on the text prompt embedding and on the control input embedding is supplied by the combination of Kreis and Ramesh (Kreis teaches at (para. [0036]) “the video diffusion model can accept text prompts that describe the desired video content” and “generating the videos can be controlled by allowing/specifying/adding bounding boxes for objects in the scene.” This discloses video generation conditioned on text and control information, but not expressly on the claimed embeddings.) Kreis also teaches conditioning signals generally (Kreis teaches at (para. [0015]) “generate the first video according to at least one of text prompts, bounding boxes, or conditioning signals.” This provides partial support for conditioning the video generation neural network on additional control input.) Although Kreis does not expressly disclose conditioning specifically on a text prompt embedding and an image/control input embedding at each time step, Ramesh supplies the known text/image embedding representations, and it would have been obvious to use Ramesh’s text and image embeddings as the conditioning signals in Kreis’s video diffusion model when generating the respective video frames at the plurality of time steps. However, the combination of Kreis and Ramesh fail to disclose the remaining limitations. Mashrabov discloses receiving a control input that comprises an image that depicts a particular instance of the object is (Mashrabov teaches at (para. [0024]) “The user of the computing device may capture, by the computing device, an image of a face or select an image of the face from a camera roll. The computing device may further generate, based on the image of the face and one of the pre-generated video templates, a personalized video.” This discloses receiving an image depicting a particular instance, namely a particular user’s face, and generating a video based on that image.); It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to incorporate the teachings of Ramesh and Mshrabov into the teachings of Kreis in order to modify Kreis’s text/conditioning-signal video diffusion system to use Ramesh’s jointly trained text/image encoder embeddings, because Kreis already teaches that video generation may be controlled by text prompts and conditioning signals, and Ramesh teaches a known way to convert text and image inputs into encoder embeddings suitable for machine-learning generation. It further would have been obvious to use Mashrabov’s particular-instance user image as the image control input so that the generated video depicts the specific object/person instance shown in the image. The predictable technical benefit would be improved controllability and instance consistency in generated video: Kreis supplies temporally coherent video generation, Ramesh supplies compact multimodal text/image conditioning representations, and Mashrabov supplies the known personalization goal of making a generated video depict a particular input-image instance As per claim 2, the combination of Kreis, Ramesh and Mashrabov disclose all the elements of claim 1 as discussed above. Ramesh also discloses wherein obtaining the control input embedding comprises processing the image using an image encoder tower of the joint encoder neural network to generate the image embedding , because Ramesh teaches a jointly trained text/image encoder arrangement in which the image side processes image input and outputs an image embedding (Ramesh teaches at ( col. 5 line 60- col.6 line 2) “Disclosed embodiments may involve jointly training an image encoder on (e.g., using, based on) the first data set and the text encoder on (e.g., using, based on) the second data set. Training (e.g., an image encoder or a text encoder) may include one or more of adjusting parameters (e.g., parameters of the image encoder or text encoder), removing parameters, adding parameters, generating functions, generating connections (e.g., neural network connecting), or any other machine training operation (e.g., as discussed regarding system 400 ).” This functionally corresponds to the claimed joint encoder neural network because the jointly trained encoder structure processes text to generate a text embedding and processes image data to generate an image embedding.) Ramesh further expressly discloses image processing by the image encoder (Ramesh teaches at (para. [Detailed Description]) “Images may be inputted into an image encoder, and the captions or text descriptions may be inputted into the text encoder. The image encoder may output an image embedding, and the text encoder may output a text embedding.” This maps to processing the image using the image encoder side/tower of the joint encoder neural network to generate the image embedding.) Ramesh does not use the exact term “image encoder tower,” but its image encoder within the jointly trained text/image encoder architecture provides the same functional structure and would render the claimed tower implementation obvious. The rationale from claim 1 is incorporated herein. As per claim 3 the combination of Kreis, Ramesh and Mashrabov discloses all the elements of claim 2 as discussed above. The combination also discloses “wherein the control input further comprises location data that defines a location of the object within the respective video frame at each of the plurality of time steps, and wherein the video depicts the particular instance of the object at the defined location within the respective video frame at each of the plurality of time steps,” Kreis discloses the control input including location-type data for controlling where objects are generated in video frames (Kreis teaches at (para. [0028]) “The AI-based video generation and synthesis can be based on, for example, text prompts that define the content of the videos, bounding boxes that define a location or area where objects are to be generated.” This discloses location data defining where an object is to be generated.) Kreis further teaches controlling generated videos by specifying object bounding boxes in the scene (Kreis teaches at (para. [0037]) “To generate videos for a particular task (e.g., driving scenarios), generating the videos can be controlled by allowing/specifying/adding bounding boxes for objects in the scene. ” This maps to the claimed control input further comprising location data that defines a location of the object within the video frame. Kreis also discloses the plurality-of-time-steps / respective-frame aspect at (para. [0030]) “In some examples, the at least one temporal attention neural network layer is updated to align multiple images generated by the image diffusion model into consecutive frames of the video.” This discloses video frames across time.) Kreis further supports time-step-based frame generation at (para. [0034]) “In some embodiments, a video can be synthesized in a stage-wise approach. The video diffusion model can include a first video diffusion model and a second video diffusion model. The first video diffusion model can be updated/trained to generate the multiple frames of a video (referred to as the third video) as described, where the multiple frames can be sparse frames of a video. In other words, the first video has a low frame rate, low FPS, or a low temporal resolution. The second video diffusion model can be updated to generate the first video by temporally up-sampling (e.g., frame interpolation) the third video to fill in at least one frame between two consecutive frames of the third video. The stage-wise approach can generate long-term and consistent videos, by generating a low frame-rate video (the third video) of any length first, and adding the intermediate/missing frames subsequently. The resulting video (the first video) can be up-sampled (relative to the third video) in terms of spatial and/or temporally resolution as described. ” This supports video frames generated at temporal positions.) However, Kreis does not expressly disclose that the object placed using the location data is the particular instance depicted in a received image. Mashrabov supplies that particular-instance aspect (Mashrabov teaches at (para. [0024]) “The user of the computing device may capture, by the computing device, an image of a face or select an image of the face from a camera roll. The computing device may further generate, based on the image of the face and one of the pre-generated video templates, a personalized video.” This discloses using an image of a particular instance to generate a video.) Mashrabov further teaches that the generated video depicts that particular instance (Mashrabov teaches at (para. [0023]) “As used herein, a facesync actor is a person whose facial landmark parameters are used, an actor is another person whose body is used in a video template and whose skin may be recolored, and a user is a person who takes an image of his face to generate a personalized video. Thus, in some embodiments, the personalized video includes the face of the user modified to have facial expressions of the facesync actor and includes a body of the actor taken from the video template and recolored to match with the color of the face of the user. ” This discloses the particular user/face instance appearing in the generated video.) Although Kreis does not expressly disclose location data for the particular input-image instance at each time step, it would have been obvious to apply Kreis’s bounding-box/location control to Mashrabov’s personalized-video instance so that the particular instance from the input image is generated at the specified locations in the video frames. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to incorporate the teachings of Ramesh and Mshrabov into the teachings of Kreis in order to combine controll where generated objects appear in video frames using bounding boxes/location conditioning, while using Mashrabov teachings to use an input image of a particular user to generate a personalized video depicting that particular instance. The combination would predictably improve controllability of personalized video generation by allowing the generated video not only to preserve the particular input-image instance, but also to place that instance at defined locations across the generated video frames Claim 15, which is similar in scope to claim 2, thus rejected under the same rationale. Claim 16, which is similar in scope to claim 3, thus rejected under the same rationale. Claim 20, is similar in scope to claim 1. However, claim 20 recites A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising. Kreis also discloses this at para.[0102] “ The computer-storage media may include both volatile and nonvolatile media and/or removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, and/or other data types. For example, the memory 704 may store computer-readable instructions (e.g., that represent a program(s) and/or a program element(s), such as an operating system. Computer-storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which may be used to store the desired information and which may be accessed by computing device 700 . As used herein, computer storage media does not comprise signals per se.” The rest of claim 20 is rejected under the same rationale as claim 1. Claim 14, which is similar in scope to claim 20, thus rejected under the same rationale . 07-21-aia AIA Claim( s) 9, 10 are r ejected under 35 U.S.C. 103 as being unpatentable over K reis in view of Hong(Hong, Wenyi, et al. "Cogvideo: Large-scale pretraining for text-to-video generation via transformers." arXiv preprint arXiv:2205.15868 (2022).), Li(Li, Yuheng, et al. "GLIGEN: Open-Set Grounded Text-to-Image Generation." arXiv preprint arXiv:2301.07093 (2023).) and Blattmann (Blattmann, Andreas, et al. "Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models." arXiv preprint arXiv:2304.08818 (2023).). A s per claim 9, Kreis discloses: a computer-implemented method of training a video generation neural network (Kreis teaches at (para. [0030]) “In some embodiments, an image diffusion model (e.g., a base model) that is pre-trained using image datasets can be fine-tuned (e.g., configured, implemented, updated, modified) into a video generator by incorporating/adding/inserting at least one temporal attention neural network layer into at least one denoising neural network of the image diffusion model. The image diffusion model can be any suitable model pre-trained to generate or synthesize images.” This teaches training/adapting a pre-trained generative neural network into a video generation neural network);, Kreis alone does not teach the remaining features of claim 9. However, Kreis in combination with Hong teaches the claimed: obtaining data specifying a pre-trained video generation neural network that has been pre-trained to generate videos conditioned on text inputs (Hong teaches in the Abstract “In this work, we present 9B-parameter transformer CogVideo, trained by inheriting a pretrained text-to-image model, CogView2. We also propose multi-frame-rate hierarchical training strategy to better align text and video clips. As (probably) the first open-source large-scale pretrained text-to-video model, CogVideo outperforms all publicly available models at a large margin in machine and human evaluations.” This discloses a pre-trained transformer-based video generation neural network for text-to-video generation); Kreis in combination with Hong and Li teaches the claimed : wherein the pre-trained video generation neural network comprises one or more original attention blocks, each original attention block comprising a self-attention layer followed by a cross-attention layer (Li teaches at Section 4.2 “The original Transformer block of LDM consists of two attention layers: The self-attention over the visual tokens, followed by cross-attention from caption tokens.” This expressly teaches the claimed self-attention-then-cross-attention block order); Kreis in combination with Hong, Li, and Blattman teaches the claimed : generating the video generation neural network from the pre-trained video generation neural network by inserting an adaptive cross-attention layer to each of the one or more original attention blocks (Kreis teaches at (para. [0030]) “...incorporating/adding/inserting at least one temporal attention neural network layer into at least one denoising neural network of the image diffusion model,” and Blattmann teaches at Section 3.1 “We thus introduce additional temporal neural network lay ers li φ, which are interleaved with the existing spatial layers li θ and learn to align individual frames in a temporally consistent manner.” These teach inserted-layer adaptation of a pre-trained generative model for video generation; Li further teaches at Section 4.2 “We freeze these two attention layers and add a new gated self-attention layer” and “Note that (8) is injected in between (6) and (7),” which supplies the closest teaching for inserting a new trainable attention/control layer into an original self-attention/cross-attention block);, and training the video generation neural network to determine trained parameter values of the adaptive cross-attention layers while holding pre-trained parameter values of the self-attention layer and the cross-attention layer included in each of the one or more original attention blocks fixed (Hong teaches at section 3.2 paragraph “All the parameters in the CogView2 are frozen in the training, and only the parameters in the newly added attention layer (See the attention plus in Figure 3) are trainable. ” Kreis teaches at (para. [0011]) “update the at least one first temporal attention layer using at least one sample sequence of images—while maintaining parameters of the at least one spatial layer unchanged.” Blattmann teaches in Figure 4 that “We turn a pre-trained LDM into a video generator by inserting temporal layers that learn to align frames into tempo rally consistent sequences. During optimization, the image back bone θ remains fixed and only the parameters φ of the temporal layers li φ are trained, cf. Eq. (2). ” Li teaches at Section 4.2 “We freeze these two attention layers and add a new gated self-attention layer to enable the spatial grounding ability; see Figure 3. ” These teachings collectively support training newly inserted attention layers while holding original pre-trained attention/backbone parameters fixed). Although, Kreis, Blattmann, Hong, and Li do not expressly disclose the exact phrase “adaptive cross-attention layer” inserted into each original attention block of a pre-trained text-conditioned video generation neural network, Kreis and Blattmann disclose inserted temporal attention layers, Hong discloses inserted spatial-temporal attention channels and Li discloses an inserted gated self-attention layer between frozen original self-attention and cross-attention layers. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to incorporate the teachings of Hong, Li and Blattmann into the teachings of Kreis in order t o implement the inserted control/attention layer as an adaptive cross-attention layer where additional conditioning information is supplied through attention, because the references collectively teach parameter-efficient adaptation of pre-trained generative models by inserting trainable attention/control layers while freezing original parameters, with the predictable benefit of adding controllability and video-specific adaptation while preserving the original text-conditioned generation capability and reducing training cost. As per claim 10 the combination of Kreis, Hong, Li and Blattmann disclose all the elements of claim 9 as discussed above. Li also discloses wherein training the video generation neural network comprises training the video generation neural network on a plurality of training inputs to optimize an objective function, wherein each training input includes (i) a training text prompt, (ii) a training image, and (iii) training location data ,because Li teaches a model continual learning based on grounding instruction inputs and an objective function (Li teaches at Section 4.2 “we use the original denoising objective as in (1) for model continual learning, based on the grounding instruction input y.” This discloses training on grounding inputs to optimize an objective function.) Li further teaches that the training input includes a text prompt/caption and grounded entities with location data (Li teaches at Section 4.1 “We define the instruction to a grounded text-to-image model as a composition of the caption and grounded entities” and “Grounding: e = [(e1, l1), · · · , (eN, lN)].” This maps to a training text prompt and location-associated grounded entity data.) Li further teaches that the location data may be a bounding box (Li teaches at Section 4.1 “we primarily study using bounding box as the grounding spatial configuration l., because of its large availability and easy annotation for users. ” This maps to training location data.) Li further teaches that the grounded entity may be represented by an image (Li teaches at Section 4.1 Image Prompt that “one may describe entity e using an image, instead of language” and “We use an image encoder to obtain feature fimage(e).” This maps to the claimed training image.) Although Li trains a grounded text-to-image diffusion model rather than a video generation neural network, it supplies the claimed training-input structure and objective-function optimization, and it would have been obvious to use such text/image/location training inputs in the video-training framework of claim 9 to provide the predictable benefit of learning controllable generation from text, reference-image, and spatial-location conditioning. The rationale of claim 9 is incorporated herein. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRIS ALEJANDRO PUNTIER whose telephone number is (703)756-1893. The examiner can normally be reached M-F 7:30-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CHRIS ALEJANDRO PUNTIER/Examiner, Art Unit 2616 /DANIEL F HAJNIK/Supervisory Patent Examiner, Art Unit 2616 Application/Control Number: 18/951,203 Page 2 Art Unit: 2616 Application/Control Number: 18/951,203 Page 3 Art Unit: 2616 Application/Control Number: 18/951,203 Page 4 Art Unit: 2616