Prosecution Insights
Last updated: August 17, 2026
Application No. 18/900,457

MEDIA ITEM CHARACTERIZATION BASED ON MULTIMODAL EMBEDDINGS

Non-Final OA §103§112§DP
Filed
Sep 27, 2024
Priority
Sep 29, 2023 — provisional 63/587,046 +1 more
Examiner
CASCAIS, JUSTIN PHILIP
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
12m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
48 granted / 64 resolved
+15.0% vs TC avg
Moderate +14% lift
Without
With
+13.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
15 currently pending
Career history
72
Total Applications
across all art units

Statute-Specific Performance

§101
10.6%
-29.4% vs TC avg
§103
60.1%
+20.1% vs TC avg
§102
14.9%
-25.1% vs TC avg
§112
10.6%
-29.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 64 resolved cases

Office Action

§103 §112 §DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant claims the benefit of US Provisional Application No. 63/587,046, filed 9/29/2023 and US Provisional Application No. 63/587,047, filed 9/29/2023 but claims 1-20 contain subject matter not disclosed in the provisional applications, thus the effective filing date of 9/27/2024 has been used. If applicant intends to have these claims afforded the benefit of the earlier filing date, applicant may: (1) amend the claim limitation(s) to avoid it/them containing subject matter not in the provisional application; or (2) present a sufficient showing that the claim limitation(s) recite(s) material from specific sections of the provisional application. Information Disclosure Statement The IDS(s) dated 1/31/2025 has been considered and placed in the application file. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-6, 8-9, 11-13, and 18-20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-5, 7-9, and 15-20 of copending Application No. 18/900,473 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because both applications claim generating or obtaining audiovisual embeddings for a media item using video/visual information and audio information, and determining a media characteristic or media-trend association based on the embeddings. In particular, the reference claims recite obtaining audiovisual embeddings representing audiovisual features of a media item, obtaining textual embeddings, providing the embeddings to an AI model trained to predict whether a media item is associated with one or more media trends of a platform, and determining whether the media item is associated with the one or more media trends. The reference claims further recite obtaining a video embedding representing visual features of at least one frame, obtaining an audio embedding representing audio features of an audio signal associated with the frame, generating ana audiovisual embedding based on fused audiovisual data comprising the video embedding and audio embedding, and updating the set of audiovisual embeddings to include the generated audiovisual embedding. The instant claims differ from the reference claims primarily in reciting that the media item comprises a sequence of video frames and that each audiovisual embedding represents a visual feature and an audio feature of a respective video frame. The reference claims differ from the instant claims primarily by expressly reciting textual embeddings and an AI model trained to predict media-trend association. These differences do not render the claims patentably distinct because they are directed to functionally overlapping implementations of the same media-item characterization workflow of generating audiovisual embeddings from visual/video and audio features of a media item and using the generated embeddings to determine a platform-relevant media characteristic, including media-trend association. This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims of this application have not in fact been patented. Claims 1-6, 8-9, and 11-20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of copending Application No. 18/900,467 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because both applications claim determining whether media items are associated with a media trend of a platform based on audiovisual, pose, embedding, distance, similarity, or alignment information. In particular, the reference claims recite identifying media items having common media characteristics, obtaining audiovisual embeddings, generating audiovisual embeddings based on video and audio embeddings, determining pose values or pose embeddings for objects depicted by media items, calculating distance scores or distances between pose values or pose embeddings of media items, and determining whether media items correspond to or are associated with a media trend of a platform. The instant claims differ from the reference claims primarily in reciting that the visual features of the sequence of video frames comprise poses of an object, identifying embeddings for an additional media item associated with a media trend, determining whether a degree of alignment between poses satisfies alignment criteria based on the audiovisual embeddings and the embeddings for the additional media item, and determining that the media item is associated with coherence scores, template media items, common audiovisual features, or similarity between audiovisual features. These differences do not render the claims patentably distinct because a distance, difference, similarity, or coherence determination between pose/audiovisual embeddings is a functionally equivalent way of determining whether visual pose features align sufficiently to associate a media item with a media trend. With respect to claim 17, the claimed “difference threshold” is not patentably distinct from the reference claims’ use of distance scores, distance calculations, similarity determinations, coherence criteria, and template/common audiovisual feature comparison to determine whether a media item is associated with a media trend. Determining whether a calculated difference, distance, or similarity satisfies a criterion is an ordinary and functionally equivalent implementation of threshold-based trend association. This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims of this application have not in fact been patented. Claims 1, 4, 7-9, 11-13, 18, and 20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of copending Application No. 19/209,412 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because both applications claim generating embeddings representing features of media items, using the embeddings to determine similarity or trend association between media items of a platform, and providing or determining an indication that the media items correspond to a media trend. The reference claims further recite audiovisual embeddings, textual embeddings, video embeddings representing visual features of a sequence of video frames, audio embeddings representing audio features of the sequence of video frames, concatenation of video/audio/text embeddings, and attention pooling to generate embeddings used for emerging media-trend identification. The instant claims differ primarily by reciting media-item characterization based on audiovisual embeddings, while the reference claims recite real-time or current-time-window identification of an emerging media trend. These differences do not render the claims patentably distinct because both sets of claims use embeddings derived from audiovisual/textual media features to determine whether media items are associated with a media trend of a platform. This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims of this application have not in fact been patented. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claim(s) 8 and 14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 8 recites the limitation "media items". The plural “media items” lacks consistency with claim 1’s “media item”. There is insufficient antecedent basis for this limitation in the claim. Claim 14 recites the limitation “a degree alignment”. The specification in at least paragraph [39] discloses “the system can determine … a degree of alignment”. It is unclear whether “a degree alignment” is “a degree of alignment” as outlined in the specification. Claim(s) 9-10 and 15-17 depend either directly or indirectly from the rejection of Claim(s) 8 and 14, therefore they are also rejected. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-5, 8-9, and 18-20 is/are rejected under 35 U.S.C. 103 as obvious over Akbari et al (Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., & Gong, B. (2021). Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34, 24206-24221, hereafter referred to as Akbari), in view of Mercea et al (Mercea, O. B., Hummel, T., Koepke, A. S., & Akata, Z. (2022, October). Temporal and cross-modal attention for audio-visual zero-shot learning. In European Conference on Computer Vision (pp. 488-505). Cham: Springer Nature Switzerland., hereafter referred to as Mercea), further in view of Kmiec et al (Kmiec, S., Bae, J., & An, R. (2018). Learnable pooling methods for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops (pp. 0-0)., hereafter referred to as Kmiec). Claim 1 Regarding Claim 1, Akbari teaches A method comprising: identifying a media item comprising a sequence of video frames (Akbari on page 2 discloses “[taking] as input the raw RGB frames of internet videos, audio waveforms, and text transcripts”); obtaining a set of video embeddings representing visual features of the sequence of video frames (Akbari on page 4 under section 3.1 discloses partitioning a video clip into patches and applying linear projection to obtain “a d-dimensional vector representation”; Figure 1 discloses “VATT linearly projects each modality into a feature vector and feeds it into a Transformer encoder”); obtaining a set of audio embeddings representing audio features of the sequence of video frames (Akbari on page 4 under section 3.1 discloses raw audio waveform input and partitioning the waveform into segments, then applying a linear projection “to get a d-dimensional vector representation”). Akbari does not explicitly teach all of generating a set of audiovisual embeddings based on the set of video embeddings and the set of audio embeddings, wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames; and determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings. However, Mercea teaches generating a set of audiovisual embeddings based on the set of video embeddings and the set of audio embeddings (Mercea in Fig. 1 discloses an audio-visual framework that “learns a multi-modal embedding … by exploiting the temporal alignment between audio and visual data in videos”. Page 7 discloses audio and visual features passed through modality-specific embedding blocks and an “audio-visual embedding” used for prediction), wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames (Mercea in page 5 under section 3.2 discloses temporal embeddings that encode “the actual point in time in the video which corresponds to an audio or visual representation”); determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings (Mercea in page 7 discloses projecting the audio-visual embedding to a label embedding space and obtaining class prediction by closet label embedding). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari by incorporating the temporally aligned audio-visual cross-attention embedding framework that is taught by Mercea, since both reference are analogous art in the field of multimodal audio-visual video representation learning and classification; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari’s multimodal Transformer framework with Mercea’s temporally aligned audio-visual cross-attention framework yields the predictable result of preserving temporal correspondence between visual and audio features while generating fused audiovisual representations, thereby improving classification and characterization of media items. Akbari in view of Mercea does not explicitly teach all of wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames. However, Kmiec teaches wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames (Kmiec in section 3 discloses “audio and video features already extracted at the frame level per second of video” and aggregation of “all local descriptors per frame” into a representation). determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings (Kmiec in section 3 discloses passing the final global video descriptor to a classifier, where “probabilities are output across possible video labels”. Under BRI, video labels/classification are “media characteristics”). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea by incorporating the frame-level video/audio descriptor aggregation and attention-pooling techniques that is taught by Kmiec, since both reference are analogous art in the field of audio-visual representation learning and classification of video/media content; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea’s audio-visual representation-learning framework with Kmiec’s frame-level descriptor and attention-pooling framework yields the predictable result of generating temporally/frame-aligned audiovisual representation suitable for media-item classification, thereby improving the accuracy and granularity of media characterization. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 2 Regarding Claim 2, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 1, wherein obtaining the set of video embeddings comprises: providing each of the sequence of video frames as an input to an image encoder (Akbari in page 4 under section 3.1 discloses the vision modality input consists of “3-channel RGB pixels of video frames” and that raw signals are tokenized and fed to Transformers); and obtaining, based on one or more outputs of the image encoder, a sequence of image tokens each representing one or more visual features of a respective video frame of the sequence of video frames (Akbari in page 4 under section 3.1 discloses partitioning a video clip into patches and applying linear projection to obtain vector representations, and feeding token sequences into the Transformer), wherein the set of video embeddings comprises the sequence of image tokens (Akbari in page 4 under section 3.1 discloses the sequence of input tokens/projected patch vectors are used as Transformer input and output representations for classification/common-space mapping). Claim 3 Regarding Claim 3, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 2, wherein the image encoder comprises a vision transformer (Akbari in Abstract discloses “VATT’s vision Transformer” and states VATT includes video, audio, and text Transformers). Claim 4 Regarding Claim 4, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 2, wherein the one or more visual features comprise at least one of: a scene depicted by the sequence of video frames, an object of the scene depicted by the sequence of video frames, at least one of an action, a motion, or a pose of the object of the scene, one or more colors included in the scene, or one or more lighting features associated with the scene (Akbari in Abstract discloses downstream “video action recognition”, which maps to the listed “action” visual feature.). Claim 5 Regarding Claim 5, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 1, wherein obtaining the set of audio embeddings comprises: extracting, from the media item, an audio signal associated with a respective video frame of the sequence of video frames (Akbari in page 4 under section 3.1 discloses audio waveform input from the same video source and partitioning raw audio waveform into segments. Mercea in page 5 under section 3.2 discloses temporal audio/visual representations tied to time in the video); providing the extracted audio signal as an input to an audio encoder (Akbari in page 4 under section 3.1 discloses raw audio waveform input and audio Transformer processing); and obtaining, based on one or more outputs of the audio encoder, an audio embedding representing audio features of the audio signal (Akbari in page 4 under section 3.1 discloses applying a linear projection to audio waveform segments to get a d-dimensional vector representation). Claim 8 Regarding Claim 8, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 1, further comprising: obtaining a set of textual embeddings representing textual features associated with content of the sequence of video frames (Akbari in page 2 and page 4 under section 3.1 discloses using text transcripts of speech audio as input. Each word is mapped to a one-hot vector followed by linear projection, “equivalent to an embedding dictionary lookup”), wherein the one or more media characteristics associated with the media items are further determined based on the set of textual embeddings (Akbari in page 5 under section 3.3 and page 6 under section 4.1 discloses video-audio-text triplets and common-space mappings involving text Transformer outputs and video embeddings, and evaluates text-to-video retrieval as a downstream task.). Claim 9 Regarding Claim 9, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 8, wherein obtaining the set of textual embeddings comprises: providing the textual features associated with content of the sequence of video frames as an input to a text encoder (Akbari in page 4 under section 3.1 discloses text input as a sequence of words and mapping each word to a vector for Transformer processing); and obtaining, based on one or more outputs of the text encoder, one or more text tokens representing the textual features (Akbari in page 4 under section 3 discloses text tokens/vectors and Transformer outputs used in the video-text common space), wherein the set of textual embeddings comprises the one or more text tokens (Akbari in page 4 under section 3.1 discloses the word vectors/embedding lookup and Transformer text outputs as textual representations). Claim 18 Regarding Claim 18, Akbari teaches A system comprising: a memory (Akbari in Abstract discloses “a framework for learning multimodal representations from unlabeled data using convolution-free Transformer architectures”); and a processing device coupled to the memory (Akbari in Abstract discloses “a framework for learning multimodal representations from unlabeled data using convolution-free Transformer architectures”), wherein the processing device is to perform operations comprising: identifying a media item comprising a sequence of video frames (Akbari on page 2 discloses “[taking] as input the raw RGB frames of internet videos, audio waveforms, and text transcripts”); obtaining a set of video embeddings representing visual features of the sequence of video frames (Akbari on page 4 under section 3.1 discloses partitioning a video clip into patches and applying linear projection to obtain “a d-dimensional vector representation”; Figure 1 discloses “VATT linearly projects each modality into a feature vector and feeds it into a Transformer encoder”); obtaining a set of audio embeddings representing audio features of the sequence of video frames (Akbari on page 4 under section 3.1 discloses raw audio waveform input and partitioning the waveform into segments, then applying a linear projection “to get a d-dimensional vector representation”). Akbari does not explicitly teach all of generating a set of audiovisual embeddings based on the set of video embeddings and the set of audio embeddings, wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames; and determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings. However, Mercea teaches generating a set of audiovisual embeddings based on the set of video embeddings and the set of audio embeddings (Mercea in Fig. 1 discloses an audio-visual framework that “learns a multi-modal embedding … by exploiting the temporal alignment between audio and visual data in videos”. Page 7 discloses audio and visual features passed through modality-specific embedding blocks and an “audio-visual embedding” used for prediction), wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames (Mercea in page 5 under section 3.2 discloses temporal embeddings that encode “the actual point in time in the video which corresponds to an audio or visual representation”); determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings (Mercea in page 7 discloses projecting the audio-visual embedding to a label embedding space and obtaining class prediction by closet label embedding). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari by incorporating the temporally aligned audio-visual cross-attention embedding framework that is taught by Mercea, since both reference are analogous art in the field of multimodal audio-visual video representation learning and classification; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari’s multimodal Transformer framework with Mercea’s temporally aligned audio-visual cross-attention framework yields the predictable result of preserving temporal correspondence between visual and audio features while generating fused audiovisual representations, thereby improving classification and characterization of media items. Akbari in view of Mercea does not explicitly teach all of wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames. However, Kmiec teaches wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames (Kmiec in section 3 discloses “audio and video features already extracted at the frame level per second of video” and aggregation of “all local descriptors per frame” into a representation). determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings (Kmiec in section 3 discloses passing the final global video descriptor to a classifier, where “probabilities are output across possible video labels”. Under BRI, video labels/classification are “media characteristics”). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea by incorporating the frame-level video/audio descriptor aggregation and attention-pooling techniques that is taught by Kmiec, since both reference are analogous art in the field of audio-visual representation learning and classification of video/media content; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea’s audio-visual representation-learning framework with Kmiec’s frame-level descriptor and attention-pooling framework yields the predictable result of generating temporally/frame-aligned audiovisual representation suitable for media-item classification, thereby improving the accuracy and granularity of media characterization. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 19 Regarding Claim 19, Akbari in view of Mercea, further in view of Kmiec teaches The system of claim 18, wherein obtaining the set of video embeddings comprises: providing each of the sequence of video frames as an input to an image encoder (Akbari in page 4 under section 3.1 discloses the vision modality input consists of “3-channel RGB pixels of video frames” and that raw signals are tokenized and fed to Transformers); and obtaining, based on one or more outputs of the image encoder, a sequence of image tokens each representing one or more visual features of a respective video frame of the sequence of video frames (Akbari in page 4 under section 3.1 discloses partitioning a video clip into patches and applying linear projection to obtain vector representations, and feeding token sequences into the Transformer), wherein the set of video embeddings comprises the sequence of image tokens (Akbari in page 4 under section 3.1 discloses the sequence of input tokens/projected patch vectors are used as Transformer input and output representations for classification/common-space mapping). Claim 20 Regarding Claim 20, Akbari teaches A non-transitory machine-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising: identifying a media item comprising a sequence of video frames (Akbari on page 2 discloses “[taking] as input the raw RGB frames of internet videos, audio waveforms, and text transcripts”); obtaining a set of video embeddings representing visual features of the sequence of video frames (Akbari on page 4 under section 3.1 discloses partitioning a video clip into patches and applying linear projection to obtain “a d-dimensional vector representation”; Figure 1 discloses “VATT linearly projects each modality into a feature vector and feeds it into a Transformer encoder”); obtaining a set of audio embeddings representing audio features of the sequence of video frames (Akbari on page 4 under section 3.1 discloses raw audio waveform input and partitioning the waveform into segments, then applying a linear projection “to get a d-dimensional vector representation”). Akbari does not explicitly teach all of generating a set of audiovisual embeddings based on the set of video embeddings and the set of audio embeddings, wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames; and determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings. However, Mercea teaches generating a set of audiovisual embeddings based on the set of video embeddings and the set of audio embeddings (Mercea in Fig. 1 discloses an audio-visual framework that “learns a multi-modal embedding … by exploiting the temporal alignment between audio and visual data in videos”. Page 7 discloses audio and visual features passed through modality-specific embedding blocks and an “audio-visual embedding” used for prediction), wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames (Mercea in page 5 under section 3.2 discloses temporal embeddings that encode “the actual point in time in the video which corresponds to an audio or visual representation”); determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings (Mercea in page 7 discloses projecting the audio-visual embedding to a label embedding space and obtaining class prediction by closet label embedding). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari by incorporating the temporally aligned audio-visual cross-attention embedding framework that is taught by Mercea, since both reference are analogous art in the field of multimodal audio-visual video representation learning and classification; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari’s multimodal Transformer framework with Mercea’s temporally aligned audio-visual cross-attention framework yields the predictable result of preserving temporal correspondence between visual and audio features while generating fused audiovisual representations, thereby improving classification and characterization of media items. Akbari in view of Mercea does not explicitly teach all of wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames. However, Kmiec teaches wherein each of the set of audiovisual embeddings represents a visual feature and an audio feature of a respective video frame of the sequence of video frames (Kmiec in section 3 discloses “audio and video features already extracted at the frame level per second of video” and aggregation of “all local descriptors per frame” into a representation). determining one or more media characteristics associated with the media item based on the set of audiovisual embeddings (Kmiec in section 3 discloses passing the final global video descriptor to a classifier, where “probabilities are output across possible video labels”. Under BRI, video labels/classification are “media characteristics”). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea by incorporating the frame-level video/audio descriptor aggregation and attention-pooling techniques that is taught by Kmiec, since both reference are analogous art in the field of audio-visual representation learning and classification of video/media content; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea’s audio-visual representation-learning framework with Kmiec’s frame-level descriptor and attention-pooling framework yields the predictable result of generating temporally/frame-aligned audiovisual representation suitable for media-item classification, thereby improving the accuracy and granularity of media characterization. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim(s) 6-7 is/are rejected under 35 U.S.C. 103 as obvious over Akbari et al (Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., & Gong, B. (2021). Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34, 24206-24221, hereafter referred to as Akbari), in view of Mercea et al (Mercea, O. B., Hummel, T., Koepke, A. S., & Akata, Z. (2022, October). Temporal and cross-modal attention for audio-visual zero-shot learning. In European Conference on Computer Vision (pp. 488-505). Cham: Springer Nature Switzerland., hereafter referred to as Mercea), further in view of Kmiec et al (Kmiec, S., Bae, J., & An, R. (2018). Learnable pooling methods for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops (pp. 0-0), hereafter referred to as Kmiec), further in view of Gong et al (Gong, Y., Chung, Y. A., & Glass, J. (2021). Ast: Audio spectrogram transformer. arXiv preprint arXiv:2104.01778, hereafter referred to as Gong). Claim 6 Regarding Claim 6, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 5. Akbari in view of Mercea, further in view of Kmiec does not explicitly teach all of wherein the audio encoder comprises an audio spectrogram transformer. However, Gong teaches wherein the audio encoder comprises an audio spectrogram transformer (Gong in Abstract discloses “Audio Spectrogram Transformer (AST); Section 2.1 discloses converting input audio waveform to log-Mel spectrogram features, splitting into patches, projecting to patch embeddings, and inputting to a Transformer). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec by incorporating the Audio Spectrogram Transformer (AST) that is taught by Gong, since both reference are analogous art in the field of audio feature extraction and classification; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec’s audio-visual representation-learning framework with Gong’s AST as the audio encoder yields the predictable result of improved audio-feature representation for downstream audiovisual media characterization. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 7 Regarding Claim 7, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 5. Akbari in view of Mercea, further in view of Kmiec does not explicitly teach all of wherein the audio encoder comprises an audio spectrogram transformer. However, Gong teaches wherein the audio features comprise at least one of: a pitch of an audio signal of the media item, a timbre of the audio signal, a rhythm of the audio signal, speech content of the audio signal, speaker characteristics associated with the audio signal, environmental sounds associated with a scene of the media item, spectral features of the audio signal, or temporal dynamics of the audio signal. (Gong in section 2.1 discloses converting waveform into log-Mel filterbank features and a spectrogram, which maps to “spectral features”; Section 2.2 and 3.2 discloses encoding positional information for time/frequency patches, supporting temporal audio dynamics). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec by incorporating spectral features and temporal audio dynamics that is taught by Gong, since both reference are analogous art in the field of audio feature extraction and classification; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec’s audio-visual representation-learning framework with Gong’s spectrogram-based audio-feature extraction yields the predictable result of generating richer audio embeddings that capture spectral and temporal characteristics of the media item, thereby improving audiovisual classification and media characterization. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim(s) 10 is/are rejected under 35 U.S.C. 103 as obvious over Akbari et al (Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., & Gong, B. (2021). Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34, 24206-24221, hereafter referred to as Akbari), in view of Mercea et al (Mercea, O. B., Hummel, T., Koepke, A. S., & Akata, Z. (2022, October). Temporal and cross-modal attention for audio-visual zero-shot learning. In European Conference on Computer Vision (pp. 488-505). Cham: Springer Nature Switzerland., hereafter referred to as Mercea), further in view of Kmiec et al (Kmiec, S., Bae, J., & An, R. (2018). Learnable pooling methods for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops (pp. 0-0)., hereafter referred to as Kmiec), further in view of Devlin (Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171-4186), hereafter referred to as Devlin). Claim 10 Regarding Claim 10, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 9. Akbari in view of Mercea, further in view of Kmiec does not explicitly teach all of wherein the text encoder comprises a bidirectional encoder representations from transformers (BERT) encoder. However, Devlin teaches wherein the text encoder comprises a bidirectional encoder representations from transformers (BERT) encoder (Devlin in Abstract discloses “BERT, which stands for Bidirectional Encoder Representations from Transformers” and states BERT is designed to pre-train “deep bidirectional representations from unlabeled text”). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec by incorporating the BERT text encoder that is taught by Devlin, since both reference are analogous art in the field of contextualized textual embeddings; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec’s audio-visual representation-learning framework with Devlin’s BERT encoder as the text encoder yields the predictable result of generating contextualized textual embeddings from transcript or textual content associated with the media item, thereby improving multimodal media-characteristic determination based on video, audio, and text information. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim(s) 11-12 is/are rejected under 35 U.S.C. 103 as obvious over Akbari et al (Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., & Gong, B. (2021). Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34, 24206-24221, hereafter referred to as Akbari), in view of Mercea et al (Mercea, O. B., Hummel, T., Koepke, A. S., & Akata, Z. (2022, October). Temporal and cross-modal attention for audio-visual zero-shot learning. In European Conference on Computer Vision (pp. 488-505). Cham: Springer Nature Switzerland., hereafter referred to as Mercea), further in view of Kmiec et al (Kmiec, S., Bae, J., & An, R. (2018). Learnable pooling methods for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops (pp. 0-0)., hereafter referred to as Kmiec), further in view of Salamon et al. (US 20210350135 A1, hereafter referred to as Salamon). Claim 11 Regarding Claim 11, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 1. Akbari in view of Mercea, further in view of Kmiec does not explicitly teach all of wherein generating the set of audiovisual embeddings comprises: performing one or more concatenation operations to concatenate a video embedding of the set of video embeddings with an audio embedding of the set of audio embeddings. However, Salamon teaches wherein generating the set of audiovisual embeddings comprises: performing one or more concatenation operations to concatenate a video embedding of the set of video embeddings with an audio embedding of the set of audio embeddings (Salamon in ¶23 discloses the audio feature vector and visual feature vector are fused/merged into an audio-visual vector, and “the audio feature vector 155 and visual feature vector 160 are concatenated in the time dimension. The concatenated feature vector is the audio-visual vector 165.”). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec by incorporating the audio-visual vector concatenation that is taught by Salamon, since both reference are analogous art in the field of audio-visual representation learning; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec’s audio-visual representation-learning framework with Salamon’s concatenation-based audio-visual vector generation yields the predictable result of fusing audio and visual embeddings into a common audiovisual representation, thereby improving downstream classification and characterization of media items. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 12 Regarding Claim 12, Akbari in view of Mercea, further in view of Kmiec, further in view of Salamon teaches The method of claim 11, further comprising: obtaining an output of the one or more concatenation operations (Salamon in ¶23 discloses that the concatenated audio/visual feature vector is the audio-visual vector); performing one or more attention pooling operations to the obtained output of the one or more concatenation operations, wherein the one or more attention pooling operations comprises an audiovisual embedding (Kmiec in Abstract discloses attention mechanisms for aggregating local video descriptors. Section 2.1 discloses the attention representation is created via a weighted sum of local descriptors, with attention clusters concatenated to form a final global representation. Section 3 discloses frame-level video/audio descriptors aggregated into a global representation for classification.); and updating the set of audiovisual embeddings to include the audiovisual embedding (Kmiec in Abstract discloses generating the attention/global representation). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as obvious over Akbari et al (Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., & Gong, B. (2021). Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34, 24206-24221, hereafter referred to as Akbari), in view of Mercea et al (Mercea, O. B., Hummel, T., Koepke, A. S., & Akata, Z. (2022, October). Temporal and cross-modal attention for audio-visual zero-shot learning. In European Conference on Computer Vision (pp. 488-505). Cham: Springer Nature Switzerland., hereafter referred to as Mercea), further in view of Kmiec et al (Kmiec, S., Bae, J., & An, R. (2018). Learnable pooling methods for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops (pp. 0-0)., hereafter referred to as Kmiec), further in view of O'Neill et al. (US 20240412542 A1, hereafter referred to as O'Neill). Claim 13 Regarding Claim 13, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 1. Akbari in view of Mercea, further in view of Kmiec does not explicitly teach all of wherein the one or more media characteristics comprise at least one of: whether the media item is associated with a media trend of a platform. a degree of interest in content of the media item by one or more users of the platform; or at least one of an image quality or an audio quality of the media item. However, O'Neill teaches wherein the one or more media characteristics comprise at least one of: whether the media item is associated with a media trend of a platform. a degree of interest in content of the media item by one or more users of the platform; or at least one of an image quality or an audio quality of the media item. (O'Neill in ¶194 discloses determining visual quality of the video component and categorizing actions or events). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec by incorporating the social-media-content attribute analysis that is taught by O'Neill, since both reference are analogous art in the field of media content and social-media attribute analysis; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec’s audio-visual representation-learning framework with O'Neill’s media-trend, engagement, and quality attribute analysis yields the predictable result of using generated audiovisual embeddings to determine platform-relevant media characteristics, thereby improving automated characterization of media items for content-sharing platforms. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim(s) 14-16 is/are rejected under 35 U.S.C. 103 as obvious over Akbari et al (Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., & Gong, B. (2021). Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34, 24206-24221, hereafter referred to as Akbari), in view of Mercea et al (Mercea, O. B., Hummel, T., Koepke, A. S., & Akata, Z. (2022, October). Temporal and cross-modal attention for audio-visual zero-shot learning. In European Conference on Computer Vision (pp. 488-505). Cham: Springer Nature Switzerland., hereafter referred to as Mercea), further in view of Kmiec et al (Kmiec, S., Bae, J., & An, R. (2018). Learnable pooling methods for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops (pp. 0-0)., hereafter referred to as Kmiec), further in view of Kundu et al. (Nath Kundu, J., Gor, M., Krishna Uppala, P., & Venkatesh Babu, R. (2018). Unsupervised Feature Learning of Human Actions as Trajectories in Pose Embedding Manifold. arXiv e-prints, arXiv-1812., hereafter referred to as Kundu), further in view of O'Neill et al. (US 20240412542 A1, hereafter referred to as O'Neill), further in view of Wang et al. (Wang, L., & Koniusz, P. (2022). Temporal-viewpoint transportation plan for skeletal few-shot action recognition. In Proceedings of the Asian conference on computer vision (pp. 4176-4193)., hereafter referred to as Wang). Claim 14 Regarding Claim 14, Akbari in view of Mercea, further in view of Kmiec teaches The method of claim 1. Akbari in view of Mercea, further in view of Kmiec does not explicitly teach all of wherein the visual features of the sequence of video frames comprise one or more poses of an object depicted by the sequence of video frames, and wherein determining one or more media characteristics associated with the media item comprises: identifying a set of embeddings for an additional media item associated with a media trend, wherein the set of embeddings represent visual features of one or more poses of an additional object depicted by an additional sequence of video frames of the additional media item; determining whether a degree alignment between the one or more poses of the object and the one or more poses of the additional object satisfy one or more alignment criteria based on the set of audiovisual embeddings for the media item and the set of embeddings for the additional media item; and responsive to determining that the degree of alignment satisfies the one or more alignment criteria, determining that the media item is associated with the media trend. However, Kundu teaches wherein the visual features of the sequence of video frames comprise one or more poses of an object depicted by the sequence of video frames (Kundu in Abstract discloses “pose-sequence representation”, “pose embeddings”, and modeling an action as a trajectory in pose embedding space), and wherein determining one or more media characteristics associated with the media item comprises: (Kundu in Abstract discloses pose embeddings / pose trajectories for action sequences). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec by incorporating the pose-embedding representation that is taught by Kundu, since both reference are analogous art in the field of visual feature extraction; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec’s audio-visual representation-learning framework with Kundu’s pose-embedding representation yields the predictable result of representing depicted object or human poses as visual features, thereby improving media characterization where pose or action information is relevant. Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu does not explicitly teach all of wherein determining one or more media characteristics associated with the media item comprises: identifying a set of embeddings for an additional media item associated with a media trend. However, O'Neill teaches wherein determining one or more media characteristics associated with the media item comprises: identifying a set of embeddings for an additional media item associated with a media trend (O'Neill in Abstract and ¶300 discloses trend analysis embeddings / media trends); determining that the media item is associated with the media trend (O'Neill in Abstract and ¶95 discloses media trends/patterns in media content). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec by the media-trend analysis that is taught by O'Neill, since both reference are analogous art in the field of audiovisual embedding characterization; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec’s audio-visual representation-learning framework with O'Neill’s media-trend analysis yields the predictable result of comparing pose-based visual features of a media item with pose-based or media-feature embeddings associated with trending media content, thereby improving identification of media items associated with platform trends. Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill does not explicitly teach all of determining whether a degree alignment between the one or more poses of the object and the one or more poses of the additional object satisfy one or more alignment criteria based on the set of audiovisual embeddings for the media item and the set of embeddings for the additional media item; and responsive to determining that the degree of alignment satisfies the one or more alignment criteria. However, Wang teaches determining whether a degree alignment between the one or more poses of the object and the one or more poses of the additional object satisfy one or more alignment criteria based on the set of audiovisual embeddings for the media item and the set of embeddings for the additional media item (Wang in Abstract discloses aligning query/support sequences of 3D body joints and an advanced Dynamic Time Warping variant for best temporal/viewpoint alignment; Page 2 discloses JEANIE starts and finishes as in DTW and SoftMin picks the smallest distance; Page 3 discloses similarity-based loss encouraging alignment of same-class sequences); and responsive to determining that the degree of alignment satisfies the one or more alignment criteria (Wang in page 3 discloses same-class alignment based on similarity/distance). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill by incorporating the pose-sequence alignment technique that is taught by Wang, since both reference are analogous art in the field of pose-based visual feature analysis; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill’s audio-visual representation-learning framework with Wang’s pose-sequence alignment yields the predictable result of determining whether pose features of a media item align with pose features associated with a media trend, thereby improving automated determination of whether the media item is associated with the trend. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 15 Regarding Claim 15, Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill further in view of Wang teaches The method of claim 14, further comprising: providing the set of audiovisual embeddings for the media item and the set of embeddings for the additional media item as an input to one or more comparison operations (Wang in Abstract discloses query/support sequence features provided to JEANIE for alignment and learning similarity/dissimilarity between pairs of sequences by finding alignment and matching distance); and obtaining, based on one or more outputs of the one or more comparison operations, the degree of alignment between the respective pose of the object and the one or more poses of the additional object (Wang in Abstract discloses jointly modeling temporal/viewpoint alignments and choosing paths based on distance/similarity), wherein the degree of alignment represents a difference between the visual features of the one or more embeddings for the media item and the set of embeddings for the additional media item (Wang in page 5 discloses matching distance between query/support sequence features and base distance for temporal-viewpoint soft-DTW alignment). Claim 16 Regarding Claim 16, Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill further in view of Wang teaches The method of claim 15, wherein the one or more comparison operations comprise a dynamic time warping function (Wang in Abstract and page 2 discloses “an advanced variant of Dynamic Time Warping”, and further teaches building on soft-DTW, a differentiable variant of DTW). Claim(s) 17 is/are rejected under 35 U.S.C. 103 as obvious over Akbari et al (Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., & Gong, B. (2021). Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34, 24206-24221, hereafter referred to as Akbari), in view of Mercea et al (Mercea, O. B., Hummel, T., Koepke, A. S., & Akata, Z. (2022, October). Temporal and cross-modal attention for audio-visual zero-shot learning. In European Conference on Computer Vision (pp. 488-505). Cham: Springer Nature Switzerland., hereafter referred to as Mercea), further in view of Kmiec et al (Kmiec, S., Bae, J., & An, R. (2018). Learnable pooling methods for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops (pp. 0-0)., hereafter referred to as Kmiec), further in view of Kundu et al. (Nath Kundu, J., Gor, M., Krishna Uppala, P., & Venkatesh Babu, R. (2018). Unsupervised Feature Learning of Human Actions as Trajectories in Pose Embedding Manifold. arXiv e-prints, arXiv-1812., hereafter referred to as Kundu), further in view of O'Neill et al. (US 20240412542 A1, hereafter referred to as O'Neill), further in view of Wang et al. (Wang, L., & Koniusz, P. (2022). Temporal-viewpoint transportation plan for skeletal few-shot action recognition. In Proceedings of the Asian conference on computer vision (pp. 4176-4193)., hereafter referred to as Wang), further in view of Reblitz-Richardson et al. (US 20140280241 A1, hereafter referred to as Reblitz-Richardson). Claim 17 Regarding Claim 17, Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill further in view of Wang teaches The method of claim 14, wherein determining whether the degree of alignment satisfies the one or more alignment criteria. Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill further in view of Wang does not explicitly teach all of determining whether a difference between the visual features of the set of audiovisual embeddings for the media item and the set of embeddings for the additional media item falls below a difference threshold. However, Reblitz-Richardson teaches determining whether a difference between the visual features of the set of audiovisual embeddings for the media item and the set of embeddings for the additional media item falls below a difference threshold (Reblitz-Richardson in ¶69-73 discloses comparing media items using a similarity threshold and grouping/parenting media items if the threshold is passed). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill further in view of Wang by incorporating the threshold-based media-item similarity determination that is taught by Reblitz-Richardson, since both reference are analogous art in the field of media-item analysis; thus, one of ordinary skilled in the art would be motivated to combine the references since Akbari in view of Mercea, further in view of Kmiec, further in view of Kundu, further in view of O'Neill further in view of Wang’s audio-visual representation-learning framework with Reblitz-Richardson’s threshold-based similarity comparison yields the predictable result of determining whether the difference or distance between pose-based visual features is sufficiently small to satisfy the alignment criteria, thereby enabling automated determination that a media item is associated with a media trend when its features are sufficiently similar to trend-associated media features. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN P CASCAIS whose telephone number is (703) 756-5576. The examiner can normally be reached Monday-Friday 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mr. O'Neal Mistry can be reached on (313) 446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.P.C./Examiner, Art Unit 2674 /ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674 Date: 7/2/2026
Read full office action

Prosecution Timeline

Sep 27, 2024
Application Filed
Jul 14, 2026
Non-Final Rejection mailed — §103, §112, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694074
METHOD AND DEVICE FOR ASCERTAINING A CLASSIFICATION AND/OR A REGRESSION RESULT WHEN MISSING SENSOR DATA
4y 3m to grant Granted Jul 28, 2026
Patent 12694608
SELECTING REPRESENTATIVE IMAGE VIEWS FOR 3D OBJECT MODELS IN SYNTHETIC CONTENT CREATION SYSTEMS AND APPLICATIONS
3y 7m to grant Granted Jul 28, 2026
Patent 12694525
GENERATIVE ADVERSARIAL NETWORK-BASED LOSSLESS IMAGE COMPRESSION MODEL FOR CROSS-SECTIONAL IMAGING
2y 6m to grant Granted Jul 28, 2026
Patent 12688703
SMART ROAD SURFACE DETECTION METHOD AND EDGE COLLECTION DEVICE, CLOUD-BASED ROAD SURFACE RECOGNITION MODULE AND SYSTEM THEREOF
3y 4m to grant Granted Jul 21, 2026
Patent 12682487
IMAGE RECOGNITION DEVICE, METHOD FOR IMAGE RECOGNITION DEVICE, AND RECORDING MEDIUM
2y 9m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
89%
With Interview (+13.7%)
2y 10m (~12m remaining)
Median Time to Grant
Low
PTA Risk
Based on 64 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month