Prosecution Insights
Last updated: October 02, 2026
Application No. 19/017,116

METHODS AND APPARATUSES FOR DETERMINING SIMILARITY BETWEEN TEXT AND VIDEO

Non-Final OA §102§103
Filed
Jan 10, 2025
Priority
Jan 11, 2024 — CN 202410044723.5
Examiner
DEPALMA, CAROLINE ELIZABETH
Art Unit
Tech Center
Assignee
Alipay.com Co., Ltd.
OA Round
1 (Non-Final)
90%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
52 granted / 58 resolved
+29.7% vs TC avg
Moderate +7% lift
Without
With
+7.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
15 currently pending
Career history
68
Total Applications
across all art units

Statute-Specific Performance

§101
13.1%
-26.9% vs TC avg
§103
42.2%
+2.2% vs TC avg
§102
18.6%
-21.4% vs TC avg
§112
21.9%
-18.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 58 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 22-24, 28, 31-34 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chen (S. Chen, Y. Zhao, Q. Jin and Q. Wu, "Fine-Grained Video-Text Retrieval With Hierarchical Graph Reasoning," 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 10635-10644, doi: 10.1109/CVPR42600.2020.01065.). Regarding claim 22, Chen discloses a computer-implemented method for determining a similarity between text and a video ([pg. 10637, 3. Hierarchical Graph Reasoning Model] video-text matching which aggregates global and local matchings at different levels to compute overall cross-modal similarities), comprising: respectively providing text and a video that are comprised in an acquired text-video pair for a text feature extraction model and a video feature extraction model (Fig. 2, [pg. 10637, 3. Hierarchical Graph Reasoning Model] Figure 2 illustrates the overview of the HGR model which consists of three blocks, including hierarchical textual encoding and hierarchical video encoding; [pg. 10637, 3.1. Hierarchical Textual Encoding] obtain hierarchical textual representations from a video description), to obtain a corresponding initial text feature and a corresponding initial video feature (Fig. 2; [pg. 10637, 3.1. Hierarchical Textual Encoding] given a video description C that consists of N words, we consider C as a global event node in the hierarchical graph, then we employ semantic role parsing toolkit to obtain verbs, noun phrases in C as well as the semantic role of each noun phrase to the corresponding verb; [pg. 10638, 3.2. Hierarchical Video Encoding] given video V as a sequence of frame-wise features, we utilize different weights to encode videos into three levels of embeddings…for the global event level we employ the attention mechanism similar to Eq 4 to obtain one global vector to represent the salient event in the video), wherein the initial text feature comprises word character features corresponding to word characters comprised in the text ([pg. 10637, 3.1. Hierarchical Textual Encoding] given a video description C that consists of N words, we consider C as a global event node in the hierarchical graph, then we employ semantic role parsing toolkit to obtain verbs, noun phrases in C as well as the semantic role of each noun phrase to the corresponding verb), and the initial video feature comprises an image feature extracted based on an image comprised in the video ([pg. 10638, 3.2. Hierarchical Video Encoding] given video V as a sequence of frame-wise features, we utilize different weights to encode videos into three levels of embeddings…for the global event level we employ the attention mechanism similar to Eq 4 to obtain one global vector to represent the salient event in the video); performing syntactic analysis on the text, to obtain a syntactic level analysis result (Fig. 3, [pg. 10641, first paragraph] In Figure 3, we present a learned pattern on how action nodes interacting with neighbor nodes in graph reasoning at different layers, which is strongly relevant to semantic roles (see also Fig. 2 and [pg. 10637, Semantic Role Graph Structure])); processing the initial text feature based on the syntactic level analysis result, to obtain text features respectively corresponding to elements in the syntactic level analysis result (Fig. 2, 3; [pg. 10637, Semantic Role Graph Structure] the verbs are considered as action nodes and connected to event node with direct edges so that temporal relations of different actions can be implicitly learned from event node in following graph reasoning, the noun phrases are entity nodes that are connected with different action nodes); constructing, based on a degree of matching between the text features respectively corresponding to the elements in the syntactic level analysis result and the initial video feature, a video level analysis result corresponding to the syntactic level analysis result (Fig. 2, [pg. 10638-10639, 3.3. Video-Text Matching] Global matching: the video and text are encoded into global vectors that capture salient event semantics with attention mechanism…global matching score…an alignment between cross-modal local components is supposed to be learned to compute overall matching score); processing, based on the video level analysis result, initial video features corresponding to elements in the video level analysis result, to obtain video features respectively corresponding to the elements in the video level analysis result (Fig. 2; [pg. 10639, Local Attentive Matching] equation 12, attention weights over video frames for each local textual node which aligns [words] to video frames, we then compute the similarity between [words] and [frame features] as weighted average of local similarities); and determining a similarity between the text and the video based on a similarity between a text feature and a video feature that respectively correspond to elements in a corresponding level (Fig. 2; [pg. 10639, Training and Inference] we take the average of cross-modal similarities at all levels as final video-text similarity (equation 13)). Regarding claim 23, Chen discloses the computer-implemented method according to claim 22 as applied above. Chen further discloses wherein the elements in the syntactic level analysis result comprise a sentence node located at a first level and an action node located at a second level (Fig. 2; [pg. 10637, 3.1. Hierarchical Textual Encoding] the overall sentence describes the global event in the video which is composed of multiple actions in temporal dimensions, and each action is composed of different entities as its arguments; [pg. 10637, Semantic Role Graph Structure] the verbs are considered as action nodes and connected to event node with direct edges, so that temporal relations of different actions can be implicitly learned from event node in following graph reasoning); and the elements in the video level analysis result comprise a video node located at a first level and a frame node located at a second level, wherein the frame node corresponds to a video frame group, and each video frame in the video frame group matches the action node (Fig. 2; [pg. 10638, 3.1. Hierarchical Video Encoding] Given video V as a sequence of frame-wise features, we utilize different weights to encode videos into three level of embeddings (see equation 11), for the global event level we employ the attention mechanism similar to Eq 4 to obtain one global vector to represent the salient event in the video, and for the action and entity level, the video representations are a sequence of frame-wise features va and vo respectively, these features will be sent to the following matching module to match with their corresponding textual features at different levels). Regarding claim 24, Chen discloses the computer-implemented method according to claim 23 as applied above. Chen further discloses wherein the elements in the syntactic level analysis result further comprise an entity node located at a third level (Fig. 2; [pg. 10637, Semantic Role Graph Structure] the noun phrases are entity nodes that are connected with different action nodes…if an entity node serves multiple semantic roles to different action nodes we duplicate the entity node for each semantic role); and the elements in the video level analysis result further comprise an image patch node located at a third level, wherein the image patch node corresponds to an image patch group, and each image patch in the image patch group matches the entity node and belongs to a video frame in a corresponding video frame group (Fig. 2; [pg. 10638, 3.1. Hierarchical Video Encoding] for the action and entity level, the video representations are a sequence of frame-wise features; see also Fig. 4 and [pg. 10642, 4.6. Qualitative Results]). Regarding claim 28, Chen discloses the computer-implemented method according to claim 23 as applied above. Chen further discloses wherein the initial video feature comprises a frame feature corresponding to a video frame ([pg. 10638, 3.2. Hierarchical Video Encoding] given video V as a sequence of frame-wise features, we utilize different weights to encode videos into three levels of embeddings); and the processing, based on the video level analysis result, initial video features corresponding to elements in the video level analysis result, to obtain video features respectively corresponding to the elements in the video level analysis result (Fig. 2; [pg. 10639, Local Attentive Matching] equation 12, attention weights over video frames for each local textual node which aligns [words] to video frames, we then compute the similarity between [words] and [frame features] as weighted average of local similarities) comprises: determining, based on a degree of matching between frame features corresponding to video frames and a text feature corresponding to the sentence node, fusion coefficients corresponding to the frame features ([pg. 10639, Local Attentive Matching] the [equation 12] is then utilized as attention weights over video frames for each local textual node i, which dynamically aligns [words] to video frames, we then compute the similarity between [words] and [video frames] as weighted average of local similarities); and fusing the frame features based on the fusion coefficients, to obtain a video feature corresponding to the video node ([pg. 10639, Local Attentive Matching] the [equation 12] is then utilized as attention weights over video frames for each local textual node i, which dynamically aligns [words] to video frames; [pg. 10642, first paragraph] due to the fusion of hierarchical levels from global to local, our model can select the more comprehensive sentence). Regarding claim 31, Chen discloses the computer-implemented method according to claim 22 as applied above. Chen further discloses wherein the determining a similarity between the text and the video based on a similarity between a text feature and a video feature that respectively correspond to elements in a corresponding level (Fig. 2; [pg. 10639, Training and Inference] we take the average of cross-modal similarities at all levels as final video-text similarity (equation 13)) comprises: determining the similarity between the text and the video by performing weighted summation on a similarity between text features and video features that respectively correspond to elements in all levels ([pg. 10639, Local Attentive Matching] equation 12, attention weights over video frames for each local textual node which aligns [words] to video frames, we then compute the similarity between [words] and [frame features] as weighted average of local similarities…the final matching similarity summarizes all local component similarities of text). Regarding claim 32, Chen discloses the computer-implemented method according to claim 31 as applied above. Chen further discloses wherein a weight corresponding to each element in each level is determined based on normalization of text features or video features corresponding to elements in the level ([pg. 10639, Local Attentive Matching] therefore we normalize [local similarities between each pair of cross-modal local components] inspired by stacked attention as [equation 12]). Regarding claim 33, Chen discloses the computer-implemented method according to claim 32 as applied above. Chen further discloses receiving query text provided by a user (Fig. 4; [pg. 10642, 4.6. Qualitative Results] text-to-video retrieval as in Figure 4…our model successfully retrieves the correct video which contains all actions and entities described in the sentence); determining a query similarity between the query text and a candidate video comprised in each query text-video pair, wherein the each query text-video pair is obtained based on the query text and each candidate video in a candidate video set (Fig. 4; [pg. 10642, 4.6. Qualitative Results] text-to-video retrieval as in Figure 4…our model successfully retrieves the correct video which contains all actions and entities described in the sentence; Fig. 2; [pg. 10639, Training and Inference] we take the average of cross-modal similarities at all levels as final video-text similarity (equation 13)); determining, from the candidate video set based on the query similarity, a matching video as a video search result (Fig. 4; [pg. 10642, 4.6. Qualitative Results] text-to-video retrieval as in Figure 4…our model successfully retrieves the correct video which contains all actions and entities described in the sentence; Fig. 2; [pg. 10639, Training and Inference] we take the average of cross-modal similarities at all levels as final video-text similarity (equation 13)); and providing the video search result for the user (Fig. 4; [pg. 10642, 4.6. Qualitative Results] text-to-video retrieval as in Figure 4…our model successfully retrieves the correct video which contains all actions and entities described in the sentence). Regarding claim 34, Chen discloses the computer-implemented method according to claim 32 as applied above. Chen further discloses receiving a query video provided by a user (Fig. 5; [pg. 10642, 4.6. Qualitative Results] in figure 5, we provide qualitative results on video-to-text retrieval as well, which demonstrate the effectiveness of our HGR model for cross-modal retrieval on both directions); determining a query similarity between the query video and candidate text comprised in each query text-video pair, wherein the each query text-video pair is obtained based on the query video and each piece of candidate text in a candidate text set (Fig. 5; [pg. 10642, 4.6. Qualitative Results] in figure 5, we provide qualitative results on video-to-text retrieval as well, which demonstrate the effectiveness of our HGR model for cross-modal retrieval on both directions; Fig. 2; [pg. 10639, Training and Inference] we take the average of cross-modal similarities at all levels as final video-text similarity (equation 13)); determining, from the candidate text set based on the query similarity, matching text as a text search result (Fig. 5; [pg. 10642, 4.6. Qualitative Results] in figure 5, we provide qualitative results on video-to-text retrieval as well, which demonstrate the effectiveness of our HGR model for cross-modal retrieval on both directions; Fig. 2; [pg. 10639, Training and Inference] we take the average of cross-modal similarities at all levels as final video-text similarity (equation 13)); and providing the text search result for the user (Fig. 5; [pg. 10642, 4.6. Qualitative Results] in figure 5, we provide qualitative results on video-to-text retrieval as well, which demonstrate the effectiveness of our HGR model for cross-modal retrieval on both directions). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 35-37, 39-41 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen (S. Chen, Y. Zhao, Q. Jin and Q. Wu, "Fine-Grained Video-Text Retrieval With Hierarchical Graph Reasoning," 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 10635-10644, doi: 10.1109/CVPR42600.2020.01065.). Regarding claim 35, Chen discloses everything claimed as applied above (see rejection of claim 22); however Chen fails to explicitly disclose one or more processors; and one or more tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more processors, perform operations. However, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Chen with generic computer components including one or more processors; and one or more tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more processors, perform operations, for the purpose of effectively implementing the method in an application setting such as video retrieval in user search/query and internet contexts such as YouTube and other social media sites (See Chen: 1. Introduction). Regarding claim 36, Chen discloses the computer-implemented system according to claim 35 as applied above. Chen further discloses everything claimed as applied above (see rejection of claim 23). Regarding claim 37, Chen discloses the computer-implemented system according to claim 36 as applied above. Chen further discloses everything claimed as applied above (see rejection of claim 24). Regarding claim 39, Chen discloses everything claimed as applied above (see rejection of claim 22); however Chen fails to explicitly disclose a non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations. It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Chen with generic computer components including non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations, for the purpose of effectively implementing the method in an application setting such as video retrieval in user search/query and internet contexts such as YouTube and other social media sites (See Chen: 1. Introduction). Regarding claim 40, Chen discloses the non-transitory, computer-readable medium according to claim 39 as applied above. Chen further discloses everything claimed as applied above (see rejection of claim 23). Regarding claim 41, Chen discloses the non-transitory, computer-readable medium according to claim 39 as applied above. Chen further discloses everything claimed as applied above (see rejection of claim 31). Allowable Subject Matter Claims 25-27, 29-30, 38 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 25, Chen discloses the computer-implemented method according to claim 24 as applied above. However Chen fails to disclose wherein the elements in the syntactic level analysis result further comprise an attribute node located at a fourth level; and the processing the initial text feature based on the syntactic level analysis result, to obtain text features respectively corresponding to elements in the syntactic level analysis result comprises: respectively extracting, from the initial text feature, initial text features corresponding to the elements in the syntactic level analysis result, to obtain text features corresponding to the sentence node and the action node; and for each entity node, performing, based on an initial text feature corresponding to an attribute node associated with the entity node, feature enhancement on an initial text feature corresponding to the entity node, to obtain a text feature corresponding to each entity node. Regarding claim 26, Chen discloses the computer-implemented method according to claim 23 as applied above. Chen further discloses wherein the initial video feature comprises a frame feature corresponding to a video frame (Fig. 2; [pg. 10638, 3.1. Hierarchical Video Encoding] Given video V as a sequence of frame-wise features, we utilize different weights to encode videos into three level of embeddings (see equation 11)). However Chen fails to disclose the constructing, based on a degree of matching between the text features respectively corresponding to the elements in the syntactic level analysis result and the initial video feature, a video level analysis result corresponding to the syntactic level analysis result comprises: for each action node, determining a degree of matching between a text feature corresponding to the action node and each time encoding feature; and providing, for a time encoding model, an obtained frame feature corresponding to a video frame, to obtain a time encoding feature that fuses with time information and that corresponds to each frame feature; and selecting a 1st quantity of video frames corresponding to time encoding features whose degrees of matching satisfy a first predetermined need, to constitute the video frame group, to obtain a frame node that is located at the second level and that corresponds to the action node. Claim 29 is dependent on claim 26 and thus similar reasoning applies. Claim 38 is directed to similar subject matter as claim 26 and thus similar reasoning applies. Regarding claim 27, Chen discloses the computer-implemented method according to claim 24 as applied above. Chen further discloses wherein the initial video feature comprises an image patch feature corresponding to an image patch obtained through division of the video frame (Fig. 2; [pg. 10638, 3.1. Hierarchical Video Encoding] Given video V as a sequence of frame-wise features, we utilize different weights to encode videos into three level of embeddings (see equation 11)); and the constructing, based on a degree of matching between the text features respectively corresponding to the elements in the syntactic level analysis result and the initial video feature, a video level analysis result corresponding to the syntactic level analysis result (Fig. 2, [pg. 10638-10639, 3.3. Video-Text Matching] Global matching: the video and text are encoded into global vectors that capture salient event semantics with attention mechanism…global matching score…an alignment between cross-modal local components is supposed to be learned to compute overall matching score) comprises: for each frame node, determining a degree of matching between a text feature corresponding to an entity node corresponding to the frame node and an image patch feature corresponding to an image patch obtained through division of each video frame in a video frame group corresponding to the frame node ([pg. 10638-10639, Attentive Matching] at the action and entity level, there are multiple local components in the video and text, therefore an alignment between cross-modal local components is supposed to be learned to compute overall matching score…such local similarities implicitly reflect the alignment between local texts and videos such as how strong a text node is relevant toa video frame). However Chen fails to disclose selecting a 2nd quantity of image patches corresponding to image patch features whose degrees of matching satisfy a second predetermined need, to constitute an image patch group, to obtain an image patch node that is located at the third level and that is connected to the frame node. Claim 30 is dependent on claim 27 and thus similar reasoning applies. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Wray (Wray, Michael, et al. "Fine-grained action retrieval through multiple parts-of-speech embeddings." Proceedings of the IEEE/CVF international conference on computer vision. 2019.) discloses determining parts of speech in video captions in order to retrieve specific actions for the purpose of cross-modal (text and video) search tasks, wherein the parts of speech may include adjectives included in the caption sentence. Zhang (CN 117609553 A) discloses a video retrieval method including obtaining encoded frame image features and text global and keyword features and using a hierarchical matching strategy. Park (US 20250078486 A1) discloses extracting features from video and text and inputting to local and global fusion modules, including features corresponding to time points or time periods, to determine similarity. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAROLINE DEPALMA whose telephone number is (571)270-0769. The examiner can normally be reached Mon-Thurs 9:00am-4pm Eastern Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CAROLINE E. DEPALMA/Examiner, Art Unit 2675 /SJ Park/Primary Examiner, Art Unit 2675
Read full office action

Prosecution Timeline

Jan 10, 2025
Application Filed
Sep 11, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12736457
SYSTEM AND METHOD OF EVALUATING FLUIDIC SUBSTANCE
2y 5m to grant Granted Sep 15, 2026
Patent 12731286
ESTIMATION OF DENSITY DISTORTION METRIC FOR PROCESSING OF POINT CLOUD GEOMETRY
3y 4m to grant Granted Sep 08, 2026
Patent 12731368
OBJECT DISCRIMINATION DEVICE
2y 9m to grant Granted Sep 08, 2026
Patent 12724476
ELECTRONIC DEVICE
2y 8m to grant Granted Sep 01, 2026
Patent 12705927
METHOD FOR SELF-MEASURING FACIAL OR CORPORAL DIMENSIONS, NOTABLY FOR THE MANUFACTURING OF PERSONALIZED APPLICATORS
3y 10m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
90%
Grant Probability
97%
With Interview (+7.3%)
2y 8m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 58 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month