DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 07/22/2025 and 02/18/2025 have been entered and considered. Initialed copies of the PTO-1449 by the Examiner are attached.
Claim Objections
Claims 1, 13 and 16 are objected to because of the following informalities: (a) Amend claim 1 as suggested below to correct the multiple run on sentences.
generating, by using the relational attention module and on the basis of query features, a salient query feature set corresponding to the query feature;
updating, by using the relational attention module on the basis of the salient query feature set, the query features;
acquiring, by using the cross-attention module and on the basis of the updated query features, predicted segment quality information corresponding to the updated query features;
and constructing a segment quality loss function according to the predicted segment quality information;
(b) Amend claim 13 as suggested below to correct the multiple run on sentences.
generating, by using the relational attention module and on the basis of query features, a salient query feature set corresponding to the query feature;
updating, by using the relational attention module on the basis of the salient query feature set, the query features;
acquiring, by using the cross-attention module and on the basis of the updated query features, predicted segment quality information corresponding to the updated query features;
and constructing a segment quality loss function according to the predicted segment quality information;
(c) Amend claim 16 as suggested below to correct the multiple run on sentences.
generating, by using the relational attention module and on the basis of query features, a salient query feature set corresponding to the query feature;
updating, by using the relational attention module on the basis of the salient query feature set, the query features;
acquiring, by using the cross-attention module and on the basis of the updated query features, predicted segment quality information corresponding to the updated query features;
and constructing a segment quality loss function according to the predicted segment quality information;
Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function.
Claims 1, 13 and 16 recites limitations that use words like “means” (or “step”) or similar terms with functional language but do invoke 35 U.S.C. 112(f):
Claims 1, 13 and 16; recites the limitation, “relational attention module …perform……,” [Line 3] in claim 1, [Line 7] in claim 13 and [Line 5] of claim 16.
Claims 1, 13 and 16; recites the limitation, “cross-attention module ...predicted ……,” [Line 7] in claim 1, [Line 11] in claim 13 and [Line 9] of claim 16.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 1, 13 and 16;
(i) “cross-attention module” (Paragraph [0067-0068][0116] – the algorithm disclose in [p][0067-0068] predicts using each query feature a corresponding temporal segment through a fully connected layer. For the reference query segment 321, a corresponding salient query feature set includes the salient similar segments 331, 332, 333, and query features in a similar feature set have features such as semantic similarity, non-redundancy in a time dimension. According to the similarity information between the query features, a similarity matrix A∈RL q ×L q is constructed, wherein Lq is a fixed number of the query features, A is a similarity matrix for characterizing a similarity between every two of the Lq query features, and each element in the similarity matrix A is a cosine similarity between two query features. On the basis of the similarity threshold γ∈[−1, 1], the similar feature set is constructed by: Esim={(i,j)❘A[i,j]-γ>0};(1-1) where A[i,j] is a similarity between an i-th query feature and a j-th query feature, γ is a pre-defined similarity threshold before training, and Esim is a similar feature set constructed according to the similarity between the features, and can correspond to the query features, so that there are a plurality of Esim. . Further in [p][0116], memory 131 is used for storing instructions, and the processor 132, which is coupled to the memory 131, is configured to implement the algorithm, on the basis of the instructions stored in the memory 131. Thus, the disclosed module have sufficient structure or material wherein is an algorithm executed by processor.
(ii) “relational attention module” (Paragraph [0089][0116] – the algorithm disclose in [p][0089] predicts sampled points in a time dimension, and obtains features of video segments by weighting and summing the sampled features, and the features of the video segments are sent into detection heads through a feedforward network. In addition to existing regression and classification heads, a segment quality head is added to estimate the quality of the segment. Further in [p][0116], memory 131 is used for storing instructions, and the processor 132, which is coupled to the memory 131, is configured to implement the algorithm, on the basis of the instructions stored in the memory 131. Thus, the disclosed module have sufficient structure or material wherein is an algorithm executed by processor.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-16 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites “using the relational attention module and on the basis of query features” in line 3; however, the limitation “using the relational attention module” is indefinite because it merely recites a use without any active, positive steps delimiting how this use is actually practiced or the structure associated with it. Ex parteErlich, 3 USPQ2d 1011 (Bd. Pat. App. & Inter. 1986).
Claim 1 recites “using the relational attention module” in line 7; however, the limitation “using the relational attention module” is indefinite because it merely recites a use without any active, positive steps delimiting how this use is actually practiced or the structure associated with it. Ex parteErlich, 3 USPQ2d 1011 (Bd. Pat. App. & Inter. 1986).
Claim 13 recites “using the relational attention module and on the basis of query features” in line 7; however, the limitation “using the relational attention module” is indefinite because it merely recites a use without any active, positive steps delimiting how this use is actually practiced or the structure associated with it. Ex parteErlich, 3 USPQ2d 1011 (Bd. Pat. App. & Inter. 1986).
Claim 13 recites “using the relational attention module” in line 11; however, the limitation “using the relational attention module” is indefinite because it merely recites a use without any active, positive steps delimiting how this use is actually practiced or the structure associated with it. Ex parteErlich, 3 USPQ2d 1011 (Bd. Pat. App. & Inter. 1986).
Claim 16 recites “using the relational attention module and on the basis of query features” in line 5; however, the limitation “using the relational attention module” is indefinite because it merely recites a use without any active, positive steps delimiting how this use is actually practiced or the structure associated with it. Ex parteErlich, 3 USPQ2d 1011 (Bd. Pat. App. & Inter. 1986).
Claim 16 recites “using the relational attention module” in line 9; however, the limitation “using the relational attention module” is indefinite because it merely recites a use without any active, positive steps delimiting how this use is actually practiced or the structure associated with it. Ex parteErlich, 3 USPQ2d 1011 (Bd. Pat. App. & Inter. 1986).
Claims 2-12, 14-15 and 17-22 are being rejected as incorporating the deficiencies of the claim upon which each respective claim depends.
Allowable Subject Matter
Independent claims 1, 13 and 16 along with its dependent claims 2-12, 14-15 and 16-22would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action.
The following is a statement of reasons for the indication of allowable subject matter:
The closest prior art of record is Xu et al (NPL titled: Multi-Task Learning with Multi-query Transformer for Dense Prediction)
Regarding claim 1, Xu teaches a method for training a decoder (we introduce MQTransformer for multitask learning of five dense prediction tasks, including semantic segmentation, depth estimation, surface normals prediction, semantic edge detection, and saliency detection – see section 1, [p][004]), wherein the decoder comprises a relational attention module (cross-scale learning – see section 1, [p][004]) and a cross-attention module (cross-task attention module – see section 3.2, [p][001] and Fig 2), and the method comprises: generating, by using the relational attention module and on the basis of query features, a salient query feature set corresponding to the query features, to perform, by using the relational attention module and on the basis of the salient query feature set, updating processing on the query features (we introduce MQTransformer for multitask learning of five dense prediction tasks, including semantic segmentation, depth estimation, surface normals prediction, semantic edge detection, and saliency detection. As illustrated in Fig. 1(c), we introduce multiple taskrelevant queries (according to task number) as the input of the encoder. The encoder outputs the learnt task-relevant query feature of each task. We show two tasks for illustration purposes. Moreover, we add cross-scale learning in modeling these task queries. Depending on the number of tasks, these query features are concatenated into two sets, i.e., query features of the same scale with different tasks are concatenated, and query features of the same task with different scales are concatenated. The concatenated query features collect representation across scale and task - see section 1, [p][004]);
However, Xu does not explicitly teach acquiring, by using the cross-attention module and on the basis of the updated query features, predicted segment quality information corresponding to the updated query features, and constructing a segment quality loss function according to the predicted segment quality information; acquiring segment relation features between predicted video segments corresponding to the query features, and constructing a segment relation loss function; and performing adjustment processing on the relational attention module and the cross-attention module according to the segment quality loss function and the segment relation loss function.
Wherein the cross-attention module and relational attention module are define as interpreted by 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph below;
(i) “cross-attention module” (Paragraph [0067-0068][0116] – the algorithm disclose in [p][0067-0068] predicts using each query feature a corresponding temporal segment through a fully connected layer. For the reference query segment 321, a corresponding salient query feature set includes the salient similar segments 331, 332, 333, and query features in a similar feature set have features such as semantic similarity, non-redundancy in a time dimension. According to the similarity information between the query features, a similarity matrix A∈RL q ×L q is constructed, wherein Lq is a fixed number of the query features, A is a similarity matrix for characterizing a similarity between every two of the Lq query features, and each element in the similarity matrix A is a cosine similarity between two query features. On the basis of the similarity threshold γ∈[−1, 1], the similar feature set is constructed by: Esim={(i,j)❘A[i,j]-γ>0};(1-1) where A[i,j] is a similarity between an i-th query feature and a j-th query feature, γ is a pre-defined similarity threshold before training, and Esim is a similar feature set constructed according to the similarity between the features, and can correspond to the query features, so that there are a plurality of Esim. . Further in [p][0116], memory 131 is used for storing instructions, and the processor 132, which is coupled to the memory 131, is configured to implement the algorithm, on the basis of the instructions stored in the memory 131. Thus, the disclosed module have sufficient structure or material wherein is an algorithm executed by processor.
(ii) “relational attention module (Paragraph [0089][0116] – the algorithm disclose in [p][0089] predicts sampled points in a time dimension, and obtains features of video segments by weighting and summing the sampled features, and the features of the video segments are sent into detection heads through a feedforward network. In addition to existing regression and classification heads, a segment quality head is added to estimate the quality of the segment. Further in [p][0116], memory 131 is used for storing instructions, and the processor 132, which is coupled to the memory 131, is configured to implement the algorithm, on the basis of the instructions stored in the memory 131. Thus, the disclosed module have sufficient structure or material wherein is an algorithm executed by processor.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Lin et al (Pub No.: 20220301298 ) discloses a methods, systems, and apparatus, including computer programs encoded on computer storage media, for training an image representation neural network.
Shi et al (NPL titled: ReAct: Temporal Action Detection with
Relational Queries) This work aims at advancing temporal action detection (TAD)
using an encoder-decoder framework with action queries, similar to DETR,
which has shown great success in object detection. However, the framework
suffers from several problems if directly applied to TAD: the insufficient
exploration of inter-query relation in the decoder, the inadequate classification training due to a limited number of training samples, and the unreliable classification scores at inference. To this end, we first propose a relational attention mechanism in the decoder, which guides the attention among queries based on their relations. Moreover, we propose two losses to facilitate and stabilize the training of action classification.
Lastly, we propose to predict the localization quality of each action query at inference in order to distinguish high-quality queries. The proposed method, named ReAct, achieves the state-of-the-art performance on THUMOS14, with much lower computational costs than previous methods. Besides, extensive ablation studies are conducted to
verify the effectiveness of each proposed component.
Lu et al (NPL titled: End-to-End Temporal Action Detection
With Transformer) discloses a temporal action detection (TAD) aims to determine the semantic label and the temporal interval of every actioninstance in an untrimmed video. It is a fundamental and challenging task in video understanding. Previous methods tackle this task with complicated pipelines. They often need to train multiple networks and involve hand-designed operations, such as non-maximal suppression and anchor generation, which limit the flexibility and prevent end-to-end learning. In this paper,
we propose an end-to-end Transformer-based method for TAD, termed TadTR. Given a small set of learnable embeddings called action queries, TadTR adaptively extracts temporal context information from the video for each query and directly predicts
action instances with the context. To adapt Transformer to TAD, we propose three improvements to enhance its locality awareness. The core is a temporal deformable attention module that selectively attends to a sparse set of key snippets in a video. A segment refinement mechanism and an actionness regression head are designed to refine the boundaries and confidence of the predicted instances, respectively.
Wang et al (Pub No.: US20230154170A1) discloses a method, apparatus, electronic device, and non-transitory computer-readable storage medium with multi-modal feature fusion are provided. The method includes generating three-dimensional (3D) feature information and two-dimensional (2D) feature information based on a color image and a depth image, generating fused feature information by fusing the 3D feature information and the 2D feature information based on an attention mechanism, and generating predicted image information by performing image processing based on the fused feature information.
Inquiries
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDRAE S ALLISON whose telephone number is (571)270-1052. The examiner can normally be reached on Monday-Friday 9am-5pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns, can be reached on (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDRAE S ALLISON/Primary Examiner, Art Unit 2673
August 20, 2026