Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01-21-2026; 06-24-2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under U.S.C 101 for containing an abstract idea without significantly more.
Regarding claim 1:
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is a process.
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
determining a particular candidate sequence as a match between the transformed candidate sequences and the encoded textual prompt based on the contrastive loss; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
each of the candidate sequences being evaluated against the textual prompt based on a contrastive loss; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No, there are no additional elements that integrate the judicial exception into a practical application. The additional elements:
one or more processors; and This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
a memory storing instructions that, when executed by the one or more processors, cause the system to perform: This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
obtaining a textual prompt; This limitation is directed to insignificant extra solution activity (mere data gathering, as per MPEP 2106.05(g))
encoding the textual prompt; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
obtaining candidate sequences of sensor data from different modalities, each sequence comprising a plurality of sequential frames, This limitation is directed to insignificant extra solution activity (mere data gathering, as per MPEP 2106.05(g))
encoding the candidate sequences of sensor data; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
embedding position information, within the encoded candidate sequences, indicating relative timestamps associated with each of the sequential frames of the sensor data; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
concatenating the encoded candidate sequences of sensor data, including the embedded position information Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
transforming the concatenated and encoded frames of sensor data to form transformed candidate sequences Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
generating a hierarchical structure that encapsulates navigation data of the particular candidate sequence. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No, there are no additional elements that amount to significantly more than the judicial exception. The additional elements are:
one or more processors; and This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
a memory storing instructions that, when executed by the one or more processors, cause the system to perform: This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
obtaining a textual prompt; This limitation is directed to insignificant extra solution activity (mere data gathering, as per MPEP 2106.05(g))
encoding the textual prompt; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
obtaining candidate sequences of sensor data from different modalities, each sequence comprising a plurality of sequential frames, This limitation is directed to insignificant extra solution activity (mere data gathering, as per MPEP 2106.05(g))
encoding the candidate sequences of sensor data; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
embedding position information, within the encoded candidate sequences, indicating relative timestamps associated with each of the sequential frames of the sensor data; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
concatenating the encoded candidate sequences of sensor data, including the embedded position information Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
transforming the concatenated and encoded frames of sensor data to form transformed candidate sequences Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
generating a hierarchical structure that encapsulates navigation data of the particular candidate sequence. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 2,
Claim 2 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the different modalities comprise a Lidar, a camera and any of a GPS or IMU. This claim merely recites a further limitation on the obtaining candidate sequences of sensor data from different modalities, each sequence comprising a plurality of sequential frames from Claim 1 which was directed to insignificant extra solution activity (mere data gathering, as per MPEP 2106.05(g))
Regarding claim 3,
Claim 3 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 2 which includes an abstract idea (see rejection for claim 2). The additional limitations:
wherein transforming of the concatenated and encoded frames of sensor data comprises mean pooling. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 4,
Claim 4 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the system comprises an encoder/decoder system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
Regarding claim 5,
Claim 5 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which includes an abstract idea (see rejection for claim 4). The additional limitations:
wherein the system comprises one or more transformers. This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
Regarding claim 6,
Claim 6 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the determining of the particular candidate sequence as a match is based on a neural network. This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Regarding claim 7,
Claim 7 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the textual prompt is in natural language format This claim merely recites a further limitation on the obtaining a textual prompt from Claim 1 which was directed to insignificant extra solution activity (mere data gathering, as per MPEP 2106.05(g))
Regarding claim 8,
Claim 8 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the encoding of the candidate sequences and the textual prompt normalizes the sensor data and the textual prompt into a common feature space. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 9,
Claim 9 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the determining of a particular candidate sequence implements zero shot learning. This claim merely recites a further limitation on the determining a particular candidate sequence as a match between the transformed candidate sequences and the encoded textual prompt based on the contrastive loss from Claim 1 which was directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Regarding claim 10,
Claim 10 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the generating of a hierarchical structure implements a self attention layer, a cross attention layer, and a feed forward layer Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 11, this claim is rejected under the same rationale with claim 1 (as shown in the rejections above), because they are analogous claims.
Regarding claim 12, this claim is rejected under the same rationale with claim 2 (as shown in the rejections above), because they are analogous claims.
Regarding claim 13, this claim is rejected under the same rationale with claim 3 (as shown in the rejections above), because they are analogous claims.
Regarding claim 14, this claim is rejected under the same rationale with claim 4 (as shown in the rejections above), because they are analogous claims.
Regarding claim 15, this claim is rejected under the same rationale with claim 5 (as shown in the rejections above), because they are analogous claims.
Regarding claim 16, this claim is rejected under the same rationale with claim 6 (as shown in the rejections above), because they are analogous claims.
Regarding claim 17, this claim is rejected under the same rationale with claim 7 (as shown in the rejections above), because they are analogous claims.
Regarding claim 18, this claim is rejected under the same rationale with claim 8 (as shown in the rejections above), because they are analogous claims.
Regarding claim 19, this claim is rejected under the same rationale with claim 9 (as shown in the rejections above), because they are analogous claims.
Regarding claim 20, this claim is rejected under the same rationale with claim 10 (as shown in the rejections above), because they are analogous claims.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 3-8, 10-11, 13-18 and 20 are rejected under 35 U.S.C. 102(a) (2) as being anticipated by Mondal et al. (US 2023/0153352 A1).
Regarding claim 1, Mondal explicitly discloses:
one or more processors; and (Mondal, ¶[0050]: “In this example, the computing system 200 includes at least one processing unit 202, such as a processor, a microprocessor, a digital signal processor,”)
a memory storing instructions that, when executed by the one or more processors, cause the system to perform: (Mondal, ¶[0054]: “The computing system 200 may include a memory 210, which may include a volatile or non-volatile memory ( e.g., a flash memory, a random access memory (RAM), and/or a read-only memory (ROM)). The non-transitory memory 210 may store instructions 212 for execution by the processing unit 202, such as to carry out example embodiments described in the present disclosure.”)
obtaining a textual prompt; (Mondal, ¶[0103]: “At 502, a word-based query for a video is received”)
encoding the textual prompt; (Mondal, ¶[0104]: “At 504, the word-based query is encoded into a query representation.”)
obtaining candidate sequences of sensor data from different modalities, each sequence comprising a plurality of sequential frames, (Mondal, ¶[0063]: “The clips are each processed by a clip feature extractor 304 to output a respective set of frame features (shown as frame features-0 to frame features-n).”)
each of the candidate sequences being evaluated against the textual prompt based on a contrastive loss; (Mondal, ¶[0089]: “At 414, the alignment loss (e.g., a contrastive loss) between the video representation (generated at 404) and the annotation representation (generated at 406) is computed. Computation of this loss enables the video-processing branch 300 and the text-processing branch 310 of the video representation generator 110 to be trained to generate video and annotation representations that correspond to each other and that can be compared in the same, common representation space… This contrastive alignment may train the video representation generator 110 so that, for a sampled video having ground-truth annotation, the cosine similarity between the video representation and the annotation representation is maximized while the cosine similarity between the video representation and a different annotation representation is minimized.”)
encoding the candidate sequences of sensor data; (Mondal, ¶[0064]: “The set of frame features of each clip is processed by a clip encoder 306 to output a respective clip representation.”)
embedding position information, within the encoded candidate sequences, indicating relative timestamps associated with each of the sequential frames of the sensor data; (Mondal, ¶[0064]: “frame features-0 of clip-0 is processed by the clip encoder 306 to output clip representation-0. Each clip representation may be a vector, and clip representation-0 to clip representation-n may all have the same vector length. In addition to clip representations-0 to clip representations-n corresponding to clips-0 to clips-n, all frame features generated from the entire video (shown as frame features-all) are also inputted to the clip encoder 306 to output a global video context representation.”, ¶[0065]: “The clip representations-0 to clip representations-n and the global video context representation are processed by a video encoder 308 to output a video representation (e.g., a feature vector).”)
concatenating the encoded candidate sequences of sensor data, including the embedded position information; (Mondal, ¶[0065]: “The video encoder 308 may be any suitable neural network that is designed to compute a cross-attention feature between the global video context representation and the clip representations-0 to clip representations-n. The video encoder 308 may be further designed to compute self-attention using all the clip representations-0 to clip representations-n and then perform averaging pooling to obtain a pooled feature for all the clips. The pooled feature may then be concatenated with the cross-attention feature to output the video representation.”)
transforming the concatenated and encoded frames of sensor data to form transformed candidate sequences; (Mondal, ¶[0064]: “The clip encoder 306 may be any suitable neural network that is designed to process a sequence (in this case, a sequence of frame features that represents a sequence of frames) and aggregate the result. For example, the clip encoder 306 may be implemented using a temporal transformer and an attention based aggregation layer. For example, the clip encoder 306 may be implemented using any standard transformer, or variants of the transformer such as Sparse Transformer or Transformer-XL… Attention-based aggregation may be performed using standard attention-based aggregation”)
determining a particular candidate sequence as a match between the transformed candidate sequences and the encoded textual prompt based on the contrastive loss; and (Mondal, ¶[0076]: “The reconstructed clip representations are then compared with the query representation to identify the most closely matched clip representation (e.g., based on vector similarity). The clip corresponding to the clip representation that most matches the query representation is identified (e.g., identified by the clip index) as the relevant clip.”, ¶[0090]: “At 416, the alignment loss (e.g., a contrastive loss) between each clip representation (generated as part of generating the video representation) and each corresponding sentence representation (generated as part of generating the annotation representation) is computed (where a clip representation generated from clip-k corresponds to the sentence representation generated from sentence-k). Computation of this loss enables the video-processing branch 300 and the text-processing branch 310 of the video representation generator 110 to be trained to generate clip and sentence representations that correspond to each other and that can be compared in the same, common representation space… The contrastive alignment may train the clip encoder 306 and the sentence encoder 316 to maximize the cosine similarity between clip and sentence representations that correspond to each other, and to minimize the cosine similarity between clip and sentence representations that do not correspond.”)
generating a hierarchical structure that encapsulates navigation data of the particular candidate sequence. (Mondal, ¶[0041]: “Examples of the present disclosure use a hierarchical representing learning model to learn video-level representations (i.e., one representation per video), which may be stored and searched.”, ¶[0056]: “the video representation generator 110 may be implemented using a hierarchical representation learning model”)
Regarding claim 3, Mondal explicitly discloses:
wherein transforming of the concatenated and encoded frames of sensor data comprises mean pooling. (Mondal, ¶[0065]: “The video encoder 308 may be further designed to compute self-attention using all the clip representations-0 to clip representations-n and then perform averaging pooling to obtain a pooled feature for all the clips”)
Regarding claim 4, Mondal explicitly discloses:
wherein the system comprises an encoder/decoder system. (Mondal, ¶[0008]: “encoding the word-based query into a query representation using a trained query encoder”, ¶[0015]: “the grounding generated for each relevant video may include the clip index of the relevant clip in the relevant video, and the grounding module may include a classifier network or clip decoder network for predicting the clip index of the relevant clip using the query representation and the similar video representation as inputs.”)
Regarding claim 5, Mondal explicitly discloses:
wherein the system comprises one or more transformers. (Mondal, ¶[0064]: “the clip encoder 306 may be implemented using a temporal transformer and an attention based aggregation layer”)
Regarding claim 6, Mondal explicitly discloses:
wherein the determining of the particular candidate sequence as a match is based on a neural network. (Mondal, ¶[0076]: “The reconstructed clip representations are then compared with the query representation to identify the most closely matched clip representation (e.g., based on vector similarity). The clip corresponding to the clip representation that most matches the query representation is identified (e.g., identified by the clip index) as the relevant clip. Cross-attention may also be computed between the reconstructed clip representation and the query representation to obtain cross-attention features, which are processed by a linear layer of the network to obtain predicted start and end timestamps.”)
Regarding claim 7, Mondal explicitly discloses:
wherein the textual prompt is in natural language format. (Mondal, ¶[0080]: “As previously described, each video in the training dataset is annotated with a ground-truth annotation including a multi-sentence text description of the video (where each sentence corresponds to a clip of the video) and clip timestamps corresponding to each sentence of the annotation.”)
Regarding claim 8, Mondal explicitly discloses:
wherein the encoding of the candidate sequences and the textual prompt normalizes the sensor data and the textual prompt into a common feature space. (Mondal, ¶[0009]: “In an example of the preceding example aspect of the method, the query representation and the video representation may be in a common representation space, and identifying the one or more similar video representations may include: computing a similarity between the query representation and each video representation stored in a video representation storage; and identifying a defined number of video representations having a highest computed similarity to the query representation.”)
Regarding claim 10, Mondal explicitly discloses:
wherein the generating of a hierarchical structure implements a self attention layer, a cross attention layer, and a feed forward layer. (Mondal, ¶[0077]: “example, the clip classifier network may be implemented using a multilayer perceptron (MLP) network, which takes the video representation and query representation as inputs and outputs softmax values over all possible clip indexes”, ¶[0065]: “The video encoder 308 may be any suitable neural network that is designed to compute a cross-attention feature between the global video context representation and the clip representations-0 to clip representations-n. The video encoder 308 may be further designed to compute self-attention using all the clip representations-0 to clip representations-n and then perform averaging pooling to obtain a pooled feature for all the clips”, ¶[0082]: “high-level clip representation may be generated for a given clip by forward propagating the clip representation corresponding to the given clip and the global video context representation through the video encoder 308 while all other clip representations are zeroed.”)
Regarding claim 11, this claim is rejected under the same rationale with claim 1 (as shown in the rejections above) because they are analogous claims.
Regarding claim 13, this claim is rejected under the same rationale with claim 3 (as shown in the rejections above) because they are analogous claims.
Regarding claim 14, this claim is rejected under the same rationale with claim 4 (as shown in the rejections above) because they are analogous claims.
Regarding claim 15, this claim is rejected under the same rationale with claim 5 (as shown in the rejections above) because they are analogous claims.
Regarding claim 16, this claim is rejected under the same rationale with claim 6 (as shown in the rejections above) because they are analogous claims.
Regarding claim 17, this claim is rejected under the same rationale with claim 7 (as shown in the rejections above) because they are analogous claims.
Regarding claim 18, this claim is rejected under the same rationale with claim 8 (as shown in the rejections above) because they are analogous claims.
Regarding claim 20, this claim is rejected under the same rationale with claim 10 (as shown in the rejections above) because they are analogous claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 2, 12 are rejected under 35 U.S.C. 103 as being unpatentable over Mondal et al. (US 2023/0153352 A1) in view of Danna (US 2021/0403036 A1).
Regarding claim 2, Mondal explicitly discloses all the limitations of claim 1 (as shown in the rejections above).
Mondal fails to disclose:
wherein the different modalities comprise a Lidar, a camera and any of a GPS or IMU.
However, Danna discloses:
wherein the different modalities comprise a Lidar, a camera and any of a GPS or IMU. (Danna, ¶[0041]: “The sensor data may include data captured by one or more sensors including optical cameras, LiDAR, radar, infrared cameras, and ultrasound equipment, to name some examples”, ¶[0112]: “In particular embodiments, the navigation system 1146 may take as input any type of sensor data from, e.g., a Global Positioning System (GPS) module, inertial measurement unit (IMU), LiDAR sensors, optical cameras, radio frequency (RF) transceivers, or any other suitable telemetry or sensory mechanisms”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Mondal and Danna. Mondal teaches systems and methods for video retrieval and grounding by applying contrastive loss. Danna teaches systems and methods for encoding and searching scenario information to apply to autonomous vehicles. One of ordinary skill would have motivation to combine Mondal and Danna because MPEP 2143 sets forth the Supreme Court rationales for obviousness including: (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results; (E): “Obvious to try” choosing from a finite number of identified, predictable solutions, with a reasonable expectation of success; (F) Known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of the ordinary skill in the art.
Regarding claim 12, this claim is rejected under the same rationale with claim 2 (as shown in the rejections above) because they are analogous claims.
Claim(s) 9, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Mondal et al. (US 2023/0153352 A1) in view of Shvetsova (“Everything at Once – Multi-modal Fusion Transformer for Video Retrieval”).
Regarding claim 9, Mondal explicitly discloses all the limitations of claim 1 (as shown in the rejections above).
Mondal fails to disclose:
wherein the determining of a particular candidate sequence implements zero shot learning.
However, Shvetsova explicitly discloses:
wherein the determining of a particular candidate sequence implements zero shot learning. (Shvetsova, Pg. 6, Col. 1, ¶[1]: “Zero-shot Text-to-video Retrieval. We use MSR-VTT [54] and YouCook2 [57] datasets to evaluate the zero-shot text-to-video retrieval capability of our model”)
The combination of Mondal and Shvetsova are analogous art because they are in the same field of training time series data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, having the teachings of Mondal and Shvetsova before them, to modify the teachings of Mondal to include the teachings of Shvetsova implementing a zero shot learning in determining the candidate sequence to enable the system to identify sensor-data sequences corresponding to previously unseen textual prompts without requiring additional labeled training data or task-specific retraining, thereby improving retrieval flexibility and reducing the time and computational resources associated with preparing labeled examples and retraining the model.
Regarding claim 19, this claim is rejected under the same rationale with claim 9 (as shown in the rejections above) because they are analogous claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMY TRAN whose telephone number is (571)270-0693. The examiner can normally be reached Monday - Friday 7:30 am - 5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AMY TRAN/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126