Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the communication filed on May 26, 2026.
Response to Amendment
Applicant’s amendment filed on May 26, 2026 with respect to claims 1-20 has been received, entered in to the record and considered.
As a result of the amendment, claims 1-3, 6, 8, 10-12, 15, 17 and 19-20 has been amended.
Claims 1-20 remain pending in this office action.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Danna (US 2021/0403036 A1), in view of Ferroni (US 2024/0257536 A1)).
As per claim 1, Danna discloses:
- a method, comprising (method and system for encoding and searching, Para [0001]),
- obtaining, by at least one processor, natural language text data associated with a query, the query comprising one or more semantic elements, wherein the natural language text data is provided at a graphical user interface (GUI) that includes a field configured to receive the natural language text data as an input (receiving a natural language query with a scenario (i.e., semantic element) and obtain data associated with the query, Para [0003], [0021], [0091], [0059], [0085], [0089], Fig. 8, item 806, natural language search module (i.e., GUI for natural language text input), Fig. 6 and 7, by a processor, Fig. 12, item 1202),
- extracting, by the at least one processor, an embedding representing the one or more semantic elements (embedding for the scenario (i.e., semantic element) is generated based on the query, Para [0003], [0041], [0053], [0059], Fig. 4, item 410, by a processor, Fig. 12, item 1202), the embedding corresponding to a first location that is associated with the query in a shared latent space (embedding reflects the location or scenario in the vector space (i.e., latent space), Para [0003], [0004], [0041], [0044], [0045], [0053], Fig. 4, item 414, 412, 418, Fig. 6A-6B),
- comparing, by the at least one processor, the embedding to at least one predetermined embedding of a set of predetermined embeddings comparing embedding in the shared latent space, the at least one predetermined embedding corresponding to a second location in the shared latent space that is associated with an image that is labeled and corresponds to the at least one predetermined embedding (comparing with the set of embedding with other embedding and find similarity in different scenarios or location in a vector space (i.e., latent space), Para [0045], [0053], by a processor, Fig. 12, item 120, Para [0054, Fig. 2, item 210, Para [0005], [0013], [0050], Fig. 4, item 412, 416, Para [0060]), and Para [0102], The machine-learning models may be trained using any suitable training algorithm, including supervised learning based on labeled training data, unsupervised learning based on unlabeled training data, and semi-supervised learning based on a mixture of labeled and unlabeled training data),
- wherein the set of predetermined embeddings in generated prior to obtaining the natural language text data associated with the query (Para [0058], [0060], Fig. 4, item 412, generating and storing embedding data in embedding store 412, (i.e., predetermined embedding prior to natural language query),
- and providing, by the at least one processor, data associated with the GUI to cause a display device to display the GUI representing a set of images comprising an image corresponding to the at least one predetermined embedding (representing (i.e., displaying) set of images corresponding to embedding, Para [0041], [0045], [0059], [0061], by a processor, communication interface (i.e., GUI), Fig. 12, item 1202),
- wherein the one or more semantic elements at least in part correspond to one or more objects represented by the image (sematic element correspond to an object, Para [0056], [0110]).
Danna does not explicitly disclose wherein the shared latent space is generated by training a text encoder and an image encoder on paired text and labeled images such that embeddings from the text encoder and the image encoder are correlated in the shared latent space; selecting, by the at least one processor, the at least one predetermined embedding based on a degree of similarity between the embedding representing the natural-language text data and the at least one predetermined embedding representing the image. However, in the same field of endeavor Ferroni in an analogous art disclose wherein the shared latent space is generated by training a text encoder and an image encoder on pairs of descriptive text indication one or more features of images and corresponding images such that embeddings from the text encoder and the image encoder are correlated in the shared latent space (Para [0118], During training, the two embeddings produced by the text and image branches are optimized such that each text-image pair from the training dataset map to the same point in the embedding space), selecting, by the at least one processor, the at least one predetermined embedding based on a degree of similarity between the embedding representing the natural-language text data and the at least one predetermined embedding representing the image (Para [0007], [0142], identifying embeddings (i.e., selecting embedding) based on similarity between embeddings (e.g. text embeddings, image embeddings, combinations of text/image embeddings, image/image embeddings, text/text embeddings).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the training text and image data on pairs and find the embedding similarity as the means to process semantic embedding as taught by Ferroni as the means to process semantic embedding in an input query and finding the similarity between embeddings in Danna, (Danna, Para [0003], Para [0021], [0059], Fig. 6 and 7, Ferroni, Para [0007], [0142]). Danna and Ferroni are analogous prior art since they both deal with processing image and text embedding and finding similarity in embeddings. A person of the ordinary skill in the art would have been motivated to make aforementioned modification to improve safety while driving a vehicle. This is because one aspect of Danna invention is to enable the vehicle to determine its surroundings so that it may safely navigate to target destinations or assist a human driver, as described at least in Para [0040]. Training text and image data on pairs and find the embedding similarity is part of this process. However, Danna doesn’t specify any particular manner in which training text and image data on pairs and find the embedding similarity as the means to process semantic embedding. This would have lead one of the ordinary skill in the art to seek and recognize training text and image data on pairs and find the embedding similarity as the means to process semantic embedding as taught by Ferroni. Ferroni describes how their techniques streaming data mining with text-image joint embeddings including for use in an autonomous vehicle (AV) to accurately, efficiently, and expeditiously search, and identify image data, (Para [0001], 0005]), as desired by Dana.
As per claim 2, rejection of claim 1 is incorporated, and further Danna discloses:
- providing, by the at least one processor, the data associated with the input to a text encoder to cause the text encoder to generate the embedding (scenario encoder (i.e., text encoder) to generate embedding, Para [0003], Fig. 2, item 208, Fig. 4, item 410, Para [0044], [0056]).
As per claim 3, rejection of claim 1 is incorporated, and further Danna discloses:
- obtaining, by the at least one processor, the set of predetermined embeddings from a database, the set of predetermined embeddings generated by the image encoder based on images (retrieving (i.e., obtaining) embedding and image encoder data from data store 220 (i.e., database), Para [0048], [0053], Fig. 2, item 220, by a processor, Fig. 12, item 1202).
As per claim 4, rejection of claim 3 is incorporated, and further Danna discloses:
- wherein the image encoder is configured to receive data associated with images generated by at least one sensor supported by at least one vehicle as input and provide embeddings associated with a latent space as output (image encoder image data generated by a sensor in a vehicle, Para [0051]).
As per claim 5, rejection of claim 1 is incorporated, and further Danna discloses:
- selecting, by the at least one processor, at least one second predetermined embedding of the set of predetermined embeddings based on a second degree of similarity between the at least one second predetermined embedding and other predetermined embeddings of the set of predetermined embeddings (second embedding based on second similarity, Para [0013], [0046], second level similarity, Para [0061).
As per claim 6, rejection of claim 1 is incorporated, and further Danna discloses:
- obtaining, by the at least one processor, data associated with a second input, the second input indicating selection of a different image of the set of images represented by the GUI (query with the second scenario (i.e., second input), Para [0053], [0054]),
- determining, by the at least one processor, at least one second predetermined embedding based on the selection of the different image represented by the GUI (second embedding based on second scenario, Para [0053], [0054]),
- and selecting, by the at least one processor, at least one third predetermined embedding based on a degree of similarity between the at least one second predetermined embedding and embeddings of the set of predetermined embeddings (selecting third embedding based similarity between embeddings, Para [0046], [0053], [0054]).
As per claim 7, rejection of claim 6 is incorporated, and further Danna discloses:
- wherein the second input indicates selection of the different image that corresponds to a different embedding of the set of predetermined embeddings (first query scenario and second query scenario are different, Para [046], [0053], [0065]).
As per claim 8, rejection of claim 1 is incorporated, and further Danna discloses:
- wherein the embedding and the set of predetermined embeddings comprise vector representations corresponding to one or more features in the shared latent space (embedding representing in a vector space, Para [0004], [0013], [0053]).
As per claim 9, rejection of claim 1 is incorporated, and further Danna discloses:
- determining that the degree of similarity between the embedding and the at least one predetermined embedding satisfies a similarity threshold (embedding similarity satisfied threshold, Para [0005], [0006]),
- and selecting the at least one predetermined embedding based on the degree of similarity satisfying the similarity threshold (satisfying similarity threshold, Para [0054], [0060]).
As per claim 10-18,
Claims 10-18 are system claims corresponding to method claims 1-9 respectively and rejected under the same reason set forth to the rejection of claims 1-9 above.
As per claims 19-20,
Claims 19-20 are computer readable medium claims corresponding to method claims 1-2 respectively and rejected under the same reason set forth to the rejection of claims 1-2 above.
Response to Arguments
8. Applicant’s arguments with respect to claims 1-20 have been considered but they are not deemed to be persuasive.
In response to the applicant’s argument in page 10, applicants argued that, the cited references do not teach at least "selecting, by the at least one processor, the at least one predetermined embedding based on a degree of similarity between the embedding representing the natural-language text data and the at least one predetermined embedding representing the image," as recited in amended claim 1.
Examiner respectfully response that, Ferroni teaches selecting, by the at least one processor, the at least one predetermined embedding based on a degree of similarity between the embedding representing the natural-language text data and the at least one predetermined embedding representing the image in Para [0007], [0142], identifying embeddings (i.e., selecting embedding) based on similarity between embeddings (e.g. text embeddings, image embeddings, combinations of text/image embeddings, image/image embeddings, text/text embeddings).
In response to applicant’s argument tin page 11, applicants argued that, neither of the cited references teach or describe comparing the "embedding representing the natural-language text data" to "a set of predetermined embeddings," where "the set of predetermined embeddings is generated prior to obtaining the natural-language text data associated with the query," as recited in amended claim 1
Examiner disagree, and respectfully response that, Danna teaches comparing scenario embedding with embedding stored in embedding storage 412, see, Para [0045], line 27-40, Para [0046], line 37-45, Para [0054], [059], Fig. 4, item 412, 416, 408.
Beside Danna, Ferroni also teaches, comparing the "embedding representing the natural-language text data" to "a set of predetermined embeddings," where "the set of predetermined embeddings is generated prior to obtaining the natural-language text data associated with the query, in Para [0019], [0140], [0167], [0162].
Therefore, examiner firmly believe that, Danna and Schulter alone or in combination reasonably teaches the argued limitation and claim 1, 10 and 19 as claimed.
In response to applicant’s argument in page 11, references do not teach or describe "selecting, by the at least one processor, at least one third predetermined embedding based on a degree of similarity between the at least one second predetermined embedding and embeddings of the set of predetermined embeddings," as recited in dependent claim 6.
Examiner disagree and respectfully response that, Danna teaches selecting, by the at least one processor, at least one third predetermined embedding based on a degree of similarity between the at least one second predetermined embedding and embeddings of the set of predetermined embeddings, at least in Para [0053],…A greater distance metric determined between the first embedding and a third embedding, which can be translated to a lower degree of similarity between the first scenario and a third scenario associated with the third embedding, indicates a higher degree of similarity between the first and second scenarios compared to the first and third scenarios. The embedding module 208 can store the embeddings in the data store 220.
Therefore, examiner firmly believe that, Danna and Ferroni alone or in combination reasonably teaches the argued limitation and claim1 as a whole, as claimed.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOHAMMED R UDDIN whose telephone number is (571)270-3138. The examiner can normally be reached M-F: 9:00 AM-5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Apu Mofiz can be reached at (571) 272-4080. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MOHAMMED R UDDIN/Primary Examiner, Art Unit 2161