Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 05/18/2026 have been fully considered and they are persuasive. However, a newly found reference Ciecko (US 2019/0272336) reads on the amended portions of the current set of claims as detailed below.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 6-9, 11, 13, 15-16, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Barut (US patent number 12,045,288) in view of Ciecko (US 2019/0272336).
As per claims 1, 8, and 18 Barut teaches, a method and method comprising: generating, using one or more multi-modal language models (Barut, fig.1 170, Multi-modal transformer model ), one or more embeddings associated with one or more portions of one or more images (Barut, fig.2 image data 202 to 208 object detector to 212 object embeddings represent one or more embeddings associated with one or more portions of one or more images as it comes in serial form from the image data); determining, based at least on a query, at least an embedding of the one or more embeddings that is associated with a portion of the one or more portions (Barut, fig.2 Natural language 220 represents a query then to word embedding data 222, and the word would represent the portion); determining, based at least on the embedding, that an image of the one or more images is associated with the query (Barut, fig.2 “Object/Entity matching” 218); and sending, to a client device associated with the query, at least one of image data representative of the image or data indicating a location of the portion within the image (Barut, fig.1 Relative location data 121 represents data indicating a location of the portion within the image).
Barut doesn’t clearly teach, determining, based at least on a query that indicates at least an object and positional information for the object, and determining, based at least on the embedding portion, that the object is represented at a location within an image of the one or more images, determining, based at least on the location within the image corresponding to the positional information indicated by the query, that the image is associated with the query.
Ciecko teaches, based at least on a query that indicates at least an object and positional information for the object (Ciecko, fig.5A 500 “query” and 540 “location and object type info”. While the specification of applicant has a more detail view of positional information it is also a location information as several examples are given in the applicant’s specification, and since no details have been given in the claim language positional can be interpreted as location and following 541 “visual image info” could also be interpreted as applicant’s examples in the specification that aren’t in the claim language. In other words for example, positional in the art could mean, pixel alignment, however according to the specification of applicant it is more of a location of the image object as well as some visual info, example third car from the left. However, no specifics in the claim language), and determining, based at least on the embedding portion, that the object is represented at a location within an image of the one or more images (Ciecko, fig.5A 545 location is being found with 550), determining, based at least on the location within the image corresponding to the positional information indicated by the query, that the image is associated with the query (Ciecko, fig.5A, 555 the files are then returned based on the query).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Barut with those of Ciecko of having the ability to have a query that identifies Object, location and other image information as part of the query and have those images be returned.
The motivation would have been to improve system time performance as taught by Ciecko ¶[0010] “In this manner, only information of particular interest to the device user is displayed for viewing, and the information retrieval process is accelerated.”.
As per claims 2 and 9, Barut in view of Ciecko teaches, the method of claim 1, wherein: further comprising receiving, from the client device, second data representative of the query, wherein the query includes one or more first words indicating the object and one or more second words indicating the positional information for the object (Ciecko, fig.5A, 542 matching location and image info would represent second words indicating the positional information for the object)
As per claims 3, 11, 13, and 19 Barut in view of Ciecko teaches, the method of claim 1, further comprising: receiving, from the client device, a selection associated with the image (Barut, fig.1 172 selected object data); determining, based at least on the selection and using at least one of the embedding or a second embedding associated with the image, a third embedding of the one or more embeddings that is associated with a second portion of the one or more portions (Barut, fig.1 123 word embedding and seeing multiple embeddings); determining, based at least on the second embedding, that a second image of the one or more images is associated with the query; and sending, to the client device, at least one of second image data representative of the second image or second data indicating a second location of the second portion within the second image (Barut, fig.1-4 as the system loops around there will be a second image and a second location of each image. In other words, as different queries get used this would inherently happen, and see fig.1 123 for multiple embeddings. Anything to do with second locations the system would loop and do the tasks creating second locations).
As per claim 4, Barut in view of Ciecko teaches, the method of claim 1, wherein the portion of the image is associated with a first object indicated by the query, and wherein the method further comprises: determining, based at least on a second object indicated by the query, at least a second embedding of the one or more embeddings that is associated with a second portion of the one or more portions; and determining, based at least on the second embedding, that the image of the one or more images is again associated with the query (Barut, claim 7 “determining second image data representing at least a second object,…” This represents second embedding in the second object in a second portion).
As per claims 6 and 15, Barut in view of Ciecko teaches, the method of claim 1, further comprising: segmenting the one or more images into the one or more portions (Barut, fig.1 object detector 115 would be segmenting ); determining one or more locations associated with the one or more portions within the one or more images (Barut, fig.1 121 relative location data); determining one or more identifiers that associate the one or more portions with the one or more images (Barut, (50) “The speech processing enabled device may also send metadata (e.g., including device identifiers, device type data, contextual data, IP address data, room location data, etc.) to the orchestrator 330.“ This represents determining one or more identifiers ); and storing, in one or more databases, the one or more embeddings, second data representative of the one or more locations, and third data representative of the one or more identifiers (Barut, (30) “Additionally, memory 103 may store data described herein, such as one or more model parameters, query data, result data, training data, etc.” This represents the storing).
As per claims 7 and 16, Barut in view of Ciecko teaches, the method of claim 1, further comprising: determining a second embedding based at least on at least one of an image, text, or an inputted embedding from the query, wherein the determining the at least the embedding from the one or more embeddings is based at least on the second embedding (Barut, fig.1 123 Word Embedding Data 123, this represents embedding based at least on at least one of text).
Allowable Subject Matter
Claims 5, 10, 12, 14, 17 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SANTIAGO GARCIA whose telephone number is (571)270-5182. The examiner can normally be reached Monday-Friday 9:30am-5:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SANTIAGO GARCIA/Primary Examiner, Art Unit 2673
/SG/