DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments filed on 06/18/2026, with respect to claim(s) 1-5, 8, 10-13, and 22-27, have been considered but they are not persuasive. Regarding independent claims 1 and 22 (and their respective deponent claims) Applicant argues that “As amended, claim 1 recites providing both (i) imaging data and (ii) query information to the model. Pirk does not disclose such configuration. Rather, Pirk describes determining object regions from image data and comparing embeddings associated with those regions to a query embedding (e.g., [0011], [0075]-[0077]). Thus, Pirk does not disclose providing the model with both imaging data and query information, as now recited. Additionally, claim 1 recites that the model has been updated using training data including natural language descriptions of relationships between parts generated using a generative transformer network. Pirk does not disclose such training. For at least these reasons, the cited reference does not anticipate claim 1 or its dependent claims, and for similar reasons, independent claim 22 or its dependent claims. Withdrawal of the rejection under 35 U.S.C. § 102 is respectfully requested.” (please see Remarks, page 11).
Examiner respectfully disagrees, First of all, as previously explained and noted by the Examiner that the way claims 1 and 22, are drafted does not requires processing of interaction, obtaining, segmenting, providing, receiving and cause performance of one or more control operations steps, for instance currently amended claim 1, recites “one or more circuits to:
obtain a query, indicating at least one of:
(i) an objective or goal to be accomplished;
(ii) an interaction with a three-dimensional (3D) object or between the 3D object and a second 3D object
(iii) an action to be performed
(iv) a semantic attribute of the 3D object or a part thereof; or
(v) a spatial relationship between two or more parts of the 3D object”, Hence as can be seen from the claim language that only one condition/option needs to be met in order to reject independent claims 1 and 22 (and their respective deponent claims). In this scenario Examiner selected (i) to reject the claims 1 and 22, and rest of claim limitations are not given a patentable weight, since rest of claim limitations relates to other conditions/options indicated in claims 1 and 22 and because of alternative language “or” said options/limitations are not required. Although Examiner cited paragraphs to some claim limitations for completeness and/or to show that even Applicant were to change “or” to “and”, the cited references would still reads on the rest of claim limitations. For instance, as Pirk paragraph 11, discloses “the query embedding can be determined based on voice input and/or based on an image of the target object. For example, “red mug” in voice input of “retrieve the red mug” can be mapped to a given point in the embedding space (e.g., through labeling of the embedding space with semantic text labels after training). Also, for example, a user can point to a “red mug” and provide a visual, verbal, and/or touch command to retrieve similar objects. Image(s) of the “red mug” can be captured (or cropped from a larger image using object recognition techniques), using the user's pointing as a queue”, Hence as can be seen from above passage that Robot first receives query red mug via voice input and then robot captures images of the red cup in order to “retrieve the red mug”. Therefore Pirk reference still reads on the argued limitations, i.e., providing the model with both imaging data and query information, as presented by the Applicant. Examiner suggests Applicant to remove alternative language “or” from the claim and also further elaborate on segmentation in order to overcome the cited references.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8, 10-13, and 22-27, is/are rejected under 35 U.S.C. 103 as being unpatentable over Pirk (US PGPUB 2021/0334599 A1) and further in view of Banipal (US PGPUB 2022/0414174 A1)
As per claim 1, Pirk discloses at least one processor (Pirk, Fig. 1:190) comprising:
one or more circuits (Pirk, paragraphs 10, 56 and 74) to:
obtain a query, indicating at least one of:
(i) an objective or goal to be accomplished ((Pirk, paragraphs 11, 24-25, 37 and 75, discloses the query embedding can be determined based on voice input and/or based on an image of an object. For example, “red mug” in voice input of “retrieve the red mug” can be mapped to a given point in the embedding space (e.g., through labeling of the embedding space with semantic text labels after training),
Furthermore please note although Examiner has cited related paragraphs of the cited reference to rest of claim limitations in a case Applicant to remove alternative claim language and to incorporate rest of claim limitations into the pending claim, however said limitations are not given a patentable weight, since these limitations are not part of the claim and are as follows:
(ii) an interaction with a three-dimensional (3D) object or between the 3D object and a second 3D object
(iii) an action to be performed
(iv) a semantic attribute of the 3D object or a part thereof; or
(v) a spatial relationship between two or more parts of the 3D object;
Obtain imaging data representing the 3D object (Pirk, paragraphs 11, and 40-41);
process the imaging data according to the query to segment the 3D object in the imaging data (Pirk, Fig. 1:152:142, and paragraphs 7, 11, 41-43, discloses Image(s) of the “red mug” can be captured (or cropped from a larger image using object recognition techniques)), the segmenting including:
providing, to a model, (i) the imaging data representing the 3D object (Pirk, Fig. 1:142:152, and paragraphs 13, 41-43) and (ii) query information including the query or one or more requests specified by or generated based at least on the query (Pirk, paragraphs 4, 7, 11, and 78), and
receiving, from the model, an identification of one or more region of the 3D object in the imaging data (Pirk, paragraphs 13, discloses processing the first image using an object recognition model to identify a plurality of first object regions in the first image)
cause performance of one or more control operations of an autonomous or semi-autonomous machine to act based at least on the identification of the one or more regions of the 3D object (Pirk, paragraphs 11, and 78, discloses robot can be controlled to interact with a target object by: determining a query embedding, in an embedding space of the object-contrastive model; processing a robot image, from a vision component of the robot, using the object-contrastive model; determining, based on the processing, a target object in a current environment of the robot; and controlling the robot to interact with the target object).
Pirk does not explicitly disclose model having been updated using training data including natural language descriptions of relationships between a plurality of parts of one or more 3D objects, the natural language descriptions generated at least using a generative transformer network that produces text in response to queries;
Banipal discloses the model having been updated using training data including natural language descriptions of relationships between a plurality of parts of one or more 3D objects (Banipal, paragraphs 43-44, discloses graphical representation generation mod 308 encompasses a generative adversarial network (GAN) tasked with rendering three-dimensional objects matching a set of descriptive terms trained using sets of images paired with descriptive text describing attributes in each image. This GAN is then provided the attributes determined at S265 as the basis for generating a three-dimensional object matching the attributes. In this simplified embodiment, graphical representation generation mod 308 generates a table with four legs, consistent with the keyword(s) previously parsed from the text-based query.), the natural language descriptions generated at least using a generative transformer network that produces text in response to queries (Banipal, paragraphs 41, 43-44 and 56);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Pirk teachings by implementing a GAN to the system of Pirk, as taught by Banipal.
The motivation would be to improve searchability with generated images based on complimentary attributes deducted from the inventory products description (paragraph 54), as taught by Banipal.
As per claim 2, Pirk in view of Banipal further discloses the at least one processor of claim 1, wherein the identification comprises a segmentation mask (Pirk, Fig. 2A:210A:250).
As per claim 3, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the identification comprises a pointwise label or set of pixels (Pirk, paragraph 11).
As per claim 4, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the causing performance of the one or more control operations of the autonomous or semi-autonomous machine comprises generating an instruction to interact with one or more parts of the 3D object (Pirk, paragraphs 11, 25 and 78).
As per claim 5, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the causing performance of the one or more control operations of the autonomous or semi-autonomous machine comprises generating an instruction to interact with the 3D object and the second 3D object (Pirk, paragraph 4, discloses generate or identify a query embedding in an embedding space of the trained model (where the query embedding represents an embedding of rich feature(s) of target object(s) to be interacted with by the robot)), based at least on one or more attributes of the 3D object or the second 3D object identified via segmenting (Pirk, paragraphs 9, 11, and 37).
As per claim 8, Pirk in view of Banipal further discloses the at least one processor of claim 1, wherein the plurality of parts of the one or more 3D objects are obtained using a dataset of 3D objects annotated with hierarchical 3D part information (Banipal, paragraph 54).
As per claim 10, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the generative transformer network provides the natural language descriptions in response to queries indicating spatial relationships, geometric attributes, functional properties, or semantic attributes of the plurality of parts of the one or more 3D objects (Banipal, paragraphs 41 and 54).
As per claim 11, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the causing performance of the one or more control operations of the autonomous or semi-autonomous machine comprises generating an instruction and transmitting the instruction to a control system of the autonomous or semi-autonomous machine to cause an interaction with the 3D object based at least on the identification of the one or more regions of the 3D object (Pirk, paragraphs 4, 11 and 78).
As per claim 12, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein obtaining the query comprises:
receiving, prior to segmenting the 3D object in the imaging data, an action to be performed with respect to the 3D object (Pirk, paragraph 4, discloses a trained model in processing vision data captured by a vision component of a robot, generate embeddings based on the processing, and control the robot based at least in part on the generated embeddings. For instance, some of those implementations can: generate or identify a query embedding in an embedding space of the trained model (where the query embedding represents an embedding of rich feature(s) of target object(s) to be interacted with by the robot)); and
generating the query based at least on the action (Pirk, paragraphs 4, 11 and 25).
As per claim 13, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the at least one processor is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational Al operations; a system implementing one or more large language models (LLMs); a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Pirk, paragraph 53).
As per claim 22, please see the analysis of claim 1.
As per claim 23, Pirk n view of Banipal further discloses the method of claim 22, wherein the method further comprises transmitting the instruction to a control system capable of performing the one or more control operation using the autonomous or semi-autonomous machine (Pirk, paragraphs 4, 11 and 78).
As per claim 24, Pirk n view of Banipal further discloses the method of claim 22, wherein the instruction is to cause the autonomous or semi-autonomous machine to interact with one or more parts of the 3D object in performing the one or more control operations (Pirk, paragraphs 11, 25 and 78).
As per claim 25, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the identification indicates one or more unlabeled parts or subparts of the 3D object or one or more unlabeled regions of the 3D object in the imaging data (Banipal, paragraph 51).
As per claim 26, Pirk n view of Banipal further discloses the at least one processor of claim 1, wherein the query information comprises one or more facts specified by or generated based at least on the query, and wherein the one or more facts comprise:
a unary fact specifying a geometric or functional property of a query part of the 3D object (Banipal, paragraph 41); and
a binary fact specifying a spatial relationship between the query part and an anchor part of the 3D object (Banipal, paragraph 41),
wherein the one or more regions identified by the model include a region corresponding to the query part (Banipal, paragraph 41).
As per claim 27, Pirk n view of Banipal further discloses the method of claim 22, wherein the identification indicates one or more unlabeled parts or subparts of the 3D object or one or more unlabeled regions of the 3D object in the imaging data (Banipal, paragraph 51).
Allowable Subject Matter
Claims 14, 16-17, and 21, are allowed.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYED Z HAIDER whose telephone number is (571)270-5169. The examiner can normally be reached MONDAY-FRIDAY 9-5:30 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, SAM K Ahn can be reached at 571-272-3044. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SYED HAIDER/Primary Examiner, Art Unit 2633